Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To extract an HTML table in Python, load the file or webpage with
HTMLDocument, select a table with query_selector("table") or collect tables with
get_elements_by_tag_name(“table”), and save the extracted content as CSV, TXT, or separate HTML files.
HTML tables commonly store reports, schedules, comparison data, product specifications, and reference values. Aspose.HTML for Python via .NET parses the source into a DOM tree, allowing an application to select table elements and read their text without processing the HTML as an unstructured string.
This article demonstrates how to extract a local HTML table to CSV, collect webpage tables in a text report, and preserve individual tables as separate HTML files.
The following example uses the source file
product-table.html. It selects the first table, reads both header and data cells, normalizes their text, and saves the rows as product-table.csv.
To extract an HTML table to CSV:
HTMLDocument in a context manager.query_selector("table") to select the first table.get_elements_by_tag_name("tr").th and td cells from every row.csv.writer. 1# Extract an HTML table to CSV in Python
2
3import csv
4import aspose.html as ah
5
6# Load the HTML file and extract the first table
7with ah.HTMLDocument("data/product-table.html") as document:
8 table = document.query_selector("table")
9
10 if table is None:
11 raise RuntimeError("The HTML document contains no table.")
12
13 extracted_rows = []
14 rows = table.get_elements_by_tag_name("tr")
15
16 # Extract header and data cells
17 for row in rows:
18 cells = row.query_selector_all("th, td")
19 values = [" ".join(cell.text_content.split()) for cell in cells]
20
21 if values:
22 extracted_rows.append(values)
23
24# Save the table as CSV
25with open("product-table.csv", "w", newline="", encoding="utf-8") as file:
26 csv.writer(file).writerows(extracted_rows)The CSV file contains the table header followed by three data rows. Using csv.writer ensures that values containing commas, quotation marks, or line breaks are escaped according to CSV rules.
Pass an absolute URL to HTMLDocument when tables should be extracted directly from a webpage. The next example loads the Aspose.HTML Supported File Formats page, processes every <table> in the returned DOM, and saves the rows as a readable text report.
To extract HTML tables from a webpage:
HTMLDocument constructor.get_elements_by_tag_name("table").tr rows and their th and td cells in every table..txt file. 1# Extract HTML tables from a webpage to TXT in Python
2
3import aspose.html as ah
4
5page_url = "https://docs.aspose.com/html/python-net/supported-file-formats/"
6
7# Load the webpage and extract its tables
8with ah.HTMLDocument(page_url) as document:
9 tables = document.get_elements_by_tag_name("table")
10
11 if tables.length == 0:
12 raise RuntimeError("The webpage contains no tables.")
13
14 lines = []
15
16 # Extract rows and cells
17 for table_index, table in enumerate(tables, start=1):
18 lines.append(f"Table {table_index}")
19 rows = table.get_elements_by_tag_name("tr")
20
21 for row in rows:
22 cells = row.query_selector_all("th, td")
23 values = [" ".join(cell.text_content.split()) for cell in cells]
24
25 if values:
26 lines.append(" | ".join(values))
27
28 lines.append("")
29
30# Save the tables as a text report
31with open("webpage-tables.txt", "w", encoding="utf-8") as file:
32 file.write("\n".join(lines))The report labels each table and writes one source row per line. Selectors and table structure vary between websites, so inspect the loaded DOM and narrow the extraction scope when a page contains unrelated layout or comparison tables.
Use outer_html when the original table markup should be preserved instead of reducing each row to plain text values. The following example loads a webpage, finds every <table>, and saves each table as a separate .htm document.
To save extracted tables as HTML files:
HTMLDocument.<table> elements with get_elements_by_tag_name("table").outer_html from each table to preserve its complete element markup.HTMLDocument from the extracted table markup. 1# Save HTML tables from a website as separate files in Python
2
3import os
4import aspose.html as ah
5
6# Prepare the output directory
7output_dir = "output"
8os.makedirs(output_dir, exist_ok=True)
9
10# Load the webpage and collect its tables
11with ah.HTMLDocument("https://docs.aspose.com/html/net/edit-html-document/") as document:
12 tables = document.get_elements_by_tag_name("table")
13
14 if tables.length > 0:
15 for i, table in enumerate(tables):
16 file_name = f"table{i}.htm"
17 file_path = os.path.join(output_dir, file_name)
18
19 # Save the table as a separate HTML document
20 with ah.HTMLDocument(table.outer_html, file_path) as table_document:
21 table_document.save(file_path)
22 else:
23 print("No tables found in the document.")The example preserves each table element, its rows, cells, attributes, and inline styles. CSS rules, fonts, images, or other resources defined outside the table are not copied automatically, so an extracted table can look different from the original webpage.
| Task | Recommended approach | Use when |
|---|---|---|
| Select the first table | document.query_selector("table") | The document contains one main data table. |
| Collect every table | document.get_elements_by_tag_name("table") | Every table in the document must be processed. |
| Select tables from main content | Select main, then call main.get_elements_by_tag_name("table") | Navigation or related page regions may contain unrelated tables. |
| Read table rows | table.get_elements_by_tag_name("tr") | Header and body rows should both be included. |
| Read header and data cells | row.query_selector_all("th, td") | Column labels must be preserved with ordinary cell values. |
| Select a table by ID or class | document.query_selector("table#report") or another focused selector | The page contains several tables but only one is required. |
| Preserve table markup | Save table.outer_html in a new HTMLDocument | Table elements and attributes must remain available in the output. |
| Issue | Cause | Recommended action |
|---|---|---|
| No table is found | The loaded document contains no <table> or the selector does not match it. | Inspect the returned DOM and begin with a broad table selector. |
| The wrong table is extracted | query_selector("table") returns only the first matching table. | Select a stable ID, class, or content container, or process every table. |
| CSV columns are shifted | The source uses rowspan or colspan. | Add application-specific logic that expands spanning cells into a rectangular grid. |
| Extra whitespace appears in cells | Cell text contains indentation, line breaks, or nested inline elements. | Normalize text_content before saving it. |
| An extracted HTML table loses its appearance | Page-level or external CSS is outside the saved table markup. | Include the required styles or resources in the new document. |
| A browser shows a table but extraction finds none | The table is absent from the DOM available after loading. | Inspect the loaded document and verify how the page produces the table. |
Yes. Pass the absolute URL to HTMLDocument, locate the required table elements, and read their rows and cells. Network access and the HTML returned by the server determine which tables are available.
Select both th and td cells when processing rows. The first example writes header values as the first row of the CSV file.
It preserves the text of each physical cell but does not expand colspan or rowspan into a normalized data grid. Add grid-processing logic when the output must reproduce the visual table layout.
Yes. The Aspose.HTML part of the workflow selects the table and extracts its cell values. Replace the CSV or TXT writing step with the JSON serializer or database client required by your application.
Yes. Read the table’s outer_html, create a new HTMLDocument from that markup, and save it to an .html or .htm file. Include any required page-level styles separately when the output should retain the original appearance.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.