Extract HTML Tables in Python

To extract an HTML table in Python, load the file or webpage with HTMLDocument, select a table with query_selector("table") or collect tables with get_elements_by_tag_name(“table”), and save the extracted content as CSV, TXT, or separate HTML files.

HTML tables commonly store reports, schedules, comparison data, product specifications, and reference values. Aspose.HTML for Python via .NET parses the source into a DOM tree, allowing an application to select table elements and read their text without processing the HTML as an unstructured string.

This article demonstrates how to extract a local HTML table to CSV, collect webpage tables in a text report, and preserve individual tables as separate HTML files.

Extract an HTML Table to CSV in Python

The following example uses the source file product-table.html. It selects the first table, reads both header and data cells, normalizes their text, and saves the rows as product-table.csv.

To extract an HTML table to CSV:

  1. Load the local HTML file with HTMLDocument in a context manager.
  2. Call query_selector("table") to select the first table.
  3. Collect its rows with get_elements_by_tag_name("tr").
  4. Select both th and td cells from every row.
  5. Normalize the text in each cell and collect non-empty rows.
  6. Write the extracted rows with Python’s standard csv.writer.
 1# Extract an HTML table to CSV in Python
 2
 3import csv
 4import aspose.html as ah
 5
 6# Load the HTML file and extract the first table
 7with ah.HTMLDocument("data/product-table.html") as document:
 8    table = document.query_selector("table")
 9
10    if table is None:
11        raise RuntimeError("The HTML document contains no table.")
12
13    extracted_rows = []
14    rows = table.get_elements_by_tag_name("tr")
15
16    # Extract header and data cells
17    for row in rows:
18        cells = row.query_selector_all("th, td")
19        values = [" ".join(cell.text_content.split()) for cell in cells]
20
21        if values:
22            extracted_rows.append(values)
23
24# Save the table as CSV
25with open("product-table.csv", "w", newline="", encoding="utf-8") as file:
26    csv.writer(file).writerows(extracted_rows)

The CSV file contains the table header followed by three data rows. Using csv.writer ensures that values containing commas, quotation marks, or line breaks are escaped according to CSV rules.

Extract Tables from a Webpage in Python

Pass an absolute URL to HTMLDocument when tables should be extracted directly from a webpage. The next example loads the Aspose.HTML Supported File Formats page, processes every <table> in the returned DOM, and saves the rows as a readable text report.

To extract HTML tables from a webpage:

  1. Pass the webpage URL to the HTMLDocument constructor.
  2. Collect all table elements with get_elements_by_tag_name("table").
  3. Stop with a clear error when the loaded page contains no tables.
  4. Process the tr rows and their th and td cells in every table.
  5. Normalize cell text and join values with a visible separator.
  6. Save the extracted rows as a .txt file.
 1# Extract HTML tables from a webpage to TXT in Python
 2
 3import aspose.html as ah
 4
 5page_url = "https://docs.aspose.com/html/python-net/supported-file-formats/"
 6
 7# Load the webpage and extract its tables
 8with ah.HTMLDocument(page_url) as document:
 9    tables = document.get_elements_by_tag_name("table")
10
11    if tables.length == 0:
12        raise RuntimeError("The webpage contains no tables.")
13
14    lines = []
15
16    # Extract rows and cells
17    for table_index, table in enumerate(tables, start=1):
18        lines.append(f"Table {table_index}")
19        rows = table.get_elements_by_tag_name("tr")
20
21        for row in rows:
22            cells = row.query_selector_all("th, td")
23            values = [" ".join(cell.text_content.split()) for cell in cells]
24
25            if values:
26                lines.append(" | ".join(values))
27
28        lines.append("")
29
30# Save the tables as a text report
31with open("webpage-tables.txt", "w", encoding="utf-8") as file:
32    file.write("\n".join(lines))

The report labels each table and writes one source row per line. Selectors and table structure vary between websites, so inspect the loaded DOM and narrow the extraction scope when a page contains unrelated layout or comparison tables.

Save Extracted Tables as Separate HTML Files

Use outer_html when the original table markup should be preserved instead of reducing each row to plain text values. The following example loads a webpage, finds every <table>, and saves each table as a separate .htm document.

To save extracted tables as HTML files:

  1. Create the output directory and load the webpage with HTMLDocument.
  2. Collect all <table> elements with get_elements_by_tag_name("table").
  3. Read outer_html from each table to preserve its complete element markup.
  4. Create a new HTMLDocument from the extracted table markup.
  5. Save each document to a separate HTML file.
  6. Print a message when the loaded webpage contains no tables.
 1# Save HTML tables from a website as separate files in Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Prepare the output directory
 7output_dir = "output"
 8os.makedirs(output_dir, exist_ok=True)
 9
10# Load the webpage and collect its tables
11with ah.HTMLDocument("https://docs.aspose.com/html/net/edit-html-document/") as document:
12    tables = document.get_elements_by_tag_name("table")
13
14    if tables.length > 0:
15        for i, table in enumerate(tables):
16            file_name = f"table{i}.htm"
17            file_path = os.path.join(output_dir, file_name)
18
19            # Save the table as a separate HTML document
20            with ah.HTMLDocument(table.outer_html, file_path) as table_document:
21                table_document.save(file_path)
22    else:
23        print("No tables found in the document.")

The example preserves each table element, its rows, cells, attributes, and inline styles. CSS rules, fonts, images, or other resources defined outside the table are not copied automatically, so an extracted table can look different from the original webpage.

Choose a Table Extraction Method

TaskRecommended approachUse when
Select the first tabledocument.query_selector("table")The document contains one main data table.
Collect every tabledocument.get_elements_by_tag_name("table")Every table in the document must be processed.
Select tables from main contentSelect main, then call main.get_elements_by_tag_name("table")Navigation or related page regions may contain unrelated tables.
Read table rowstable.get_elements_by_tag_name("tr")Header and body rows should both be included.
Read header and data cellsrow.query_selector_all("th, td")Column labels must be preserved with ordinary cell values.
Select a table by ID or classdocument.query_selector("table#report") or another focused selectorThe page contains several tables but only one is required.
Preserve table markupSave table.outer_html in a new HTMLDocumentTable elements and attributes must remain available in the output.

Common HTML Table Extraction Issues

IssueCauseRecommended action
No table is foundThe loaded document contains no <table> or the selector does not match it.Inspect the returned DOM and begin with a broad table selector.
The wrong table is extractedquery_selector("table") returns only the first matching table.Select a stable ID, class, or content container, or process every table.
CSV columns are shiftedThe source uses rowspan or colspan.Add application-specific logic that expands spanning cells into a rectangular grid.
Extra whitespace appears in cellsCell text contains indentation, line breaks, or nested inline elements.Normalize text_content before saving it.
An extracted HTML table loses its appearancePage-level or external CSS is outside the saved table markup.Include the required styles or resources in the new document.
A browser shows a table but extraction finds noneThe table is absent from the DOM available after loading.Inspect the loaded document and verify how the page produces the table.

FAQ

Can Aspose.HTML extract a table from a webpage URL in Python?

Yes. Pass the absolute URL to HTMLDocument, locate the required table elements, and read their rows and cells. Network access and the HTML returned by the server determine which tables are available.

How do I preserve table headers in the extracted data?

Select both th and td cells when processing rows. The first example writes header values as the first row of the CSV file.

Does the example preserve colspan and rowspan?

It preserves the text of each physical cell but does not expand colspan or rowspan into a normalized data grid. Add grid-processing logic when the output must reproduce the visual table layout.

Can I export an HTML table to JSON or a database?

Yes. The Aspose.HTML part of the workflow selects the table and extracts its cell values. Replace the CSV or TXT writing step with the JSON serializer or database client required by your application.

Can I save each extracted table as a separate HTML file?

Yes. Read the table’s outer_html, create a new HTMLDocument from that markup, and save it to an .html or .htm file. Include any required page-level styles separately when the output should retain the original appearance.

Related Articles