Extract HTML Tables in Java

To extract an HTML table in Java, load the file or webpage with HTMLDocument, locate a table with querySelector() or all tables with querySelectorAll(), select their tr, th, and td elements, and write the extracted values to CSV, TXT, or another data store.

HTML tables commonly contain reports, schedules, comparison data, product specifications, and reference values. Aspose.HTML for Java parses the page into a DOM tree, allowing an application to select table elements and read their text and attributes without processing the HTML as an unstructured string.

This article demonstrates how to extract one table from a local HTML file into CSV and how to collect every table from a webpage URL into a text report.

Extract an HTML Table to CSV in Java

The following example uses extract-html-table-java.html. It selects the first <table>, reads both header and data cells, escapes their values for CSV, and saves the result as html-table.csv.

To extract an HTML table to CSV:

  1. Open the local HTML file in an HTMLDocument using a try-with-resources statement.
  2. Call querySelector("table") to select the first table.
  3. Select every table row with querySelectorAll("tr").
  4. Select both th and td cells within each row.
  5. Read and trim the text of every cell.
  6. Escape quotation marks, quote each value, and join the values with commas.
  7. Write the accumulated text to a UTF-8 CSV file with Files.write().
 1import com.aspose.html.HTMLDocument;
 2import com.aspose.html.collections.NodeList;
 3import com.aspose.html.dom.Element;
 4
 5import java.io.IOException;
 6import java.nio.charset.StandardCharsets;
 7import java.nio.file.Files;
 8import java.nio.file.Paths;
 9
10// Specify input and output file paths
11String inputPath = "extract-html-table-java.html";
12String outputPath = "html-table.csv";
13
14// Load the HTML document
15try (HTMLDocument document = new HTMLDocument(inputPath)) {
16
17    // Select the first table in the document
18    Element table = document.querySelector("table");
19
20    if (table == null) {
21        throw new IllegalStateException("The HTML document contains no table.");
22    }
23
24    // Select all table rows
25    NodeList rows = table.querySelectorAll("tr");
26    StringBuilder csv = new StringBuilder();
27
28    // Extract text from table cells
29    rows.forEach(rowNode -> {
30        Element row = (Element) rowNode;
31        NodeList cells = row.querySelectorAll("th, td");
32
33        final boolean[] firstCell = {true};
34
35        cells.forEach(cellNode -> {
36            if (!firstCell[0]) {
37                csv.append(",");
38            }
39
40            String value = cellNode.getTextContent().trim();
41
42            // Escape double quotes and add the value to the CSV row
43            csv.append("\"")
44                    .append(value.replace("\"", "\"\""))
45                    .append("\"");
46
47            firstCell[0] = false;
48        });
49
50        csv.append(System.lineSeparator());
51    });
52
53    // Save the extracted table data to a CSV file
54    Files.write(
55            Paths.get(outputPath),
56            csv.toString().getBytes(StandardCharsets.UTF_8)
57    );
58}

The first CSV row contains the table headers. The remaining rows contain the extracted workflow, input, result, and Java API values. Quoting every value and doubling embedded quotation marks prevents commas and quotation marks inside cell text from breaking the basic CSV structure.

Extract All Tables from a Webpage URL

Pass a webpage URL to the HTMLDocument constructor when the tables should be extracted directly from a website. The next example selects all tables from the Java Data Extraction page and saves their row and cell text to webpage-tables.txt.

To extract all HTML tables from a webpage:

  1. Open the webpage URL in an HTMLDocument using a try-with-resources statement.
  2. Call querySelectorAll("table") to collect every table in the loaded DOM.
  3. Process the rows and cells within each table.
  4. Join the cell values with a visible separator for a readable text report.
  5. Save the accumulated report to a UTF-8 text file with Files.write().
 1import com.aspose.html.HTMLDocument;
 2import com.aspose.html.collections.NodeList;
 3import com.aspose.html.dom.Element;
 4
 5import java.io.IOException;
 6import java.nio.charset.StandardCharsets;
 7import java.nio.file.Files;
 8import java.nio.file.Paths;
 9
10// Specify the web page URL and output file path
11String pageUrl = "https://docs.aspose.com/html/java/data-extraction/";
12String outputPath = "webpage-tables.txt";
13
14// Load an HTML document from the URL
15try (HTMLDocument document = new HTMLDocument(pageUrl)) {
16
17    // Select all tables on the web page
18    NodeList tables = document.querySelectorAll("table");
19    StringBuilder report = new StringBuilder();
20
21    // Extract data from each table
22    tables.forEach(tableNode -> {
23        report.append("Table")
24                .append(System.lineSeparator());
25
26        Element table = (Element) tableNode;
27        NodeList rows = table.querySelectorAll("tr");
28
29        // Extract data from each table row
30        rows.forEach(rowNode -> {
31            Element row = (Element) rowNode;
32            NodeList cells = row.querySelectorAll("th, td");
33
34            final boolean[] firstCell = {true};
35
36            // Extract text from header and data cells
37            cells.forEach(cellNode -> {
38                if (!firstCell[0]) {
39                    report.append(" | ");
40                }
41
42                report.append(cellNode.getTextContent().trim());
43                firstCell[0] = false;
44            });
45
46            report.append(System.lineSeparator());
47        });
48
49        report.append(System.lineSeparator());
50    });
51
52    // Save the extracted table data to a text file
53    Files.write(
54            Paths.get(outputPath),
55            report.toString().getBytes(StandardCharsets.UTF_8)
56    );
57}

The example reports every table available in the loaded page DOM. For a page containing unrelated layout or navigation tables, narrow the selector to a stable content area, such as main table or article table.

Choose Selectors for Table Extraction

TaskSelector and methodUse it when
Select the first tablequerySelector("table")The document contains one main data table.
Select every tablequerySelectorAll("table")Every table in the document must be processed.
Select tables in main contentquerySelectorAll("main table")Navigation or other page regions may contain unrelated tables.
Select table rowstable.querySelectorAll("tr")Header and body rows should both be included.
Select header and data cellsrow.querySelectorAll("th, td")Header labels must be preserved with ordinary cell values.
Select links inside cellstable.querySelectorAll("th a[href], td a[href]")Both link text and destination URLs are required.

Common HTML Table Extraction Issues

IssueCauseFix
No table is foundThe selector does not match the DOM loaded from the file or URL.Inspect the loaded HTML and test a broad table selector before adding conditions.
The wrong table is extractedquerySelector("table") returns only the first match.Use a stable ID, class, parent container, or querySelectorAll() when the page contains several tables.
CSV columns are shiftedThe source table uses rowspan or colspan.Add grid-normalization logic that expands spanning cells before writing tabular output.
A link URL is missing from a cellgetTextContent() returns the anchor text but not its href.Select the nested anchor and read getAttribute("href") separately.
A table visible in a browser is absentClient-side code may create it after the source HTML is loaded.Inspect the DOM available to HTMLDocument and verify that the required table exists before extraction.
Rows from a nested table are includedA descendant tr selector also finds rows inside nested tables.Detect nested tables or use extraction logic limited to the required table sections and direct row ownership.

FAQ

Can Aspose.HTML for Java extract tables from a webpage URL?

Yes. Load the URL with HTMLDocument, select the required tables from the DOM, and read their rows and cells. Network access and the actual HTML returned by the website determine which tables are available.

How do I export an HTML table to CSV in Java?

Read th and td text from each row, escape each value according to the required CSV rules, join the fields with commas, and write one CSV line per table row. The first example includes basic escaping for commas, quotation marks, and line breaks inside quoted fields.

Does the example preserve colspan and rowspan?

It preserves the text of each physical cell but does not expand colspan or rowspan into a rectangular data grid. Add application-specific normalization when the output must reproduce the visual table layout.

Can I extract links or images from table cells?

Yes. After selecting a cell, query its descendant a[href] or img[src] elements and read their attributes. getTextContent() alone returns text rather than linked resource URLs.

Related Data Extraction Articles