Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To extract an HTML table in Java, load the file or webpage with
HTMLDocument, locate a table with
querySelector() or all tables with
querySelectorAll(), select their tr, th, and td elements, and write the extracted values to CSV, TXT, or another data store.
HTML tables commonly contain reports, schedules, comparison data, product specifications, and reference values. Aspose.HTML for Java parses the page into a DOM tree, allowing an application to select table elements and read their text and attributes without processing the HTML as an unstructured string.
This article demonstrates how to extract one table from a local HTML file into CSV and how to collect every table from a webpage URL into a text report.
The following example uses
extract-html-table-java.html. It selects the first <table>, reads both header and data cells, escapes their values for CSV, and saves the result as html-table.csv.
To extract an HTML table to CSV:
HTMLDocument using a try-with-resources statement.querySelector("table") to select the first table.querySelectorAll("tr").th and td cells within each row.Files.write(). 1import com.aspose.html.HTMLDocument;
2import com.aspose.html.collections.NodeList;
3import com.aspose.html.dom.Element;
4
5import java.io.IOException;
6import java.nio.charset.StandardCharsets;
7import java.nio.file.Files;
8import java.nio.file.Paths;
9
10// Specify input and output file paths
11String inputPath = "extract-html-table-java.html";
12String outputPath = "html-table.csv";
13
14// Load the HTML document
15try (HTMLDocument document = new HTMLDocument(inputPath)) {
16
17 // Select the first table in the document
18 Element table = document.querySelector("table");
19
20 if (table == null) {
21 throw new IllegalStateException("The HTML document contains no table.");
22 }
23
24 // Select all table rows
25 NodeList rows = table.querySelectorAll("tr");
26 StringBuilder csv = new StringBuilder();
27
28 // Extract text from table cells
29 rows.forEach(rowNode -> {
30 Element row = (Element) rowNode;
31 NodeList cells = row.querySelectorAll("th, td");
32
33 final boolean[] firstCell = {true};
34
35 cells.forEach(cellNode -> {
36 if (!firstCell[0]) {
37 csv.append(",");
38 }
39
40 String value = cellNode.getTextContent().trim();
41
42 // Escape double quotes and add the value to the CSV row
43 csv.append("\"")
44 .append(value.replace("\"", "\"\""))
45 .append("\"");
46
47 firstCell[0] = false;
48 });
49
50 csv.append(System.lineSeparator());
51 });
52
53 // Save the extracted table data to a CSV file
54 Files.write(
55 Paths.get(outputPath),
56 csv.toString().getBytes(StandardCharsets.UTF_8)
57 );
58}The first CSV row contains the table headers. The remaining rows contain the extracted workflow, input, result, and Java API values. Quoting every value and doubling embedded quotation marks prevents commas and quotation marks inside cell text from breaking the basic CSV structure.
Pass a webpage URL to the HTMLDocument constructor when the tables should be extracted directly from a website. The next example selects all tables from the Java Data Extraction page and saves their row and cell text to webpage-tables.txt.
To extract all HTML tables from a webpage:
HTMLDocument using a try-with-resources statement.querySelectorAll("table") to collect every table in the loaded DOM.Files.write(). 1import com.aspose.html.HTMLDocument;
2import com.aspose.html.collections.NodeList;
3import com.aspose.html.dom.Element;
4
5import java.io.IOException;
6import java.nio.charset.StandardCharsets;
7import java.nio.file.Files;
8import java.nio.file.Paths;
9
10// Specify the web page URL and output file path
11String pageUrl = "https://docs.aspose.com/html/java/data-extraction/";
12String outputPath = "webpage-tables.txt";
13
14// Load an HTML document from the URL
15try (HTMLDocument document = new HTMLDocument(pageUrl)) {
16
17 // Select all tables on the web page
18 NodeList tables = document.querySelectorAll("table");
19 StringBuilder report = new StringBuilder();
20
21 // Extract data from each table
22 tables.forEach(tableNode -> {
23 report.append("Table")
24 .append(System.lineSeparator());
25
26 Element table = (Element) tableNode;
27 NodeList rows = table.querySelectorAll("tr");
28
29 // Extract data from each table row
30 rows.forEach(rowNode -> {
31 Element row = (Element) rowNode;
32 NodeList cells = row.querySelectorAll("th, td");
33
34 final boolean[] firstCell = {true};
35
36 // Extract text from header and data cells
37 cells.forEach(cellNode -> {
38 if (!firstCell[0]) {
39 report.append(" | ");
40 }
41
42 report.append(cellNode.getTextContent().trim());
43 firstCell[0] = false;
44 });
45
46 report.append(System.lineSeparator());
47 });
48
49 report.append(System.lineSeparator());
50 });
51
52 // Save the extracted table data to a text file
53 Files.write(
54 Paths.get(outputPath),
55 report.toString().getBytes(StandardCharsets.UTF_8)
56 );
57}The example reports every table available in the loaded page DOM. For a page containing unrelated layout or navigation tables, narrow the selector to a stable content area, such as main table or article table.
| Task | Selector and method | Use it when |
|---|---|---|
| Select the first table | querySelector("table") | The document contains one main data table. |
| Select every table | querySelectorAll("table") | Every table in the document must be processed. |
| Select tables in main content | querySelectorAll("main table") | Navigation or other page regions may contain unrelated tables. |
| Select table rows | table.querySelectorAll("tr") | Header and body rows should both be included. |
| Select header and data cells | row.querySelectorAll("th, td") | Header labels must be preserved with ordinary cell values. |
| Select links inside cells | table.querySelectorAll("th a[href], td a[href]") | Both link text and destination URLs are required. |
| Issue | Cause | Fix |
|---|---|---|
| No table is found | The selector does not match the DOM loaded from the file or URL. | Inspect the loaded HTML and test a broad table selector before adding conditions. |
| The wrong table is extracted | querySelector("table") returns only the first match. | Use a stable ID, class, parent container, or querySelectorAll() when the page contains several tables. |
| CSV columns are shifted | The source table uses rowspan or colspan. | Add grid-normalization logic that expands spanning cells before writing tabular output. |
| A link URL is missing from a cell | getTextContent() returns the anchor text but not its href. | Select the nested anchor and read getAttribute("href") separately. |
| A table visible in a browser is absent | Client-side code may create it after the source HTML is loaded. | Inspect the DOM available to HTMLDocument and verify that the required table exists before extraction. |
| Rows from a nested table are included | A descendant tr selector also finds rows inside nested tables. | Detect nested tables or use extraction logic limited to the required table sections and direct row ownership. |
Yes. Load the URL with HTMLDocument, select the required tables from the DOM, and read their rows and cells. Network access and the actual HTML returned by the website determine which tables are available.
Read th and td text from each row, escape each value according to the required CSV rules, join the fields with commas, and write one CSV line per table row. The first example includes basic escaping for commas, quotation marks, and line breaks inside quoted fields.
It preserves the text of each physical cell but does not expand colspan or rowspan into a rectangular data grid. Add application-specific normalization when the output must reproduce the visual table layout.
Yes. After selecting a cell, query its descendant a[href] or img[src] elements and read their attributes. getTextContent() alone returns text rather than linked resource URLs.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.