Extract Text from HTML in Java

To extract text from HTML in Java, load the source with HTMLDocument, read all body text with getTextContent(), or use querySelectorAll() before reading text from selected elements. Display or pass the resulting strings to another application component.

Aspose.HTML for Java parses HTML into a DOM tree, so you can extract plain text without using regular expressions to remove tags. Read the complete document body or select only meaningful content such as headings, paragraphs, and list items.

This article demonstrates how to extract all body text from a local HTML file and how to exclude navigation and footer content by selecting specific elements.

Extract Body Text from HTML

The getBody() method returns the document <body> element. Calling getTextContent() on that element returns the text of the body and its descendants without serializing the HTML tags.

The following example uses extract-text-article.html, extracts its body text, and prints the result to the console:

  1. Open the local HTML file in an HTMLDocument using a try-with-resources statement.
  2. Access the <body> element and read its text with getTextContent().
  3. Trim whitespace at the beginning and end of the extracted string.
  4. Print the extracted text to the console.
 1import com.aspose.html.HTMLDocument;
 2
 3// Load an HTML document
 4try (HTMLDocument document = new HTMLDocument("extract-text-article.html")) {
 5
 6    // Extract text from the HTML body
 7    String text = document.getBody().getTextContent().trim();
 8
 9    System.out.println(text);
10}

The output contains the body text without <main>, <h1>, or <p> tags. getTextContent() follows the DOM text nodes rather than the browser’s visual layout, so source indentation and whitespace may still require normalization for a particular output format.

Extract Text from Selected HTML Elements

Reading the complete body can also include navigation, sidebars, footer notices, and other template content. Use a CSS selector when only a specific content area or set of elements should be extracted.

The next example uses extract-selected-content-java.html. It selects headings, paragraphs, and list items inside the article while excluding the page navigation and footer.

To extract selected text from HTML in Java:

  1. Open the HTML file in an HTMLDocument using a try-with-resources statement.
  2. Call querySelectorAll() with a selector limited to the required content container.
  3. Read and trim getTextContent() from every matched node.
  4. Skip elements whose normalized text is empty.
  5. Print every extracted value on a separate console line.
 1import com.aspose.html.HTMLDocument;
 2import com.aspose.html.collections.NodeList;
 3
 4// Load an HTML document from a file
 5try (HTMLDocument document = new HTMLDocument("extract-selected-content-java.html")) {
 6
 7    // Select headings, paragraphs, and list items from the article
 8    NodeList elements = document.querySelectorAll(
 9            "article.content h1, article.content h2, "
10                    + "article.content p, article.content li"
11    );
12
13    // Extract and display text from the selected elements
14    elements.forEach(node -> {
15        String text = node.getTextContent().trim();
16
17        if (!text.isEmpty()) {
18            System.out.println(text);
19        }
20    });
21}

The console output contains the article heading, paragraphs, and list items. It does not contain the Home, Products, Support, or copyright text from outside article.content.

Choose a Text Extraction Method

Text extraction taskRecommended approachNotes
Extract all body textdocument.getBody().getTextContent()Suitable when the complete body contains useful content.
Extract the main articlequerySelector("main") or querySelector("article")Read getTextContent() from the selected container.
Extract headings or repeated contentquerySelectorAll("h1, h2, h3") or another focused selectorProcess each match separately when structure must be preserved.
Extract attributes together with textSelect the element, then combine getTextContent() with getAttribute()Required for values such as link URLs, image sources, or custom data attributes.

Common HTML Text Extraction Issues

IssueCauseFix
Navigation or footer text appears in the resultText was read from the complete body.Select main, article, or another stable content container before calling getTextContent().
The output contains unexpected spaces or line breaksgetTextContent() returns DOM text and preserves whitespace present in text nodes.Apply whitespace normalization that matches the required TXT or data format.
Text from hidden elements is includedDOM text extraction does not determine whether an element is visually displayed.Exclude known hidden elements with selectors or inspect their attributes and styles before collecting text.
Link destinations are missinggetTextContent() returns anchor text, not the href attribute.Select a[href] elements and read getAttribute("href") separately.
Expected text is absentThe content may not exist in the DOM available after loading.Inspect the loaded document and verify whether scripts or another external process add the content later.

FAQ

Does getTextContent() include HTML tags?

No. It returns the text stored in an element and its descendant nodes. Use getInnerHTML() or getOuterHTML() when markup must be preserved.

Does getTextContent() return only visible text?

No. It reads text from the DOM and can include text from elements hidden by HTML attributes or CSS. It is not the same as copying visually rendered text from a browser.

Can I extract text directly from a webpage URL?

Yes. Create an HTMLDocument from the page URL and apply the same body or selector-based extraction. Inspect the loaded DOM first because the page structure and availability of script-generated content vary between websites.

Can extracted HTML text be saved as a TXT file?

Yes. After extracting the required strings with Aspose.HTML, write them to a TXT file with standard Java I/O APIs such as Files.write() or a BufferedWriter.

Related Data Extraction Articles