Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To extract text from HTML in Java, load the source with
HTMLDocument, read all body text with
getTextContent(), or use
querySelectorAll() before reading text from selected elements. Display or pass the resulting strings to another application component.
Aspose.HTML for Java parses HTML into a DOM tree, so you can extract plain text without using regular expressions to remove tags. Read the complete document body or select only meaningful content such as headings, paragraphs, and list items.
This article demonstrates how to extract all body text from a local HTML file and how to exclude navigation and footer content by selecting specific elements.
The
getBody() method returns the document <body> element. Calling getTextContent() on that element returns the text of the body and its descendants without serializing the HTML tags.
The following example uses extract-text-article.html, extracts its body text, and prints the result to the console:
HTMLDocument using a try-with-resources statement.<body> element and read its text with getTextContent(). 1import com.aspose.html.HTMLDocument;
2
3// Load an HTML document
4try (HTMLDocument document = new HTMLDocument("extract-text-article.html")) {
5
6 // Extract text from the HTML body
7 String text = document.getBody().getTextContent().trim();
8
9 System.out.println(text);
10}The output contains the body text without <main>, <h1>, or <p> tags. getTextContent() follows the DOM text nodes rather than the browser’s visual layout, so source indentation and whitespace may still require normalization for a particular output format.
Reading the complete body can also include navigation, sidebars, footer notices, and other template content. Use a CSS selector when only a specific content area or set of elements should be extracted.
The next example uses extract-selected-content-java.html. It selects headings, paragraphs, and list items inside the article while excluding the page navigation and footer.
To extract selected text from HTML in Java:
HTMLDocument using a try-with-resources statement.querySelectorAll() with a selector limited to the required content container.getTextContent() from every matched node. 1import com.aspose.html.HTMLDocument;
2import com.aspose.html.collections.NodeList;
3
4// Load an HTML document from a file
5try (HTMLDocument document = new HTMLDocument("extract-selected-content-java.html")) {
6
7 // Select headings, paragraphs, and list items from the article
8 NodeList elements = document.querySelectorAll(
9 "article.content h1, article.content h2, "
10 + "article.content p, article.content li"
11 );
12
13 // Extract and display text from the selected elements
14 elements.forEach(node -> {
15 String text = node.getTextContent().trim();
16
17 if (!text.isEmpty()) {
18 System.out.println(text);
19 }
20 });
21}The console output contains the article heading, paragraphs, and list items. It does not contain the Home, Products, Support, or copyright text from outside article.content.
| Text extraction task | Recommended approach | Notes |
|---|---|---|
| Extract all body text | document.getBody().getTextContent() | Suitable when the complete body contains useful content. |
| Extract the main article | querySelector("main") or querySelector("article") | Read getTextContent() from the selected container. |
| Extract headings or repeated content | querySelectorAll("h1, h2, h3") or another focused selector | Process each match separately when structure must be preserved. |
| Extract attributes together with text | Select the element, then combine getTextContent() with getAttribute() | Required for values such as link URLs, image sources, or custom data attributes. |
| Issue | Cause | Fix |
|---|---|---|
| Navigation or footer text appears in the result | Text was read from the complete body. | Select main, article, or another stable content container before calling getTextContent(). |
| The output contains unexpected spaces or line breaks | getTextContent() returns DOM text and preserves whitespace present in text nodes. | Apply whitespace normalization that matches the required TXT or data format. |
| Text from hidden elements is included | DOM text extraction does not determine whether an element is visually displayed. | Exclude known hidden elements with selectors or inspect their attributes and styles before collecting text. |
| Link destinations are missing | getTextContent() returns anchor text, not the href attribute. | Select a[href] elements and read getAttribute("href") separately. |
| Expected text is absent | The content may not exist in the DOM available after loading. | Inspect the loaded document and verify whether scripts or another external process add the content later. |
No. It returns the text stored in an element and its descendant nodes. Use getInnerHTML() or getOuterHTML() when markup must be preserved.
No. It reads text from the DOM and can include text from elements hidden by HTML attributes or CSS. It is not the same as copying visually rendered text from a browser.
Yes. Create an HTMLDocument from the page URL and apply the same body or selector-based extraction. Inspect the loaded DOM first because the page structure and availability of script-generated content vary between websites.
Yes. After extracting the required strings with Aspose.HTML, write them to a TXT file with standard Java I/O APIs such as Files.write() or a BufferedWriter.
href values or resolve relative links to absolute URLs.TreeWalker, XPath, and CSS selectors.Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.