Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To extract text from an HTML file or website in Python, load the local path or URL with
HTMLDocument and read document.body.text_content. When only part of the page is required, select elements with
query_selector_all() and read the text_content of each match.
Aspose.HTML for Python via .NET parses HTML into a Document Object Model (DOM), allowing you to extract text from the entire document or selected elements. You can read the complete document body or select only headings, paragraphs, list items, and other relevant elements.
This article demonstrates how to extract body text from a local HTML file, exclude navigation and footer content with a CSS selector, and collect text directly from a website URL.
If the complete source should be written directly to a TXT file, use
HTML-to-text conversion with Converter.convert_html() and TextSaveOptions. The DOM workflow on this page is intended for selecting and processing text before the application stores or uses it.
The body property provides access to the document <body> element. Its text_content property returns the combined text of that element and its descendants without serializing HTML tags.
The following example uses
extract-text-article.html, extracts its body text, and saves the result as extracted-text.txt:
HTMLDocument in a context manager.document.body and read its text_content property.strip(). 1# Extract plain text from an HTML file in Python
2
3import aspose.html as ah
4
5# Load the HTML file and extract body text
6with ah.HTMLDocument("data/extract-text-article.html") as document:
7 text = document.body.text_content.strip()
8
9# Save the extracted text
10with open("extracted-text.txt", "w", encoding="utf-8") as file:
11 file.write(text)The saved TXT file contains the report heading and paragraphs without the <main>, <h1>, or <p> tags. text_content follows DOM text nodes rather than the visual layout of a browser, so indentation and whitespace from the source may remain in the resulting string.
Reading the complete body can include navigation, sidebars, footer notices, and other template content. Use a CSS selector when only a particular content area or set of elements should be extracted.
The next example uses
extract-selected-content.html. It selects headings, paragraphs, and list items inside article.content, while excluding the navigation and footer.
To extract selected text from HTML in Python:
HTMLDocument in a context manager.query_selector_all() with a selector limited to the required content area.text_content values.strip() and skip empty strings. 1# Extract text from selected HTML elements in Python
2
3import aspose.html as ah
4
5# Define a selector for article text elements
6selector = (
7 "article.content h1, article.content h2, "
8 "article.content p, article.content li"
9)
10
11# Load the HTML file and extract selected text
12with ah.HTMLDocument("data/extract-selected-content.html") as document:
13 elements = document.query_selector_all(selector)
14
15 for element in elements:
16 text = element.text_content.strip()
17
18 if text:
19 print(text)The output contains only the heading, paragraph, and list items located inside article.content. The page navigation and footer are excluded because the selector does not target elements outside that container.
Pass an absolute webpage URL to HTMLDocument when the HTML should be loaded directly from a website. For most pages, selecting the main content is more useful than reading the complete body because the body can also contain menus, banners, and footer text.
The following example extracts headings and paragraphs from the main content of the Aspose.HTML for Python via .NET product page:
HTMLDocument constructor.text_content from each matched element.split() and join(). 1# Extract text from a webpage in Python
2
3import aspose.html as ah
4
5page_url = "https://products.aspose.com/html/python-net/"
6
7# Load the webpage and extract its main text
8with ah.HTMLDocument(page_url) as document:
9 elements = document.query_selector_all("main h1, main h2, main p")
10
11 for element in elements:
12 text = " ".join(element.text_content.split())
13
14 if text:
15 print(text)CSS selectors depend on the structure of the target website. Inspect the loaded DOM and adjust main h1, main h2, main p when another page uses different containers or elements. Network access, authentication, and the HTML returned by the server also affect the result.
| Task | Recommended approach | Notes |
|---|---|---|
| Extract all body text | document.body.text_content | Use when the complete body contains relevant content. |
| Extract the main content area | Select main or article, then read text_content | Excludes page regions outside the selected container. |
| Extract headings or repeated items | Use query_selector_all() with a focused selector | Processes each matched element separately and preserves logical groups. |
| Extract text from a webpage | Load the absolute URL and select elements inside main or another content container | The selector must match the DOM returned by the website. |
| Extract attributes with text | Combine text_content with get_attribute() | Required for link destinations, image sources, and custom data attributes. |
| Issue | Cause | Recommended action |
|---|---|---|
| Navigation or footer text appears | Text was read from the complete body. | Select main, article, or another stable content container first. |
| The output contains extra spaces or line breaks | text_content preserves whitespace found in DOM text nodes. | Normalize whitespace according to the required output format. |
| Text from hidden elements is included | DOM text extraction does not determine whether content is visually displayed. | Exclude known hidden elements or inspect their attributes and styles. |
| Link destinations are missing | text_content returns anchor text rather than its href value. | Select a[href] elements and read get_attribute("href") separately. |
| Website extraction returns no text | The URL is unavailable, requires authentication, or the selector does not match the returned HTML. | Verify network access and inspect the loaded DOM before changing the extraction code. |
| Expected text is absent | The required content is not present in the loaded DOM. | Inspect the loaded document and confirm that the selector matches its actual structure. |
No. It returns the combined text stored in an element and its descendants. Use inner_html or outer_html when the HTML markup must be preserved.
No. It reads DOM text and can include content from elements hidden through HTML attributes or CSS. It is not equivalent to copying visually rendered text from a browser.
Yes. Pass the absolute URL to HTMLDocument and use the same body or selector-based workflow. The website example above selects headings and paragraphs from the page’s <main> element. The returned content depends on the DOM available after loading.
Yes. After extracting the required string, write it to a text file with Python’s standard file operations. To convert the complete HTML source directly to TXT instead, use Converter.convert_html() with TextSaveOptions.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.