Extract Text from HTML in Python

To extract text from an HTML file or website in Python, load the local path or URL with HTMLDocument and read document.body.text_content. When only part of the page is required, select elements with query_selector_all() and read the text_content of each match.

Aspose.HTML for Python via .NET parses HTML into a Document Object Model (DOM), allowing you to extract text from the entire document or selected elements. You can read the complete document body or select only headings, paragraphs, list items, and other relevant elements.

This article demonstrates how to extract body text from a local HTML file, exclude navigation and footer content with a CSS selector, and collect text directly from a website URL.

If the complete source should be written directly to a TXT file, use HTML-to-text conversion with Converter.convert_html() and TextSaveOptions. The DOM workflow on this page is intended for selecting and processing text before the application stores or uses it.

Extract Plain Text from an HTML File

The body property provides access to the document <body> element. Its text_content property returns the combined text of that element and its descendants without serializing HTML tags.

The following example uses extract-text-article.html, extracts its body text, and saves the result as extracted-text.txt:

  1. Load the local HTML file with HTMLDocument in a context manager.
  2. Access document.body and read its text_content property.
  3. Remove leading and trailing whitespace with strip().
  4. Write the resulting string to a UTF-8 text file.
 1# Extract plain text from an HTML file in Python
 2
 3import aspose.html as ah
 4
 5# Load the HTML file and extract body text
 6with ah.HTMLDocument("data/extract-text-article.html") as document:
 7    text = document.body.text_content.strip()
 8
 9# Save the extracted text
10with open("extracted-text.txt", "w", encoding="utf-8") as file:
11    file.write(text)

The saved TXT file contains the report heading and paragraphs without the <main>, <h1>, or <p> tags. text_content follows DOM text nodes rather than the visual layout of a browser, so indentation and whitespace from the source may remain in the resulting string.

Extract Text from Selected HTML Elements

Reading the complete body can include navigation, sidebars, footer notices, and other template content. Use a CSS selector when only a particular content area or set of elements should be extracted.

The next example uses extract-selected-content.html. It selects headings, paragraphs, and list items inside article.content, while excluding the navigation and footer.

To extract selected text from HTML in Python:

  1. Load the source file with HTMLDocument in a context manager.
  2. Call query_selector_all() with a selector limited to the required content area.
  3. Iterate through the returned nodes and read their text_content values.
  4. Apply strip() and skip empty strings.
  5. Print each extracted value on a separate line.
 1# Extract text from selected HTML elements in Python
 2
 3import aspose.html as ah
 4
 5# Define a selector for article text elements
 6selector = (
 7    "article.content h1, article.content h2, "
 8    "article.content p, article.content li"
 9)
10
11# Load the HTML file and extract selected text
12with ah.HTMLDocument("data/extract-selected-content.html") as document:
13    elements = document.query_selector_all(selector)
14
15    for element in elements:
16        text = element.text_content.strip()
17
18        if text:
19            print(text)

The output contains only the heading, paragraph, and list items located inside article.content. The page navigation and footer are excluded because the selector does not target elements outside that container.

Extract Text from a Webpage in Python

Pass an absolute webpage URL to HTMLDocument when the HTML should be loaded directly from a website. For most pages, selecting the main content is more useful than reading the complete body because the body can also contain menus, banners, and footer text.

The following example extracts headings and paragraphs from the main content of the Aspose.HTML for Python via .NET product page:

  1. Pass the webpage URL to the HTMLDocument constructor.
  2. Select the content elements required from the relevant page container.
  3. Read text_content from each matched element.
  4. Normalize consecutive whitespace with split() and join().
  5. Skip empty values and print the extracted webpage text.
 1# Extract text from a webpage in Python
 2
 3import aspose.html as ah
 4
 5page_url = "https://products.aspose.com/html/python-net/"
 6
 7# Load the webpage and extract its main text
 8with ah.HTMLDocument(page_url) as document:
 9    elements = document.query_selector_all("main h1, main h2, main p")
10
11    for element in elements:
12        text = " ".join(element.text_content.split())
13
14        if text:
15            print(text)

CSS selectors depend on the structure of the target website. Inspect the loaded DOM and adjust main h1, main h2, main p when another page uses different containers or elements. Network access, authentication, and the HTML returned by the server also affect the result.

Choose a Text Extraction Method

TaskRecommended approachNotes
Extract all body textdocument.body.text_contentUse when the complete body contains relevant content.
Extract the main content areaSelect main or article, then read text_contentExcludes page regions outside the selected container.
Extract headings or repeated itemsUse query_selector_all() with a focused selectorProcesses each matched element separately and preserves logical groups.
Extract text from a webpageLoad the absolute URL and select elements inside main or another content containerThe selector must match the DOM returned by the website.
Extract attributes with textCombine text_content with get_attribute()Required for link destinations, image sources, and custom data attributes.

Common HTML Text Extraction Issues

IssueCauseRecommended action
Navigation or footer text appearsText was read from the complete body.Select main, article, or another stable content container first.
The output contains extra spaces or line breakstext_content preserves whitespace found in DOM text nodes.Normalize whitespace according to the required output format.
Text from hidden elements is includedDOM text extraction does not determine whether content is visually displayed.Exclude known hidden elements or inspect their attributes and styles.
Link destinations are missingtext_content returns anchor text rather than its href value.Select a[href] elements and read get_attribute("href") separately.
Website extraction returns no textThe URL is unavailable, requires authentication, or the selector does not match the returned HTML.Verify network access and inspect the loaded DOM before changing the extraction code.
Expected text is absentThe required content is not present in the loaded DOM.Inspect the loaded document and confirm that the selector matches its actual structure.

FAQ

Does text_content include HTML tags?

No. It returns the combined text stored in an element and its descendants. Use inner_html or outer_html when the HTML markup must be preserved.

Does text_content return only visible text?

No. It reads DOM text and can include content from elements hidden through HTML attributes or CSS. It is not equivalent to copying visually rendered text from a browser.

Can I extract text from a webpage URL?

Yes. Pass the absolute URL to HTMLDocument and use the same body or selector-based workflow. The website example above selects headings and paragraphs from the page’s <main> element. The returned content depends on the DOM available after loading.

Can I save the extracted text as a TXT file?

Yes. After extracting the required string, write it to a text file with Python’s standard file operations. To convert the complete HTML source directly to TXT instead, use Converter.convert_html() with TextSaveOptions.

Related Articles