Convert HTML to Plain Text in Python

To convert HTML to text in Python, pass an HTML file, HTML string, HTMLDocument, or Url to Converter.convert_html() with TextSaveOptions and a .txt output path.

Aspose.HTML for Python via .NET converts an HTML document to plain text without HTML tags or CSS presentation. Use this workflow when the complete document should become a TXT file. To collect text from selected elements, exclude navigation, or process separate content blocks, use the DOM-based text extraction workflow instead.

Convert HTML to Text in Python

To convert an HTML source to TXT:

  1. Specify the HTML file, string, document, or URL to convert.
  2. Create a TextSaveOptions object.
  3. Call Converter.convert_html() with the source, options, and TXT output path.

Convert an HTML File to TXT

The following example converts extract-text-article.html to a text file. It uses the file-path overload, so the source does not need to be loaded into an HTMLDocument first:

 1import os
 2import aspose.html.converters as conv
 3import aspose.html.saving as sav
 4
 5data_dir = "data"
 6output_dir = "output"
 7os.makedirs(output_dir, exist_ok=True)
 8
 9input_path = os.path.join(data_dir, "extract-text-article.html")
10output_path = os.path.join(output_dir, "article.txt")
11
12options = sav.TextSaveOptions()
13conv.Converter.convert_html(input_path, options, output_path)

The resulting article.txt contains the document’s textual content without elements such as <main>, <h1>, or <p>. Paragraphs and other block-level content are separated in the text output.

Convert an HTML String and Keep List Markers

Set enable_list_item_markers to True when list items in the TXT output should retain a marker. The string overload also requires a base URI, which identifies the location used to resolve any relative references in the HTML. This self-contained example uses the current working directory:

 1import os
 2import aspose.html.converters as conv
 3import aspose.html.saving as sav
 4
 5output_dir = "output"
 6os.makedirs(output_dir, exist_ok=True)
 7output_path = os.path.join(output_dir, "checklist.txt")
 8
 9html_code = """
10<h1>Release Checklist</h1>
11<ul>
12    <li>Review the document.</li>
13    <li>Verify the output.</li>
14</ul>
15"""
16
17options = sav.TextSaveOptions()
18options.enable_list_item_markers = True
19
20conv.Converter.convert_html(
21    html_code,
22    os.getcwd() + os.sep,
23    options,
24    output_path,
25)

The generated TXT file contains the heading and both list items. With list markers enabled, each list item is preceded by a marker instead of being written as unmarked text.

Convert a Webpage URL to Text

Use a Url object when the source HTML should be loaded directly from a website:

 1import os
 2import aspose.html as ah
 3import aspose.html.converters as conv
 4import aspose.html.saving as sav
 5
 6output_dir = "output"
 7os.makedirs(output_dir, exist_ok=True)
 8output_path = os.path.join(output_dir, "webpage.txt")
 9
10page_url = ah.Url("https://docs.aspose.com/html/files/aspose.html")
11options = sav.TextSaveOptions()
12
13conv.Converter.convert_html(page_url, options, output_path)

The application must have network access to the URL. This workflow converts the complete loaded document; it does not isolate the main article from navigation, banners, or footer content.

Configure Text Output with TextSaveOptions

TextSaveOptions represents the settings for HTML-to-text conversion. Its text-specific property controls whether list markers are included:

PropertyUse it to control
enable_list_item_markersWhether markers are written before list items. The default value is False.

TXT output does not preserve fonts, colors, borders, images, or page layout because a plain-text file has no equivalent representation for visual styling.

TextSaveOptions vs. text_content

Both approaches produce text, but they solve different tasks:

RequirementRecommended approach
Convert a complete HTML file, string, document, or URL directly to TXTConverter.convert_html() with TextSaveOptions
Keep list markers in converted textSet TextSaveOptions.enable_list_item_markers to True
Extract text only from main, article, or selected elementsSelect the elements and read text_content
Exclude navigation, footer, or other page regionsUse CSS selectors or XPath before collecting text
Read attributes such as href, src, or custom data valuesUse DOM methods such as get_attribute()

text_content returns the text nodes contained by the selected DOM element. It gives the application control over what is collected, but Python code must write the resulting string to a TXT file. TextSaveOptions performs complete document-to-text conversion and writes the output through the Converter API.

Common HTML to Text Conversion Issues

ProblemCause and solution
List items have no markersMarkers are disabled by default. Set enable_list_item_markers to True.
Navigation and footer text appear in the outputTextSaveOptions converts the complete document. Use DOM selectors and text_content when only a content region is required.
The TXT file does not look like the webpagePlain text does not support HTML layout, CSS formatting, images, or fonts. Use PDF or an image format when visual appearance must be preserved.
The output contains more line breaks than expectedHTML blocks are separated in the text output. Normalize the generated text afterward when another whitespace convention is required.
URL conversion failsConfirm that the address is accessible from the application environment and does not require unsupported authentication or interaction.

FAQ

Does HTML-to-text conversion remove HTML tags?

Yes. The generated TXT file contains textual content rather than serialized HTML markup. Use HTMLDocument.save() when the markup must remain HTML.

What is the difference between HTML to TXT and HTML to Markdown?

TXT contains plain text without markup for headings, links, emphasis, or tables. Markdown retains supported document structure through Markdown syntax. Use HTML to Markdown conversion when that structure is required.

Can I convert a webpage URL to a TXT file?

Yes. Pass an absolute Url and TextSaveOptions to Converter.convert_html(). The result contains text from the complete loaded document. Use DOM extraction when only selected webpage content is required.

Related Articles