Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To convert HTML to text in Python, pass an HTML file, HTML string,
HTMLDocument, or
Url to
Converter.convert_html() with
TextSaveOptions and a .txt output path.
Aspose.HTML for Python via .NET converts an HTML document to plain text without HTML tags or CSS presentation. Use this workflow when the complete document should become a TXT file. To collect text from selected elements, exclude navigation, or process separate content blocks, use the DOM-based text extraction workflow instead.
To convert an HTML source to TXT:
TextSaveOptions object.Converter.convert_html() with the source, options, and TXT output path.The following example converts
extract-text-article.html to a text file. It uses the file-path overload, so the source does not need to be loaded into an HTMLDocument first:
1import os
2import aspose.html.converters as conv
3import aspose.html.saving as sav
4
5data_dir = "data"
6output_dir = "output"
7os.makedirs(output_dir, exist_ok=True)
8
9input_path = os.path.join(data_dir, "extract-text-article.html")
10output_path = os.path.join(output_dir, "article.txt")
11
12options = sav.TextSaveOptions()
13conv.Converter.convert_html(input_path, options, output_path)The resulting article.txt contains the document’s textual content without elements such as <main>, <h1>, or <p>. Paragraphs and other block-level content are separated in the text output.
Set enable_list_item_markers to True when list items in the TXT output should retain a marker. The string overload also requires a base URI, which identifies the location used to resolve any relative references in the HTML. This self-contained example uses the current working directory:
1import os
2import aspose.html.converters as conv
3import aspose.html.saving as sav
4
5output_dir = "output"
6os.makedirs(output_dir, exist_ok=True)
7output_path = os.path.join(output_dir, "checklist.txt")
8
9html_code = """
10<h1>Release Checklist</h1>
11<ul>
12 <li>Review the document.</li>
13 <li>Verify the output.</li>
14</ul>
15"""
16
17options = sav.TextSaveOptions()
18options.enable_list_item_markers = True
19
20conv.Converter.convert_html(
21 html_code,
22 os.getcwd() + os.sep,
23 options,
24 output_path,
25)The generated TXT file contains the heading and both list items. With list markers enabled, each list item is preceded by a marker instead of being written as unmarked text.
Use a Url object when the source HTML should be loaded directly from a website:
1import os
2import aspose.html as ah
3import aspose.html.converters as conv
4import aspose.html.saving as sav
5
6output_dir = "output"
7os.makedirs(output_dir, exist_ok=True)
8output_path = os.path.join(output_dir, "webpage.txt")
9
10page_url = ah.Url("https://docs.aspose.com/html/files/aspose.html")
11options = sav.TextSaveOptions()
12
13conv.Converter.convert_html(page_url, options, output_path)The application must have network access to the URL. This workflow converts the complete loaded document; it does not isolate the main article from navigation, banners, or footer content.
TextSaveOptions represents the settings for HTML-to-text conversion. Its text-specific property controls whether list markers are included:
| Property | Use it to control |
|---|---|
| enable_list_item_markers | Whether markers are written before list items. The default value is False. |
TXT output does not preserve fonts, colors, borders, images, or page layout because a plain-text file has no equivalent representation for visual styling.
Both approaches produce text, but they solve different tasks:
| Requirement | Recommended approach |
|---|---|
| Convert a complete HTML file, string, document, or URL directly to TXT | Converter.convert_html() with TextSaveOptions |
| Keep list markers in converted text | Set TextSaveOptions.enable_list_item_markers to True |
Extract text only from main, article, or selected elements | Select the elements and read text_content |
| Exclude navigation, footer, or other page regions | Use CSS selectors or XPath before collecting text |
Read attributes such as href, src, or custom data values | Use DOM methods such as get_attribute() |
text_content returns the text nodes contained by the selected DOM element. It gives the application control over what is collected, but Python code must write the resulting string to a TXT file. TextSaveOptions performs complete document-to-text conversion and writes the output through the Converter API.
| Problem | Cause and solution |
|---|---|
| List items have no markers | Markers are disabled by default. Set enable_list_item_markers to True. |
| Navigation and footer text appear in the output | TextSaveOptions converts the complete document. Use DOM selectors and text_content when only a content region is required. |
| The TXT file does not look like the webpage | Plain text does not support HTML layout, CSS formatting, images, or fonts. Use PDF or an image format when visual appearance must be preserved. |
| The output contains more line breaks than expected | HTML blocks are separated in the text output. Normalize the generated text afterward when another whitespace convention is required. |
| URL conversion fails | Confirm that the address is accessible from the application environment and does not require unsupported authentication or interaction. |
Yes. The generated TXT file contains textual content rather than serialized HTML markup. Use HTMLDocument.save() when the markup must remain HTML.
TXT contains plain text without markup for headings, links, emphasis, or tables. Markdown retains supported document structure through Markdown syntax. Use HTML to Markdown conversion when that structure is required.
Yes. Pass an absolute Url and TextSaveOptions to Converter.convert_html(). The result contains text from the complete loaded document. Use DOM extraction when only selected webpage content is required.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.