Create and Load HTML Documents in Python

The HTMLDocument class represents an HTML document as a Document Object Model (DOM) tree. The document-tree model and HTML document structure are defined by the WHATWG DOM and HTML Living Standard. You can create an empty document or load HTML from a local file, URL, string, or binary stream and then inspect, edit, save, or convert it.

Use HTMLDocument() to create an empty HTML document. To load existing content, pass a file path or URL to a supported constructor. When loading HTML from a string or stream, also provide a base URI so that relative CSS, images, scripts, and links can be resolved correctly.

Create an HTML Document

Aspose.HTML for Python via .NET creates the standard <html>, <head>, and <body> structure for a new HTMLDocument. You can save that structure immediately or add DOM nodes before saving it.

Create an Empty HTML Document

The following example creates an empty document and saves it as output/document-empty.html:

  1. Create the output directory and destination path.
  2. Initialize HTMLDocument without constructor arguments.
  3. Call save() to write the document to an HTML file.
 1# Create an empty HTML document using Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Setup an output directory and prepare a path to save the document
 7output_dir = "output"
 8if not os.path.exists(output_dir):
 9    os.makedirs(output_dir)
10save_path = os.path.join(output_dir, "document-empty.html")
11
12# Initialize an empty HTML document
13document = ah.HTMLDocument()
14
15# Work with the document here...
16
17# Save the document to a file
18document.save(save_path)

The saved file contains an empty HTML document structure. Add elements to the DOM before calling save() when the file should contain page content.

Create HTML Content with the DOM

This example creates a document, adds a Hello, World! text node to its <body>, and saves output/create-new-document.html:

  1. Open a new HTMLDocument with a Python context manager.
  2. Create a text node with create_text_node().
  3. Append the node to document.body.
  4. Save the populated document.
 1# Create an HTML document using Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Prepare the output path to save a document
 7output_dir = "output"
 8if not os.path.exists(output_dir):
 9    os.makedirs(output_dir)
10document_path = os.path.join(output_dir, "create-new-document.html")
11
12# Initialize an empty HTML document
13with ah.HTMLDocument() as document:
14    # Create a text node and add it to the document
15    text = document.create_text_node("Hello, World!")
16    document.body.append_child(text)
17
18    # Save the document to a file
19    document.save(document_path)

Continue with Edit an HTML Document for examples that create elements, set attributes, and modify CSS.

Load HTML from Different Sources

Choose the constructor that matches where the HTML content comes from. A file path or URL identifies an external source, while a string or stream supplies the markup directly.

Load HTML from a File

The following example loads data/document.html and saves it to output/document-edited.html:

  1. Build the input and output paths.
  2. Load the local file with HTMLDocument(document_path).
  3. Make any required DOM changes where indicated in the example.
  4. Save the document to the output path.
 1# Load HTML from a file using Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Setup directories and define paths
 7output_dir = "output"
 8input_dir = "data"
 9if not os.path.exists(output_dir):
10    os.makedirs(output_dir)
11
12document_path = os.path.join(input_dir, "document.html")
13save_path = os.path.join(output_dir, "document-edited.html")
14
15# Initialize a document from a file
16document = ah.HTMLDocument(document_path)
17
18# Work with the document
19
20# Save the document to a file
21document.save(save_path)

The current example does not modify the DOM; it demonstrates the load-and-save workflow. Add editing operations before document.save() when the output must differ from the source.

Load HTML from a URL

HTMLDocument can load a web page directly from an absolute URL. The constructor performs a synchronous load and processes linked resources required by the document.

The example performs these steps:

  1. Pass the web-page URL to HTMLDocument.
  2. Access the root element through document_element.
  3. Print the serialized HTML through outer_html.
1# Load HTML from a URL using Python
2
3import aspose.html as ah
4
5# Load a document from the specified web page
6document = ah.HTMLDocument("https://docs.aspose.com/html/files/aspose.html")
7
8# Write the document content to the output stream
9print(document.document_element.outer_html)

This example writes the loaded markup to the console; it does not save an output file. If the page cannot be loaded, the constructor raises a loading error, so production code should handle network and resource-loading failures.

Load HTML from a String

Use an HTMLDocument constructor that accepts HTML markup and a base URI when the source already exists in memory. The base URI determines how relative resource URLs are resolved.

The following example:

  1. Defines <p>Hello, World!</p> as a Python string.
  2. Creates a document from the string and uses . as the base URI.
  3. Saves the result as output/create-html-from-string.html.
 1# Create HTML from a string using Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Prepare HTML code
 7html_code = "<p>Hello, World!</p>"
 8
 9# Setup output directory
10output_dir = "output"
11if not os.path.exists(output_dir):
12    os.makedirs(output_dir)
13
14# Initialize a document from the string variable
15document = ah.HTMLDocument(html_code, ".")
16
17# Save the document to disk
18document.save(os.path.join(output_dir, "create-html-from-string.html"))

Load HTML from a Stream

Use io.BytesIO when HTML is already available as bytes and should be loaded without first writing a source file to disk.

The stream example:

  1. Encodes the HTML markup as a BytesIO stream.
  2. Supplies the stream and base URI to HTMLDocument.
  3. Saves the loaded document as output/load-from-stream.html.
 1# Load HTML from a stream using Python
 2
 3import os
 4import io
 5import aspose.html as ah
 6
 7# Prepare an output path for saving the document
 8output_dir = "output"
 9if not os.path.exists(output_dir):
10    os.makedirs(output_dir)
11
12# Use BytesIO instead of StringIO
13content_stream = io.BytesIO(b"<p>Hello, World!</p>")
14base_uri = "."
15
16# Initialize a document from the content stream
17document = ah.HTMLDocument(content_stream, base_uri)
18
19# Save the document to a disk
20document.save(os.path.join(output_dir, "load-from-stream.html"))

Load an SVG File in Python

Use SVGDocument from the aspose.html.dom.svg module for standalone SVG content. Like HTML, SVG is exposed through a DOM, but SVG-specific classes and methods belong to the SVG API. The SVG 2 specification defines the SVG document model, elements, and rendering behavior.

Load SVG from a File and Access Its DOM

The following example loads the local load-svg.svg file and reads an attribute from its <circle> element. Download the sample and place it in the data directory before running the code.

  1. Save the sample SVG as data/load-svg.svg.
  2. Pass the SVG file path to the SVGDocument constructor.
  3. Find the <circle> element with get_elements_by_tag_name().
  4. Read its fill attribute through the SVG DOM.
1# Load SVG from a file and access its DOM using Python
2
3import aspose.html.dom.svg as ahsvg
4
5# Load an SVG file from a local path
6with ahsvg.SVGDocument("data/load-svg.svg") as document:
7    # Access an element through the SVG DOM
8    circle = document.get_elements_by_tag_name("circle")[0]
9    print(circle.get_attribute("fill"))

The code prints #2F80ED, the fill color defined for the circle. The same DOM can be used to inspect elements, change attributes, or save the modified SVG.

Load SVG from a Stream

Use a stream when SVG markup is already available in memory. The existing example:

  1. Defines SVG markup containing a circle with a radius of 40 pixels.
  2. Encodes the markup and stores it in a BytesIO stream.
  3. Creates SVGDocument with the stream and a base URI.
  4. Prints the serialized SVG and saves it as load-from-stream.svg.
 1# Load SVG from a string using Python
 2
 3import io
 4import aspose.html.dom.svg as ahsvg
 5
 6# Initialize an SVG document from a string object
 7svg_content = "<svg xmlns='http://www.w3.org/2000/svg'><circle cx='50' cy='50' r='40'/></svg>"
 8base_uri = "."
 9content_stream = io.BytesIO(svg_content.encode('utf-8'))
10
11document = ahsvg.SVGDocument(content_stream, base_uri)
12
13# Write the document content to the output stream
14print(document.document_element.outer_html)
15
16# Save the document to a disk
17document.save("load-from-stream.svg")

For additional SVG-specific workflows, see the Aspose.SVG for Python via .NET documentation.

Work with MHTML and EPUB Sources

MHTML stores HTML and associated resources in a web archive, while EPUB packages publication content and resources. These formats are supported as conversion sources, but they are not loaded with an HTMLDocument constructor for DOM editing.

Common Document Loading Issues

IssueLikely CauseRecommended Action
Relative images or styles do not loadThe string or stream was created without a suitable base URI.Pass a base URI that relative resource paths can be resolved against.
Loading from a URL failsThe page or one of its required resources is unavailable or blocked by network settings.Verify the URL and network access, and handle loading errors in application code.
An empty document contains no visible contentHTMLDocument() creates the document structure but does not add page content.Create and append elements or text nodes before saving.
MHTML or EPUB cannot be edited with HTMLDocumentThese formats use source-specific conversion workflows.Use the MHTML or EPUB converter for the required output format.

Related Articles