Convert HTML to MHTML in Python

To convert HTML to MHTML in Python, call Converter.convert_html() with an HTML file, HTMLDocument, or Url, MHTMLSaveOptions, and an output path. MHTML packages the HTML document and handled resources such as CSS and images into one web archive.

MHTML, or MIME HTML, is a web archive format described by RFC 2557. It stores HTML markup and handled related resources as MIME parts in one .mht or .mhtml file. This makes MHTML useful for offline copies, document exchange, reference snapshots, and archiving webpages that would otherwise depend on separate resource files.

HTML vs MHTML

HTML and MHTML can represent the same webpage, but they store its resources differently:

FeatureHTMLMHTML
Main contentHTML markup in an .html or .htm fileHTML markup stored as a MIME part in an .mht or .mhtml archive
CSS, images, fonts, and scriptsEmbedded in HTML or referenced as separate files and URLsHandled resources can be packaged as additional MIME parts in the same archive
Offline useExternal files must remain available in the expected locationsIncluded resources can be read without their original files or network locations
EditingDesigned for direct editing and web publishingPrimarily intended for archiving and exchange
Viewer supportSupported by web browsers and HTML toolsDisplay depends on whether the selected browser or application supports MHTML

An MHTML file is not a screenshot. It preserves handled markup and resources rather than converting the page into a fixed visual image. Dynamic behavior, authenticated resources, unavailable files, and viewer differences can affect the archived result.

Convert HTML to MHTML in Python

Convert an HTML File to MHTML

The following example converts a local HTML file and packages its related stylesheet and SVG image into the MHTML archive:

  1. Specify the source HTML file and output MHTML path.
  2. Create MHTMLSaveOptions.
  3. Call Converter.convert_html() to create the archive.
 1import os
 2import aspose.html.converters as conv
 3import aspose.html.saving as sav
 4
 5input_dir = "data"
 6output_dir = "output"
 7os.makedirs(output_dir, exist_ok=True)
 8
 9input_path = os.path.join(input_dir, "html-with-resources.html")
10output_path = os.path.join(output_dir, "html-with-resources.mhtml")
11
12options = sav.MHTMLSaveOptions()
13conv.Converter.convert_html(input_path, options, output_path)

Place html-with-resources.html, html-resources.css, and html-resources.svg in the same input directory before running the example. Because the HTML file provides a source location, relative CSS and image URLs are resolved against its directory and the handled resources are stored in the MHTML file.

Convert an HTML String to MHTML with a Base URI

An HTML string has no file location from which relative resource URLs can be resolved. Provide a base URI when creating HTMLDocument so that referenced CSS, images, fonts, and scripts can be found.

  1. Define the HTML string and its base directory or URL.
  2. Create an HTMLDocument with the HTML and base URI.
  3. Convert the document with MHTMLSaveOptions.
 1import os
 2import aspose.html as ah
 3import aspose.html.converters as conv
 4import aspose.html.saving as sav
 5
 6input_dir = "data"
 7output_dir = "output"
 8os.makedirs(output_dir, exist_ok=True)
 9output_path = os.path.join(output_dir, "string-with-resources.mhtml")
10
11html = """
12<link rel="stylesheet" href="html-resources.css">
13<h1>Monthly Report</h1>
14<img src="html-resources.svg" alt="Monthly sales chart">
15"""
16
17base_uri = os.path.abspath(input_dir) + os.sep
18
19with ah.HTMLDocument(html, base_uri) as document:
20    conv.Converter.convert_html(
21        document,
22        sav.MHTMLSaveOptions(),
23        output_path
24    )

The base URI points to the directory containing the existing CSS and SVG files. Without it, the relative URLs have no reliable location and those resources may be missing from the archive.

Archive a Webpage URL as MHTML

Use the URL overload to load an online page and save the handled page resources in an MHTML archive:

  1. Create a Url for the webpage.
  2. Create MHTMLSaveOptions.
  3. Pass the URL, options, and output path to Converter.convert_html().
 1import os
 2import aspose.html as ah
 3import aspose.html.converters as conv
 4import aspose.html.saving as sav
 5
 6output_dir = "output"
 7os.makedirs(output_dir, exist_ok=True)
 8output_path = os.path.join(output_dir, "webpage.mhtml")
 9
10url = ah.Url("https://docs.aspose.com/html/files/aspose.html")
11options = sav.MHTMLSaveOptions()
12
13conv.Converter.convert_html(url, options, output_path)

The application must have network access to the page and its allowed resources. By default, resources on another host can be excluded by resource_url_restriction, so a webpage that loads CSS, images, fonts, or scripts from a CDN may require an intentionally broader restriction.

Configure MHTML Resources and Linked Pages

Resources required by the current page and links to other HTML pages are handled differently:

The default resource handling behavior is SAVE, the default resource URL restriction is the same host, and the default linked-page depth is 0.

Include Linked Pages with MHTMLSaveOptions

Set max_handling_depth when the archive should include HTML pages reached through links. The following example includes the page directly linked from the source document and restricts handled pages and resources to the same host:

  1. Specify the main HTML file and output path.
  2. Create MHTMLSaveOptions.
  3. Set max_handling_depth to 1 for directly linked pages.
  4. Apply same-host URL restrictions.
  5. Convert the main document to MHTML.
 1import os
 2import aspose.html.converters as conv
 3import aspose.html.saving as sav
 4
 5input_dir = "data"
 6output_dir = "output"
 7os.makedirs(output_dir, exist_ok=True)
 8
 9input_path = os.path.join(input_dir, "save-with-linked-page.html")
10output_path = os.path.join(output_dir, "linked-pages.mhtml")
11
12options = sav.MHTMLSaveOptions()
13options.resource_handling_options.max_handling_depth = 1
14options.resource_handling_options.page_url_restriction = (
15    sav.UrlRestriction.SAME_HOST
16)
17options.resource_handling_options.resource_url_restriction = (
18    sav.UrlRestriction.SAME_HOST
19)
20
21conv.Converter.convert_html(input_path, options, output_path)

Place save-with-linked-page.html and linked-page.html in the same input directory. A depth of 0 processes only the current document, 1 includes directly linked pages, and -1 removes the depth limit. Avoid an unlimited depth for untrusted or uncontrolled websites because it can retrieve an unexpectedly large page set. URL restrictions still apply at every depth.

MHTMLSaveOptions and ResourceHandlingOptions

MHTMLSaveOptions provides resource_handling_options, which returns a ResourceHandlingOptions object.

PropertyWhat it controlsDefault
defaultHandling of CSS, images, fonts, and other general resourcesSAVE
java_scriptHandling of external JavaScriptSAVE
resource_url_restrictionWhich CSS, image, font, script, and other resource URLs may be processedSame host
page_url_restrictionWhich linked HTML page URLs may be processedRoot location and subfolders
max_handling_depthMaximum depth of linked HTML pages0

Use URL restrictions and depth together. ROOT_AND_SUB_FOLDERS is the narrowest normal page scope, SAME_HOST permits other paths on the same host, and NONE removes the URL restriction. Broaden these settings only for trusted sources and only when the archive is expected to include cross-folder or cross-host content.

Common HTML to MHTML Conversion Issues

ProblemCheck or fix
CSS or images are missingVerify the source location or base URI, resource availability, and resource_url_restriction. The default same-host restriction can exclude CDN resources.
A linked HTML page is not includedIncrease max_handling_depth and verify page_url_restriction. This setting affects linked pages rather than the current page’s CSS and images.
Too many pages are archivedReduce max_handling_depth and narrow page_url_restriction. Do not use -1 for an uncontrolled website.
The archive still needs network accessSome resources were unavailable, excluded, dynamically requested, or not handled. Inspect the source URLs and resource restrictions.
The MHTML file opens differently in another applicationMHTML support and rendering behavior vary by browser and viewer. Test the archive in the intended application.
Interactive behavior is missingMHTML is an archive, not a recording of user interaction or server-side application state.

Related Articles

Other Platforms

FAQ

Are MHT and MHTML the same format?

Yes. .mht and .mhtml are commonly used file extensions for the MIME HTML web archive format.

Does MHTML include CSS, images, and fonts?

It can include handled resources that are available and allowed by ResourceHandlingOptions. Resources that cannot be loaded or are excluded by URL restrictions are not packaged successfully.

Why is a base URI required for an HTML string?

A string has no source file or webpage location. The base URI supplies the reference location needed to resolve relative URLs such as styles/site.css or images/photo.png.

Does max_handling_depth control images and CSS?

No. It controls how many levels of linked HTML pages are processed. CSS, images, fonts, scripts, and other resources required by the current page are controlled by resource handling behavior and resource_url_restriction.

Can I convert HTML to MHTML for free?

Use the free online HTML to MHTML Converter below for occasional manual conversion. For programmatic conversion, you can evaluate the library without a license or follow the licensing guide to request a temporary license.

Try Online HTML to MHTML Conversion

                
            

Use the free online HTML to MHTML Converter for quick manual conversion without writing code. Use Aspose.HTML for Python via .NET when webpage archiving must run in a Python application, service, or content-preservation workflow.

Free Online HTML to MHTML Converter