Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To convert HTML to MHTML in Python, call Converter.convert_html() with an HTML file, HTMLDocument, or Url, MHTMLSaveOptions, and an output path. MHTML packages the HTML document and handled resources such as CSS and images into one web archive.
MHTML, or MIME HTML, is a web archive format described by
RFC 2557. It stores HTML markup and handled related resources as MIME parts in one .mht or .mhtml file. This makes MHTML useful for offline copies, document exchange, reference snapshots, and archiving webpages that would otherwise depend on separate resource files.
HTML and MHTML can represent the same webpage, but they store its resources differently:
| Feature | HTML | MHTML |
|---|---|---|
| Main content | HTML markup in an .html or .htm file | HTML markup stored as a MIME part in an .mht or .mhtml archive |
| CSS, images, fonts, and scripts | Embedded in HTML or referenced as separate files and URLs | Handled resources can be packaged as additional MIME parts in the same archive |
| Offline use | External files must remain available in the expected locations | Included resources can be read without their original files or network locations |
| Editing | Designed for direct editing and web publishing | Primarily intended for archiving and exchange |
| Viewer support | Supported by web browsers and HTML tools | Display depends on whether the selected browser or application supports MHTML |
An MHTML file is not a screenshot. It preserves handled markup and resources rather than converting the page into a fixed visual image. Dynamic behavior, authenticated resources, unavailable files, and viewer differences can affect the archived result.
The following example converts a local HTML file and packages its related stylesheet and SVG image into the MHTML archive:
MHTMLSaveOptions.Converter.convert_html() to create the archive. 1import os
2import aspose.html.converters as conv
3import aspose.html.saving as sav
4
5input_dir = "data"
6output_dir = "output"
7os.makedirs(output_dir, exist_ok=True)
8
9input_path = os.path.join(input_dir, "html-with-resources.html")
10output_path = os.path.join(output_dir, "html-with-resources.mhtml")
11
12options = sav.MHTMLSaveOptions()
13conv.Converter.convert_html(input_path, options, output_path)Place html-with-resources.html, html-resources.css, and html-resources.svg in the same input directory before running the example. Because the HTML file provides a source location, relative CSS and image URLs are resolved against its directory and the handled resources are stored in the MHTML file.
An HTML string has no file location from which relative resource URLs can be resolved. Provide a base URI when creating HTMLDocument so that referenced CSS, images, fonts, and scripts can be found.
HTMLDocument with the HTML and base URI.MHTMLSaveOptions. 1import os
2import aspose.html as ah
3import aspose.html.converters as conv
4import aspose.html.saving as sav
5
6input_dir = "data"
7output_dir = "output"
8os.makedirs(output_dir, exist_ok=True)
9output_path = os.path.join(output_dir, "string-with-resources.mhtml")
10
11html = """
12<link rel="stylesheet" href="html-resources.css">
13<h1>Monthly Report</h1>
14<img src="html-resources.svg" alt="Monthly sales chart">
15"""
16
17base_uri = os.path.abspath(input_dir) + os.sep
18
19with ah.HTMLDocument(html, base_uri) as document:
20 conv.Converter.convert_html(
21 document,
22 sav.MHTMLSaveOptions(),
23 output_path
24 )The base URI points to the directory containing the existing CSS and SVG files. Without it, the relative URLs have no reliable location and those resources may be missing from the archive.
Use the URL overload to load an online page and save the handled page resources in an MHTML archive:
Url for the webpage.MHTMLSaveOptions.Converter.convert_html(). 1import os
2import aspose.html as ah
3import aspose.html.converters as conv
4import aspose.html.saving as sav
5
6output_dir = "output"
7os.makedirs(output_dir, exist_ok=True)
8output_path = os.path.join(output_dir, "webpage.mhtml")
9
10url = ah.Url("https://docs.aspose.com/html/files/aspose.html")
11options = sav.MHTMLSaveOptions()
12
13conv.Converter.convert_html(url, options, output_path)The application must have network access to the page and its allowed resources. By default, resources on another host can be excluded by resource_url_restriction, so a webpage that loads CSS, images, fonts, or scripts from a CDN may require an intentionally broader restriction.
Resources required by the current page and links to other HTML pages are handled differently:
resource_url_restriction.<a href="page.html"> link points to another page. Linked pages are not included by default because max_handling_depth is 0.max_handling_depth affects linked HTML pages, not the CSS and images required to display the current page.The default resource handling behavior is SAVE, the default resource URL restriction is the same host, and the default linked-page depth is 0.
Set max_handling_depth when the archive should include HTML pages reached through links. The following example includes the page directly linked from the source document and restricts handled pages and resources to the same host:
MHTMLSaveOptions.max_handling_depth to 1 for directly linked pages. 1import os
2import aspose.html.converters as conv
3import aspose.html.saving as sav
4
5input_dir = "data"
6output_dir = "output"
7os.makedirs(output_dir, exist_ok=True)
8
9input_path = os.path.join(input_dir, "save-with-linked-page.html")
10output_path = os.path.join(output_dir, "linked-pages.mhtml")
11
12options = sav.MHTMLSaveOptions()
13options.resource_handling_options.max_handling_depth = 1
14options.resource_handling_options.page_url_restriction = (
15 sav.UrlRestriction.SAME_HOST
16)
17options.resource_handling_options.resource_url_restriction = (
18 sav.UrlRestriction.SAME_HOST
19)
20
21conv.Converter.convert_html(input_path, options, output_path)Place
save-with-linked-page.html and
linked-page.html in the same input directory. A depth of 0 processes only the current document, 1 includes directly linked pages, and -1 removes the depth limit. Avoid an unlimited depth for untrusted or uncontrolled websites because it can retrieve an unexpectedly large page set. URL restrictions still apply at every depth.
MHTMLSaveOptions provides
resource_handling_options, which returns a
ResourceHandlingOptions object.
| Property | What it controls | Default |
|---|---|---|
| default | Handling of CSS, images, fonts, and other general resources | SAVE |
| java_script | Handling of external JavaScript | SAVE |
| resource_url_restriction | Which CSS, image, font, script, and other resource URLs may be processed | Same host |
| page_url_restriction | Which linked HTML page URLs may be processed | Root location and subfolders |
| max_handling_depth | Maximum depth of linked HTML pages | 0 |
Use URL restrictions and depth together. ROOT_AND_SUB_FOLDERS is the narrowest normal page scope, SAME_HOST permits other paths on the same host, and NONE removes the URL restriction. Broaden these settings only for trusted sources and only when the archive is expected to include cross-folder or cross-host content.
| Problem | Check or fix |
|---|---|
| CSS or images are missing | Verify the source location or base URI, resource availability, and resource_url_restriction. The default same-host restriction can exclude CDN resources. |
| A linked HTML page is not included | Increase max_handling_depth and verify page_url_restriction. This setting affects linked pages rather than the current page’s CSS and images. |
| Too many pages are archived | Reduce max_handling_depth and narrow page_url_restriction. Do not use -1 for an uncontrolled website. |
| The archive still needs network access | Some resources were unavailable, excluded, dynamically requested, or not handled. Inspect the source URLs and resource restrictions. |
| The MHTML file opens differently in another application | MHTML support and rendering behavior vary by browser and viewer. Test the archive in the intended application. |
| Interactive behavior is missing | MHTML is an archive, not a recording of user interaction or server-side application state. |
Yes. .mht and .mhtml are commonly used file extensions for the MIME HTML web archive format.
It can include handled resources that are available and allowed by ResourceHandlingOptions. Resources that cannot be loaded or are excluded by URL restrictions are not packaged successfully.
A string has no source file or webpage location. The base URI supplies the reference location needed to resolve relative URLs such as styles/site.css or images/photo.png.
No. It controls how many levels of linked HTML pages are processed. CSS, images, fonts, scripts, and other resources required by the current page are controlled by resource handling behavior and resource_url_restriction.
Use the free online HTML to MHTML Converter below for occasional manual conversion. For programmatic conversion, you can evaluate the library without a license or follow the licensing guide to request a temporary license.
Use the free online HTML to MHTML Converter for quick manual conversion without writing code. Use Aspose.HTML for Python via .NET when webpage archiving must run in a Python application, service, or content-preservation workflow.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.