Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To convert MHTML to PDF in Python, open the MHTML file as a binary stream, create PdfSaveOptions, and call Converter.convert_mhtml() with the stream, options, and output PDF path.
MHTML, also called MIME HTML, can package HTML markup and related resources such as stylesheets and images in one .mhtml or .mht archive. Converting MHTML to PDF creates a fixed-layout document suitable for printing, sharing, archiving, and document workflows that do not use web archive files directly.
The basic workflow uses the default PDF save settings:
PdfSaveOptions object.Converter.convert_mhtml(). 1# Convert MHTML to PDF in Python
2
3import aspose.html.converters as conv
4import aspose.html.saving as sav
5
6# Open the MHTML file
7with open("document.mht", "rb") as stream:
8
9 # Configure PDF output
10 options = sav.PdfSaveOptions()
11
12 # Convert MHTML to PDF
13 conv.Converter.convert_mhtml(stream, options, "document.pdf")The example opens an MHTML archive, converts it with the default PDF settings, and saves one PDF document containing all rendered pages. The with statement closes the input stream after conversion.
Use PdfSaveOptions when the PDF requires a specific page size, margins, background, metadata, image-compression setting, or encryption. Configure only the properties required by the output.
The following workflow sets custom page dimensions and margins during conversion:
PdfSaveOptions for the PDF output.page_setup.Converter.convert_mhtml(). 1import os
2import aspose.html.converters as conv
3import aspose.html.drawing as dr
4import aspose.html.saving as sav
5
6input_dir = "data"
7output_dir = "output"
8os.makedirs(output_dir, exist_ok=True)
9
10input_path = os.path.join(input_dir, "document.mht")
11output_path = os.path.join(output_dir, "mhtml-options.pdf")
12
13options = sav.PdfSaveOptions()
14options.page_setup.any_page = dr.Page(
15 dr.Size(800, 600),
16 dr.Margin(20, 20, 20, 20)
17)
18with open(input_path, "rb") as stream:
19 conv.Converter.convert_mhtml(stream, options, output_path)The numeric page and margin values are measured in pixels. Use explicit length units when the physical page dimensions must remain predictable across different output settings.
The most relevant PdfSaveOptions properties for MHTML conversion are:
| Property | Use it to control |
|---|---|
| page_setup | PDF page size, margins, and page layout. |
| css | CSS processing, including the media type used for media queries. |
| background_color | The color that fills each PDF page. The default is transparent. |
| jpeg_quality | JPEG compression quality when JPEG compression is used. It does not increase the resolution of source images. |
| document_info | PDF metadata such as title, author, subject, and keywords. |
| encryption | PDF passwords, permissions, and encryption settings. |
| is_tagged_pdf | Whether tagged PDF structure is created in the output. |
For additional rendering controls, see Fine-Tuning Converters.
An MHTML archive normally stores the main HTML document and captured resources as MIME parts. During conversion, Aspose.HTML resolves those parts and renders the resulting HTML and CSS into PDF pages. Images, styles, and fonts can be reproduced when they are present in the archive and their references are valid.
Some MHTML files still refer to external resources instead of embedding every dependency. Those resources must remain accessible during conversion. If an image, stylesheet, or font is missing from the PDF, inspect the archive contents and its resource URLs before changing PDF save settings.
PDF uses fixed pages, while the archived webpage may have been designed for a browser viewport. Page size, margins, fixed-width elements, styles, and available fonts can therefore change line wrapping and page breaks.
| Issue | Likely cause and solution |
|---|---|
| Images or styles are missing | The resource is absent from the archive, its MIME part is invalid, or an external reference cannot be accessed. Check the MHTML source and its resource URLs. |
| The PDF looks different from the webpage | PDF uses fixed pages, while webpages commonly use responsive layouts. Check page dimensions, margins, fonts, and fixed-width page styles. |
| Text uses a different font | The archived font is unavailable or was not embedded. Make the required font accessible in the conversion environment. |
| Content is clipped or surrounded by excessive space | The configured PDF page does not match the content layout. Adjust page_setup, margins, or fixed-width CSS rules. |
| The PDF has more pages than expected | Browser content is being paginated into a fixed-layout document. Page size, font substitution, CSS page-break rules, and margins can change the page count. |
You can evaluate Aspose.HTML for Python via .NET without a license, but evaluation output has limitations such as a watermark and processing limits. Request a free temporary license for unrestricted testing, or use the free online MHTML to PDF Converter for an occasional manual conversion. See Licensing for license setup.
The converter renders images and CSS stored in valid MHTML parts. A resource that was not embedded in the archive must still be accessible through its external reference during conversion.
.mht and .mhtml are commonly used file extensions for the same MIME HTML web archive format. Open either file type as a binary stream and pass it to Converter.convert_mhtml().
Use the free online MHTML to PDF Converter for quick manual conversion without writing code. Use Aspose.HTML for Python via .NET when MHTML-to-PDF conversion must run programmatically in an application, service, or batch workflow.
Download complete examples and data files from GitHub.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.