Convert MHTML to PDF in Python

To convert MHTML to PDF in Python, open the MHTML file as a binary stream, create PdfSaveOptions, and call Converter.convert_mhtml() with the stream, options, and output PDF path.

MHTML, also called MIME HTML, can package HTML markup and related resources such as stylesheets and images in one .mhtml or .mht archive. Converting MHTML to PDF creates a fixed-layout document suitable for printing, sharing, archiving, and document workflows that do not use web archive files directly.

Convert MHTML to PDF in Python

The basic workflow uses the default PDF save settings:

  1. Open the source MHTML file as a binary stream.
  2. Create a PdfSaveOptions object.
  3. Pass the stream, options, and output path to Converter.convert_mhtml().
 1# Convert MHTML to PDF in Python
 2
 3import aspose.html.converters as conv
 4import aspose.html.saving as sav
 5
 6# Open the MHTML file
 7with open("document.mht", "rb") as stream:
 8
 9    # Configure PDF output
10    options = sav.PdfSaveOptions()
11
12    # Convert MHTML to PDF
13    conv.Converter.convert_mhtml(stream, options, "document.pdf")

The example opens an MHTML archive, converts it with the default PDF settings, and saves one PDF document containing all rendered pages. The with statement closes the input stream after conversion.

Customize MHTML to PDF with PdfSaveOptions

Use PdfSaveOptions when the PDF requires a specific page size, margins, background, metadata, image-compression setting, or encryption. Configure only the properties required by the output.

The following workflow sets custom page dimensions and margins during conversion:

  1. Create PdfSaveOptions for the PDF output.
  2. Configure the required page dimensions and margins through page_setup.
  3. Open the MHTML source as a binary stream and call Converter.convert_mhtml().
 1import os
 2import aspose.html.converters as conv
 3import aspose.html.drawing as dr
 4import aspose.html.saving as sav
 5
 6input_dir = "data"
 7output_dir = "output"
 8os.makedirs(output_dir, exist_ok=True)
 9
10input_path = os.path.join(input_dir, "document.mht")
11output_path = os.path.join(output_dir, "mhtml-options.pdf")
12
13options = sav.PdfSaveOptions()
14options.page_setup.any_page = dr.Page(
15    dr.Size(800, 600),
16    dr.Margin(20, 20, 20, 20)
17)
18with open(input_path, "rb") as stream:
19    conv.Converter.convert_mhtml(stream, options, output_path)

The numeric page and margin values are measured in pixels. Use explicit length units when the physical page dimensions must remain predictable across different output settings.

The most relevant PdfSaveOptions properties for MHTML conversion are:

PropertyUse it to control
page_setupPDF page size, margins, and page layout.
cssCSS processing, including the media type used for media queries.
background_colorThe color that fills each PDF page. The default is transparent.
jpeg_qualityJPEG compression quality when JPEG compression is used. It does not increase the resolution of source images.
document_infoPDF metadata such as title, author, subject, and keywords.
encryptionPDF passwords, permissions, and encryption settings.
is_tagged_pdfWhether tagged PDF structure is created in the output.

For additional rendering controls, see Fine-Tuning Converters.

How MHTML Resources Are Converted to PDF

An MHTML archive normally stores the main HTML document and captured resources as MIME parts. During conversion, Aspose.HTML resolves those parts and renders the resulting HTML and CSS into PDF pages. Images, styles, and fonts can be reproduced when they are present in the archive and their references are valid.

Some MHTML files still refer to external resources instead of embedding every dependency. Those resources must remain accessible during conversion. If an image, stylesheet, or font is missing from the PDF, inspect the archive contents and its resource URLs before changing PDF save settings.

PDF uses fixed pages, while the archived webpage may have been designed for a browser viewport. Page size, margins, fixed-width elements, styles, and available fonts can therefore change line wrapping and page breaks.

Common MHTML to PDF Issues

IssueLikely cause and solution
Images or styles are missingThe resource is absent from the archive, its MIME part is invalid, or an external reference cannot be accessed. Check the MHTML source and its resource URLs.
The PDF looks different from the webpagePDF uses fixed pages, while webpages commonly use responsive layouts. Check page dimensions, margins, fonts, and fixed-width page styles.
Text uses a different fontThe archived font is unavailable or was not embedded. Make the required font accessible in the conversion environment.
Content is clipped or surrounded by excessive spaceThe configured PDF page does not match the content layout. Adjust page_setup, margins, or fixed-width CSS rules.
The PDF has more pages than expectedBrowser content is being paginated into a fixed-layout document. Page size, font substitution, CSS page-break rules, and margins can change the page count.

Related MHTML Conversion Guides

Other Platforms

FAQ

Can I convert MHTML to PDF for free?

You can evaluate Aspose.HTML for Python via .NET without a license, but evaluation output has limitations such as a watermark and processing limits. Request a free temporary license for unrestricted testing, or use the free online MHTML to PDF Converter for an occasional manual conversion. See Licensing for license setup.

Does MHTML to PDF conversion include images and CSS?

The converter renders images and CSS stored in valid MHTML parts. A resource that was not embedded in the archive must still be accessible through its external reference during conversion.

Are MHT and MHTML different input formats?

.mht and .mhtml are commonly used file extensions for the same MIME HTML web archive format. Open either file type as a binary stream and pass it to Converter.convert_mhtml().

Try Online MHTML to PDF Conversion

                
            

Use the free online MHTML to PDF Converter for quick manual conversion without writing code. Use Aspose.HTML for Python via .NET when MHTML-to-PDF conversion must run programmatically in an application, service, or batch workflow.

Download complete examples and data files from GitHub.

Free Online MHTML to PDF Converter