Save File from URL in Python

Use Aspose.HTML for Python via .NET network APIs when an application has a direct resource URL and needs to save the response body locally. The workflow applies to images, stylesheets, scripts, documents, and other downloadable files.

To save a file from a URL in Python, create an HTMLDocument, build a RequestMessage for the resource URL, and call context.network.send(). If the response is successful, write response.content.read_as_byte_array() to a local binary file.

Download a File from a URL in Python

The following example downloads message-handlers.png from an absolute URL and saves it in the output directory. An empty HTMLDocument supplies the document context and its network service; it is not the content being downloaded.

To download and save a file from a URL:

  1. Create the output directory and an empty HTMLDocument.
  2. Create a Url for the resource to download.
  3. Pass the URL to RequestMessage and send it through document.context.network.
  4. Check response.is_success before reading the response body.
  5. Derive the local file name from url.pathname and write the returned byte array in binary mode.
 1# Download a file from a URL in Python
 2
 3import os
 4import aspose.html as ah
 5import aspose.html.net as ahnet
 6
 7# Prepare the output directory
 8output_dir = "output"
 9os.makedirs(output_dir, exist_ok=True)
10
11# Define the file URL
12url = ah.Url("https://docs.aspose.com/html/images/handlers/message-handlers.png")
13
14# Download and save the file
15with ah.HTMLDocument() as document:
16    with ahnet.RequestMessage(url) as request:
17        with document.context.network.send(request) as response:
18            if response.is_success:
19                file_path = os.path.join(output_dir, os.path.basename(url.pathname))
20                with open(file_path, "wb") as file:
21                    file.write(response.content.read_as_byte_array())

The example reads the entire response into memory before writing it. For a large file, account for memory consumption and use a streaming workflow appropriate to the application.

Choose the Correct URL Workflow

GoalRecommended workflow
Download one file from a known URLSend RequestMessage and save the response bytes as shown in this article.
Find file URLs in an HTML documentNavigate the DOM or use XPath or CSS selectors in HTML Navigation.
Save HTML together with CSS and imagesUse Save HTML with Resources in Python.
Convert a web page to PDF, DOCX, XPS, or an imageUse the appropriate workflow in the HTML Converter.

Common File Download Issues

IssueCause and recommended action
No output file is createdresponse.is_success is false, so the save block is skipped. Inspect the response and handle unsuccessful requests explicitly.
The output file name is emptyA URL ending in / has no file name in its path. Provide an explicit local file name instead of relying on os.path.basename(url.pathname).
The saved file has unexpected contentThe URL returned an error document or another media type. Validate the response status and content type before saving.
Memory usage is highread_as_byte_array() loads the complete response body into memory. Use a suitable streaming approach for large files.
The request failsThe URL is unavailable, blocked, or unsupported by the active network configuration. Verify the URL, protocol, permissions, and network access.

FAQ

Why does the example create an empty HTMLDocument?

The document provides access to document.context.network, which sends the RequestMessage. The empty document itself is not saved or used as the downloaded content.

Can this example download files other than images?

Yes. The response body is read as bytes, so the same workflow can save other downloadable file types. Use an appropriate local extension and validate the response before writing it.

Does this example save an entire web page?

No. It downloads one resource from a known URL. To preserve HTML with its linked CSS and images, use the resource-saving workflow linked above.

Other Platforms

Related Articles