Extract Images from a Website in Python

Web pages commonly reference images through <img src> elements and page icons through <link rel="icon">. Aspose.HTML for Python via .NET can load the page DOM, locate these elements, resolve their URLs, and request the associated files.

To extract images from a website in Python, load the page with HTMLDocument, find image elements, read their src or href attributes, and resolve relative references with Url. Send a RequestMessage through the document network service and save successful responses as binary files.

Extract Images from img Elements

The first example loads a web page and downloads the files referenced by its <img src> attributes. A Python set removes duplicate source values before requests are sent.

To extract images from a website:

  1. Create the output directory and load the page URL into an HTMLDocument.
  2. Call get_elements_by_tag_name(“img”) to collect image elements.
  3. Read each src attribute with get_attribute() and keep distinct values in a set.
  4. Resolve every source against document.base_uri to create an absolute URL.
  5. Send a RequestMessage through document.context.network.
  6. If the response is successful, derive the file name from the URL path and write the returned bytes to the output directory.
 1# Extract images from a website in Python
 2
 3import os
 4import aspose.html as ah
 5import aspose.html.net as ahnet
 6
 7# Prepare the output directory
 8output_dir = "output"
 9os.makedirs(output_dir, exist_ok=True)
10
11# Load the webpage and collect image URLs
12with ah.HTMLDocument("https://docs.aspose.com/html/net/tutorial/html-colors/") as document:
13    # Collect all <img> elements
14    images = document.get_elements_by_tag_name("img")
15
16    urls = {image.get_attribute("src") for image in images if image.get_attribute("src")}
17
18    # Download each image
19    for image_url in urls:
20        url = ah.Url(image_url, document.base_uri)
21        file_name = os.path.basename(url.pathname)
22
23        if not file_name:
24            continue
25
26        with ahnet.RequestMessage(url) as request:
27            with document.context.network.send(request) as response:
28                if response.is_success:
29                    with open(os.path.join(output_dir, file_name), "wb") as file:
30                        file.write(response.content.read_as_byte_array())

Download and reuse website images only when permitted by the website terms, applicable licenses, and copyright law.

Extract Icons from a Website

Page icons are commonly referenced by <link> elements in the document <head>. The second example selects links whose rel attribute is exactly icon, resolves each href, and saves successful responses under output/icons.

To extract page icons:

  1. Load the website URL into an HTMLDocument.
  2. Collect all <link> elements and keep those for which rel == "icon".
  3. Read distinct href values and resolve them against the document base URI.
  4. Send a request for each absolute icon URL.
  5. Save every successful response to the output/icons directory.
 1# Extract icons from a website in Python
 2
 3import os
 4import aspose.html as ah
 5import aspose.html.net as ahnet
 6
 7# Prepare the output directory
 8output_dir = os.path.join("output", "icons")
 9os.makedirs(output_dir, exist_ok=True)
10
11# Load the webpage and collect icon URLs
12with ah.HTMLDocument("https://docs.aspose.com/html/python-net/") as document:
13    links = document.get_elements_by_tag_name("link")
14    icon_urls = {
15        link.get_attribute("href")
16        for link in links
17        if "icon" in (link.get_attribute("rel") or "").lower().split()
18        and link.get_attribute("href")
19    }
20
21    # Download each icon
22    for icon_url in icon_urls:
23        url = ah.Url(icon_url, document.base_uri)
24        file_name = os.path.basename(url.pathname)
25
26        if not file_name:
27            continue
28
29        with ahnet.RequestMessage(url) as request:
30            with document.context.network.send(request) as response:
31                if response.is_success:
32                    file_path = os.path.join(output_dir, file_name)
33                    with open(file_path, "wb") as file:
34                        file.write(response.content.read_as_byte_array())

Which Website Images Do These Examples Extract?

Image sourceCovered by the examples?
<img src="image.png">Yes, the first example reads src.
<link rel="icon" href="favicon.ico">Yes, the second example reads href when rel is exactly icon.
srcset, <picture>, or lazy-load attributes such as data-srcNo. Read and process these attributes separately when the page uses them.
CSS background-imageNo. Inspect stylesheets or computed style data separately.
Inline or external SVGUse Extract SVG from Website.

Common Image Extraction Issues

IssueCause and recommended action
An image URL is emptyAn <img> element has no usable src. Filter empty values before constructing Url objects.
A data URI cannot be saved with a useful file nameA data: reference has no normal URL path. Detect and decode data URIs separately or skip them.
An output file has an empty nameThe resource URL ends in / or has no path file name. Assign an explicit local name.
One downloaded image overwrites anotherDifferent URLs can have the same basename. Generate unique output names when collisions are possible.
Some visible images are missingThe page uses srcset, lazy loading, CSS backgrounds, scripts, or another image source not covered by img[src]. Inspect the loaded DOM and page-specific attributes.
An icon is not foundThe code checks for an exact rel="icon" value. A page may use another token combination, such as shortcut icon. Parse rel as a token list when broader coverage is required.
A saved file is not an imageA request returned an error page or another media type. Check the response and content type before writing the file.

FAQ

Does this code extract every image from a website?

No. The first example downloads resources referenced by <img src>. Images defined through srcset, <picture>, CSS, lazy-load attributes, scripts, or data URIs require additional page-specific handling.

How are relative image URLs resolved?

The examples pass each relative reference and document.base_uri to Url, producing the absolute URL required by the network request.

Does the code remove duplicate image URLs?

Yes. A Python set removes repeated src or href values. Different URLs with the same file name can still overwrite one another in the output directory.

Other Platforms

Related Articles