Extract Images from a Website in Java

Aspose.HTML for Java can load a web page, inspect its DOM, resolve relative resource URLs, and request image files. The examples on this page extract images referenced by <img src> and icons referenced by <link rel="icon">.

To extract images from a website in Java, load the page into an HTMLDocument, find the relevant <img> or <link> elements, read their src or href attributes with getAttribute(), resolve relative references against document.getBaseURI(), and download each resource through the document network service.

Extract Images from HTML Image Elements

Many images in an HTML document are referenced by the src attribute of an <img> element. The following example collects those attribute values in a Set, converts relative references to absolute URLs, requests each resource, and saves successful responses.

  1. Create an HTMLDocument from the web page URL.
  2. Call getElementsByTagName(“img”) to collect the <img> elements.
  3. Read each element’s src attribute and collect distinct values.
  4. Resolve every relative reference against the document base URI with the Url class.
  5. Create and send a RequestMessage for each absolute URL.
  6. For a successful ResponseMessage, derive an output name and write the response bytes to the destination directory.
 1// Extract images from website using Java
 2
 3// Open a document you want to download images from
 4final HTMLDocument document = new HTMLDocument("https://docs.aspose.com/svg/net/drawing-basics/svg-shapes/");
 5
 6// Collect all <img> elements
 7HTMLCollection images = document.getElementsByTagName("img");
 8
 9// Create a distinct collection of relative image URLs
10Iterator<Element> iterator = images.iterator();
11java.util.Set<String> urls = new HashSet<>();
12for (Element e : images) {
13    urls.add(e.getAttribute("src"));
14}
15
16// Create absolute image URLs
17java.util.List<Url> absUrls = urls.stream()
18        .map(src -> new Url(src, document.getBaseURI()))
19        .collect(Collectors.toList());
20
21// foreach to while statements conversion
22for (Url url : absUrls) {
23    // Create an image request message
24    final RequestMessage request = new RequestMessage(url);
25
26    // Extract image
27    final ResponseMessage response = document.getContext().getNetwork().send(request);
28
29    // Check whether a response is successful
30    if (response.isSuccess()) {
31        String[] split = url.getPathname().split("/");
32        String path = split[split.length - 1];
33
34        // Save file to a local file system
35        FileHelper.writeAllBytes(path, response.getContent().readAsByteArray());
36    }
37}

The example finds only values of <img src>. It does not parse srcset, <picture><source>, CSS background-image, or custom lazy-loading attributes such as data-src.

Always respect the website’s terms of use, copyright, and permissions when downloading or reusing image files.

Extract Icons from Website

Website icons can be referenced by <link> elements. The current example selects links whose rel attribute is exactly "icon", reads their href values, resolves the URLs, and downloads the resources.

  1. Load the source web page into an HTMLDocument.
  2. Collect its <link> elements.
  3. Keep elements whose rel value equals "icon".
  4. Read distinct href values and resolve them against document.getBaseURI().
  5. Send a request for each resulting URL.
  6. Save the bytes returned by successful responses.
 1// Download icons from website using Java
 2
 3// Open a document you want to download icons from
 4final HTMLDocument document = new HTMLDocument("https://docs.aspose.com/html/net/message-handlers/");
 5
 6// Collect all <link> elements
 7HTMLCollection links = document.getElementsByTagName("link");
 8
 9// Leave only "icon" elements
10java.util.Set<Element> icons = new HashSet<>();
11for (Element link : links) {
12    if ("icon".equals(link.getAttribute("rel"))) {
13        icons.add(link);
14    }
15}
16
17// Create a distinct collection of relative icon URLs
18java.util.Set<String> urls = new HashSet<>();
19for (Element icon : icons) {
20    urls.add(icon.getAttribute("href"));
21}
22
23// Create absolute image URLs
24java.util.List<Url> absUrls = urls.stream()
25        .map(src -> new Url(src, document.getBaseURI()))
26        .collect(Collectors.toList());
27
28// foreach to while statements conversion
29for (Url url : absUrls) {
30    // Create a downloading request
31    final RequestMessage request = new RequestMessage(url);
32
33    // Extract icon
34    final ResponseMessage response = document.getContext().getNetwork().send(request);
35
36    // Check whether a response is successful
37    if (response.isSuccess()) {
38        String[] split = url.getPathname().split("/");
39        String path = split[split.length - 1];
40
41        // Save file to a local file system
42        FileHelper.writeAllBytes(path, response.getContent().readAsByteArray());
43    }
44}

The exact comparison with "icon" does not cover every icon declaration. An HTML rel attribute can contain multiple tokens, and websites may use values such as shortcut icon or apple-touch-icon.

Image Sources Covered by the Examples

Image sourceCovered?Notes
<img src="image.png">YesRelative and absolute src values are collected.
<link rel="icon" href="favicon.ico">YesThe icon example requires the exact rel="icon" value.
<img srcset="..."> or <picture>NoParse srcset and <source> elements separately.
CSS background-imageNoThe image URL is stored in CSS rather than an HTML src attribute.
Lazy-loading attributes such as data-srcNoAttribute names are site-specific and require additional extraction logic.
Inline data: imagesNot filteredDecide whether to decode them or skip them before creating network requests.

Common Image Extraction Issues

IssueLikely causeFix
The page URL is downloaded as though it were an imageAn <img> has an empty src, which resolves to the document URL.Remove null, empty, and whitespace-only attribute values before resolving URLs.
A data URI causes a request or file-name problemAn image is embedded directly in a data: URL.Skip data URIs or decode them separately.
One downloaded image overwrites anotherDifferent URLs have the same final path segment.Generate unique output names instead of relying only on the URL file name.
Some visible images are missingThey are defined through srcset, CSS, JavaScript, or a custom lazy-loading attribute.Inspect the loaded DOM and styles, then add extraction logic for the required source type.
A large image consumes substantial memoryreadAsByteArray() buffers the complete response before saving.Use an appropriate streaming workflow for large resources.

FAQ

Does the example extract every image visible on a website?

No. It extracts resources referenced by <img src> and exact <link rel="icon"> values. Responsive images, CSS backgrounds, data URIs, and site-specific lazy-loaded images require additional handling.

Can I extract image URLs from a local HTML file?

Yes. Load the local file into an HTMLDocument and use the same DOM queries. Relative references are resolved from the document base URI, and the referenced resources must remain accessible if the application also downloads them.

Related Aspose.HTML Articles

Other Platforms

Try HTML Web Applications

Aspose.HTML provides free online HTML Web Applications for converting files, extracting web data, generating HTML, and analyzing pages without installing additional software.

Text “HTML Web Applications”