Extract Links from HTML in Java

To extract links from HTML in Java, load the file or webpage with HTMLDocument, select anchor elements with querySelectorAll("a[href]"), and read the link text and href attribute. Resolve relative references against the document base URI when absolute URLs are required.

An HTML link usually consists of visible anchor text and an href value. Aspose.HTML for Java parses the source into a DOM tree, allowing an application to select links and read their text and attributes without searching the HTML source with regular expressions.

This article demonstrates how to extract anchor text and original href values from a local HTML file and how to collect unique absolute HTTP and HTTPS links from a webpage.

The following example uses extract-links-java.html. It selects every <a> element that has an href attribute and prints its text together with the original attribute value.

To extract links from an HTML file in Java:

  1. Open the local HTML file in an HTMLDocument using a try-with-resources statement.
  2. Call querySelectorAll("a[href]") to select anchor elements that have an href attribute.
  3. Cast each matching node to Element.
  4. Read and trim the anchor text with getTextContent().
  5. Read the original link destination with getAttribute("href").
  6. Print or pass the extracted values to another application component.
 1import com.aspose.html.HTMLDocument;
 2import com.aspose.html.collections.NodeList;
 3import com.aspose.html.dom.Element;
 4
 5String inputPath = "extract-links-java.html";
 6
 7// Load an HTML document from a file
 8try (HTMLDocument document = new HTMLDocument(inputPath)) {
 9
10    // Select all anchor elements that have an href attribute
11    NodeList links = document.querySelectorAll("a[href]");
12
13    // Extract the text and original href value from each link
14    links.forEach(node -> {
15        Element link = (Element) node;
16        String text = link.getTextContent().trim();
17        String href = link.getAttribute("href").trim();
18
19        System.out.println(text + ": " + href);
20    });
21}

The console output includes relative, absolute, fragment, and email references exactly as they appear in the source attributes:

Getting Started: /html/java/getting-started/
Java API Reference: https://reference.aspose.com/html/java/
Support information: #support
Contact Support: mailto:support@aspose.com

Links collected from a webpage can contain relative paths, fragment references, email addresses, telephone links, or JavaScript URLs. The next example resolves relative references, keeps only HTTP and HTTPS destinations, skips same-page fragments and malformed values, and removes exact duplicate URL strings.

To extract absolute links from a webpage in Java:

  1. Open the webpage URL in an HTMLDocument using a try-with-resources statement.
  2. Read the document base URI with getBaseURI() and create a Java URI from it.
  3. Select every a[href] element in the loaded DOM.
  4. Skip empty values and references that begin with #.
  5. Resolve each remaining href against the base URI.
  6. Keep HTTP and HTTPS URLs in a LinkedHashSet to remove exact duplicates while preserving discovery order.
  7. Print or process the resulting absolute URLs.
 1import com.aspose.html.HTMLDocument;
 2import com.aspose.html.collections.NodeList;
 3import com.aspose.html.dom.Element;
 4
 5import java.net.URI;
 6import java.util.LinkedHashSet;
 7import java.util.Set;
 8
 9String pageUrl = "https://docs.aspose.com/html/java/data-extraction/";
10
11// Load an HTML document from a URL
12try (HTMLDocument document = new HTMLDocument(pageUrl)) {
13
14    // Get the base URI of the loaded document
15    URI baseUri = URI.create(document.getBaseURI());
16    Set<String> absoluteLinks = new LinkedHashSet<>();
17
18    // Select all links with an href attribute
19    NodeList links = document.querySelectorAll("a[href]");
20
21    links.forEach(node -> {
22        Element link = (Element) node;
23        String href = link.getAttribute("href").trim();
24
25        // Skip empty links and page fragments
26        if (href.isEmpty() || href.startsWith("#")) {
27            return;
28        }
29
30        try {
31            // Resolve relative URLs against the document base URI
32            URI absoluteUri = baseUri.resolve(href);
33            String scheme = absoluteUri.getScheme();
34
35            // Collect only HTTP and HTTPS links
36            if ("http".equalsIgnoreCase(scheme)
37                    || "https".equalsIgnoreCase(scheme)) {
38                absoluteLinks.add(absoluteUri.toString());
39            }
40        } catch (IllegalArgumentException ignored) {
41            // Skip malformed URI values
42        }
43    });
44
45    // Display unique absolute links
46    absoluteLinks.forEach(System.out::println);
47}

The result contains the unique HTTP and HTTPS destinations available in the DOM loaded from the page. URLs that differ by fragments, trailing slashes, letter case, or query-string order remain distinct because the example does not perform full URL canonicalization.

TaskCSS selectorResult
Extract every hyperlinka[href]Anchor elements that have an href attribute
Extract links from main contentmain a[href] or article a[href]Links inside the selected content container
Extract download linksa[download][href]Links that include both download and href attributes
Extract links to PDF filesa[href$='.pdf']Links whose href value ends with lowercase .pdf
Extract external absolute linksa[href^='http://'], a[href^='https://']Links already written as absolute HTTP or HTTPS URLs

Attribute selectors inspect the source value. For example, a[href$='.pdf'] does not match an uppercase extension or a PDF URL followed by query parameters. Select a[href] and inspect normalized URLs in Java when broader detection is required.

Common Link Extraction Issues

IssueCauseFix
A relative URL is returnedgetAttribute("href") returns the source attribute rather than resolving it.Resolve the value against document.getBaseURI() when an absolute URL is required.
Duplicate links remainThe same destination appears in multiple anchor elements, or equivalent URLs use different spellings.Use a Set for exact duplicates and add application-specific URL canonicalization when equivalent URLs must be combined.
Email, telephone, or JavaScript links appearThe selector matches any anchor with href, regardless of its URI scheme.Resolve and keep only the schemes required by the application, such as HTTP and HTTPS.
A link visible in a browser is missingClient-side code may insert it after the source document is loaded.Inspect the DOM available to HTMLDocument and verify that the anchor exists before adjusting the selector.
Anchor text is emptyThe link may contain only an image, icon, or empty content.Read useful attributes or descendant content, such as an image alt value, when a text label is required.

FAQ

How do I get the href value of a link in Java?

Select the anchor as an Element and call getAttribute("href"). The method returns the attribute value from the HTML source, which may be relative rather than absolute.

How do I convert relative links to absolute URLs?

Use document.getBaseURI() as the base and resolve each extracted href with a URI-resolution API. The second example uses java.net.URI.resolve() and then filters the resulting URLs by scheme.

Should I use getAnchors() to extract every hyperlink?

No. The getAnchors() collection represents legacy named anchors rather than every hyperlink with an href value. Use querySelectorAll("a[href]") to extract hyperlinks.

Does extracting a link download its destination?

No. Reading an href value only extracts the destination string. Follow or download a URL separately, and respect the target website’s access rules, terms of use, and rate limits.

Related Data Extraction Articles