Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To extract links from HTML in Java, load the file or webpage with
HTMLDocument, select anchor elements with
querySelectorAll("a[href]"), and read the link text and href attribute. Resolve relative references against the document base URI when absolute URLs are required.
An HTML link usually consists of visible anchor text and an href value. Aspose.HTML for Java parses the source into a DOM tree, allowing an application to select links and read their text and attributes without searching the HTML source with regular expressions.
This article demonstrates how to extract anchor text and original href values from a local HTML file and how to collect unique absolute HTTP and HTTPS links from a webpage.
The following example uses
extract-links-java.html. It selects every <a> element that has an href attribute and prints its text together with the original attribute value.
To extract links from an HTML file in Java:
HTMLDocument using a try-with-resources statement.querySelectorAll("a[href]") to select anchor elements that have an href attribute.Element.getTextContent().getAttribute("href"). 1import com.aspose.html.HTMLDocument;
2import com.aspose.html.collections.NodeList;
3import com.aspose.html.dom.Element;
4
5String inputPath = "extract-links-java.html";
6
7// Load an HTML document from a file
8try (HTMLDocument document = new HTMLDocument(inputPath)) {
9
10 // Select all anchor elements that have an href attribute
11 NodeList links = document.querySelectorAll("a[href]");
12
13 // Extract the text and original href value from each link
14 links.forEach(node -> {
15 Element link = (Element) node;
16 String text = link.getTextContent().trim();
17 String href = link.getAttribute("href").trim();
18
19 System.out.println(text + ": " + href);
20 });
21}The console output includes relative, absolute, fragment, and email references exactly as they appear in the source attributes:
Getting Started: /html/java/getting-started/Java API Reference: https://reference.aspose.com/html/java/Support information: #supportContact Support: mailto:support@aspose.com
Links collected from a webpage can contain relative paths, fragment references, email addresses, telephone links, or JavaScript URLs. The next example resolves relative references, keeps only HTTP and HTTPS destinations, skips same-page fragments and malformed values, and removes exact duplicate URL strings.
To extract absolute links from a webpage in Java:
HTMLDocument using a try-with-resources statement.getBaseURI() and create a Java URI from it.a[href] element in the loaded DOM.#.href against the base URI.LinkedHashSet to remove exact duplicates while preserving discovery order. 1import com.aspose.html.HTMLDocument;
2import com.aspose.html.collections.NodeList;
3import com.aspose.html.dom.Element;
4
5import java.net.URI;
6import java.util.LinkedHashSet;
7import java.util.Set;
8
9String pageUrl = "https://docs.aspose.com/html/java/data-extraction/";
10
11// Load an HTML document from a URL
12try (HTMLDocument document = new HTMLDocument(pageUrl)) {
13
14 // Get the base URI of the loaded document
15 URI baseUri = URI.create(document.getBaseURI());
16 Set<String> absoluteLinks = new LinkedHashSet<>();
17
18 // Select all links with an href attribute
19 NodeList links = document.querySelectorAll("a[href]");
20
21 links.forEach(node -> {
22 Element link = (Element) node;
23 String href = link.getAttribute("href").trim();
24
25 // Skip empty links and page fragments
26 if (href.isEmpty() || href.startsWith("#")) {
27 return;
28 }
29
30 try {
31 // Resolve relative URLs against the document base URI
32 URI absoluteUri = baseUri.resolve(href);
33 String scheme = absoluteUri.getScheme();
34
35 // Collect only HTTP and HTTPS links
36 if ("http".equalsIgnoreCase(scheme)
37 || "https".equalsIgnoreCase(scheme)) {
38 absoluteLinks.add(absoluteUri.toString());
39 }
40 } catch (IllegalArgumentException ignored) {
41 // Skip malformed URI values
42 }
43 });
44
45 // Display unique absolute links
46 absoluteLinks.forEach(System.out::println);
47}The result contains the unique HTTP and HTTPS destinations available in the DOM loaded from the page. URLs that differ by fragments, trailing slashes, letter case, or query-string order remain distinct because the example does not perform full URL canonicalization.
| Task | CSS selector | Result |
|---|---|---|
| Extract every hyperlink | a[href] | Anchor elements that have an href attribute |
| Extract links from main content | main a[href] or article a[href] | Links inside the selected content container |
| Extract download links | a[download][href] | Links that include both download and href attributes |
| Extract links to PDF files | a[href$='.pdf'] | Links whose href value ends with lowercase .pdf |
| Extract external absolute links | a[href^='http://'], a[href^='https://'] | Links already written as absolute HTTP or HTTPS URLs |
Attribute selectors inspect the source value. For example, a[href$='.pdf'] does not match an uppercase extension or a PDF URL followed by query parameters. Select a[href] and inspect normalized URLs in Java when broader detection is required.
| Issue | Cause | Fix |
|---|---|---|
| A relative URL is returned | getAttribute("href") returns the source attribute rather than resolving it. | Resolve the value against document.getBaseURI() when an absolute URL is required. |
| Duplicate links remain | The same destination appears in multiple anchor elements, or equivalent URLs use different spellings. | Use a Set for exact duplicates and add application-specific URL canonicalization when equivalent URLs must be combined. |
| Email, telephone, or JavaScript links appear | The selector matches any anchor with href, regardless of its URI scheme. | Resolve and keep only the schemes required by the application, such as HTTP and HTTPS. |
| A link visible in a browser is missing | Client-side code may insert it after the source document is loaded. | Inspect the DOM available to HTMLDocument and verify that the anchor exists before adjusting the selector. |
| Anchor text is empty | The link may contain only an image, icon, or empty content. | Read useful attributes or descendant content, such as an image alt value, when a text label is required. |
Select the anchor as an Element and call
getAttribute("href"). The method returns the attribute value from the HTML source, which may be relative rather than absolute.
Use document.getBaseURI() as the base and resolve each extracted href with a URI-resolution API. The second example uses java.net.URI.resolve() and then filters the resulting URLs by scheme.
No. The getAnchors() collection represents legacy named anchors rather than every hyperlink with an href value. Use querySelectorAll("a[href]") to extract hyperlinks.
No. Reading an href value only extracts the destination string. Follow or download a URL separately, and respect the target website’s access rules, terms of use, and rate limits.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.