Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
To extract links from HTML in Python, load the file or webpage with
HTMLDocument, collect <a> elements with
get_elements_by_tag_name(“a”), and read their text_content and href attributes. Resolve relative references against
document.base_uri when absolute URLs are required.
An HTML link usually consists of visible anchor text and an href value. Aspose.HTML for Python via .NET parses the source into a DOM tree, allowing an application to inspect both values and decide which URL schemes, page areas, or duplicate links should be included.
This article demonstrates how to extract original link text and href values from a local HTML file and how to collect unique absolute HTTP and HTTPS links from a webpage.
The following example reuses
save-with-linked-page.html, a platform-neutral source file containing a relative hyperlink. It reads the link label and original href value without changing or resolving the attribute.
To extract links from an HTML file:
HTMLDocument in a context manager.<a> elements with get_elements_by_tag_name("a").href attribute with
get_attribute() and skip anchors where it is missing or empty.text_content value.href to another application component. 1# Extract links from a local HTML file in Python
2
3import aspose.html as ah
4
5# Load the HTML file and print non-empty links
6with ah.HTMLDocument("data/save-with-linked-page.html") as document:
7 links = document.get_elements_by_tag_name("a")
8
9 for link in links:
10 href = link.get_attribute("href")
11
12 if href is None:
13 continue
14
15 href = href.strip()
16
17 if not href:
18 continue
19
20 text = " ".join(link.text_content.split())
21 print(f"{text}: {href}")The example prints Open the linked page: linked-page.html. The relative value is returned exactly as it appears in the source attribute.
Webpage links can contain relative paths, fragments, email addresses, telephone links, or other URI schemes. The next example selects anchors from the page’s <main> element, resolves relative references, keeps HTTP and HTTPS links, removes exact duplicates, and saves the result as webpage-links.txt.
To extract absolute links from a webpage:
HTMLDocument constructor.urljoin() with document.base_uri to resolve relative references. 1# Extract links from a webpage in Python
2
3from urllib.parse import urljoin, urlparse
4import aspose.html as ah
5
6page_url = "https://products.aspose.com/html/python-net/"
7
8# Load the webpage and collect unique HTTP links
9with ah.HTMLDocument(page_url) as document:
10 main = document.query_selector("main")
11
12 if main is None:
13 raise RuntimeError("The webpage contains no main element.")
14
15 links = main.get_elements_by_tag_name("a")
16 absolute_links = []
17 seen = set()
18
19 for link in links:
20 href = link.get_attribute("href")
21
22 if href is None:
23 continue
24
25 href = href.strip()
26
27 if not href or href.startswith("#"):
28 continue
29
30 absolute_url = urljoin(document.base_uri, href)
31 scheme = urlparse(absolute_url).scheme.lower()
32
33 if scheme in ("http", "https") and absolute_url not in seen:
34 seen.add(absolute_url)
35 absolute_links.append(absolute_url)
36
37# Save the absolute links
38with open("webpage-links.txt", "w", encoding="utf-8") as file:
39 file.write("\n".join(absolute_links))The result contains unique HTTP and HTTPS destinations available inside the loaded main content. URLs that differ by fragments, trailing slashes, letter case, or query-string order remain distinct because the example removes exact duplicates rather than performing URL canonicalization.
| Task | Recommended approach | Result |
|---|---|---|
| Inspect every anchor | document.get_elements_by_tag_name("a") | All <a> elements, including anchors without href |
| Extract links from main content | Select main, then call main.get_elements_by_tag_name("a") | Anchors outside navigation and footer areas when the page uses <main> correctly |
| Keep original link values | Read get_attribute("href") | Relative, absolute, fragment, email, and other references as written in the source |
| Produce absolute URLs | Resolve each href against document.base_uri | Absolute destinations suitable for further URL processing |
| Keep web links only | Check for the http or https scheme | Excludes mailto, tel, javascript, and unsupported schemes |
| Issue | Cause | Recommended action |
|---|---|---|
| A relative URL is returned | get_attribute("href") reads the original source attribute. | Resolve it against document.base_uri when an absolute URL is required. |
Calling strip() fails | An anchor has no href, so get_attribute("href") returns None. | Check for None before trimming or processing the value. |
| Duplicate links remain | Equivalent destinations use different fragments, casing, trailing slashes, or query strings. | Add application-specific URL canonicalization when exact deduplication is insufficient. |
| Email or telephone links appear | All anchors were collected without filtering their URI schemes. | Keep only the schemes required by the application. |
| Link text is empty | The anchor contains an image, icon, or no text node. | Inspect descendant content or an image alt attribute when a label is required. |
| A browser shows a link but extraction misses it | The link is absent from the DOM available after loading or lies outside the selected container. | Inspect the loaded DOM and verify the extraction scope. |
Retrieve the anchor as an element and call get_attribute("href"). The result can be an absolute URL, a relative path, a fragment, another URI scheme, an empty string, or None when the attribute is absent.
Use document.base_uri as the base and resolve the extracted href with urllib.parse.urljoin(). The webpage example applies this step before filtering by URI scheme.
No. Reading an href extracts only the destination string. Download or follow the URL separately and observe the target website’s access rules and terms of use.
Yes. After collecting the link text and URL values, pass them to Python’s csv or json module instead of writing one URL per text line.
Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.