Extract Links from HTML in Python

To extract links from HTML in Python, load the file or webpage with HTMLDocument, collect <a> elements with get_elements_by_tag_name(“a”), and read their text_content and href attributes. Resolve relative references against document.base_uri when absolute URLs are required.

An HTML link usually consists of visible anchor text and an href value. Aspose.HTML for Python via .NET parses the source into a DOM tree, allowing an application to inspect both values and decide which URL schemes, page areas, or duplicate links should be included.

This article demonstrates how to extract original link text and href values from a local HTML file and how to collect unique absolute HTTP and HTTPS links from a webpage.

The following example reuses save-with-linked-page.html, a platform-neutral source file containing a relative hyperlink. It reads the link label and original href value without changing or resolving the attribute.

To extract links from an HTML file:

  1. Load the local file with HTMLDocument in a context manager.
  2. Collect all <a> elements with get_elements_by_tag_name("a").
  3. Read the href attribute with get_attribute() and skip anchors where it is missing or empty.
  4. Normalize the anchor’s text_content value.
  5. Print or pass the extracted text and original href to another application component.
 1# Extract links from a local HTML file in Python
 2
 3import aspose.html as ah
 4
 5# Load the HTML file and print non-empty links
 6with ah.HTMLDocument("data/save-with-linked-page.html") as document:
 7    links = document.get_elements_by_tag_name("a")
 8
 9    for link in links:
10        href = link.get_attribute("href")
11
12        if href is None:
13            continue
14
15        href = href.strip()
16
17        if not href:
18            continue
19
20        text = " ".join(link.text_content.split())
21        print(f"{text}: {href}")

The example prints Open the linked page: linked-page.html. The relative value is returned exactly as it appears in the source attribute.

Webpage links can contain relative paths, fragments, email addresses, telephone links, or other URI schemes. The next example selects anchors from the page’s <main> element, resolves relative references, keeps HTTP and HTTPS links, removes exact duplicates, and saves the result as webpage-links.txt.

To extract absolute links from a webpage:

  1. Pass the webpage URL to the HTMLDocument constructor.
  2. Select the relevant content area and collect its descendant anchors.
  3. Skip missing, empty, and same-page fragment references.
  4. Use urljoin() with document.base_uri to resolve relative references.
  5. Keep only URLs whose schemes are HTTP or HTTPS.
  6. Use a set to remove exact duplicates while retaining discovery order in a list.
  7. Save the absolute URLs to a UTF-8 text file.
 1# Extract links from a webpage in Python
 2
 3from urllib.parse import urljoin, urlparse
 4import aspose.html as ah
 5
 6page_url = "https://products.aspose.com/html/python-net/"
 7
 8# Load the webpage and collect unique HTTP links
 9with ah.HTMLDocument(page_url) as document:
10    main = document.query_selector("main")
11
12    if main is None:
13        raise RuntimeError("The webpage contains no main element.")
14
15    links = main.get_elements_by_tag_name("a")
16    absolute_links = []
17    seen = set()
18
19    for link in links:
20        href = link.get_attribute("href")
21
22        if href is None:
23            continue
24
25        href = href.strip()
26
27        if not href or href.startswith("#"):
28            continue
29
30        absolute_url = urljoin(document.base_uri, href)
31        scheme = urlparse(absolute_url).scheme.lower()
32
33        if scheme in ("http", "https") and absolute_url not in seen:
34            seen.add(absolute_url)
35            absolute_links.append(absolute_url)
36
37# Save the absolute links
38with open("webpage-links.txt", "w", encoding="utf-8") as file:
39    file.write("\n".join(absolute_links))

The result contains unique HTTP and HTTPS destinations available inside the loaded main content. URLs that differ by fragments, trailing slashes, letter case, or query-string order remain distinct because the example removes exact duplicates rather than performing URL canonicalization.

TaskRecommended approachResult
Inspect every anchordocument.get_elements_by_tag_name("a")All <a> elements, including anchors without href
Extract links from main contentSelect main, then call main.get_elements_by_tag_name("a")Anchors outside navigation and footer areas when the page uses <main> correctly
Keep original link valuesRead get_attribute("href")Relative, absolute, fragment, email, and other references as written in the source
Produce absolute URLsResolve each href against document.base_uriAbsolute destinations suitable for further URL processing
Keep web links onlyCheck for the http or https schemeExcludes mailto, tel, javascript, and unsupported schemes

Common Link Extraction Issues

IssueCauseRecommended action
A relative URL is returnedget_attribute("href") reads the original source attribute.Resolve it against document.base_uri when an absolute URL is required.
Calling strip() failsAn anchor has no href, so get_attribute("href") returns None.Check for None before trimming or processing the value.
Duplicate links remainEquivalent destinations use different fragments, casing, trailing slashes, or query strings.Add application-specific URL canonicalization when exact deduplication is insufficient.
Email or telephone links appearAll anchors were collected without filtering their URI schemes.Keep only the schemes required by the application.
Link text is emptyThe anchor contains an image, icon, or no text node.Inspect descendant content or an image alt attribute when a label is required.
A browser shows a link but extraction misses itThe link is absent from the DOM available after loading or lies outside the selected container.Inspect the loaded DOM and verify the extraction scope.

FAQ

How do I get the href value of a link in Python?

Retrieve the anchor as an element and call get_attribute("href"). The result can be an absolute URL, a relative path, a fragment, another URI scheme, an empty string, or None when the attribute is absent.

How do I convert a relative link to an absolute URL?

Use document.base_uri as the base and resolve the extracted href with urllib.parse.urljoin(). The webpage example applies this step before filtering by URI scheme.

Does extracting a link download its destination?

No. Reading an href extracts only the destination string. Download or follow the URL separately and observe the target website’s access rules and terms of use.

Can I save extracted links as CSV or JSON?

Yes. After collecting the link text and URL values, pass them to Python’s csv or json module instead of writing one URL per text line.

Related Articles