Select HTML Nodes with XPath in Python

Load HTML with HTMLDocument, pass an XPath expression to evaluate(), request XPathResultType.ANY, and retrieve matching nodes by repeatedly calling iterate_next().

XPath is useful when a selection depends on document structure, position, attributes, or text. Unlike a CSS selector, an XPath expression can also return attribute nodes or values rather than their owning elements.

The examples below first evaluate an expression against HTML created from a string. They then use XPath to isolate photo images in an existing file and return their src attributes directly.

Run an XPath Query in Python

The call evaluate(expression, context_node, resolver, type, result) evaluates an XPath expression relative to a context node. For an ordinary document-wide HTML query, pass the document as the context node, None as the namespace resolver, XPathResultType.ANY as the requested type, and None when there is no existing result object to reuse.

The following example selects only paragraphs whose data-status value is published:

  1. Load the HTML source into an HTMLDocument.
  2. Write an XPath expression for the required structural, attribute, position, or text conditions.
  3. Pass the expression and context node to evaluate() and request an appropriate result type.
  4. Retrieve and process the returned nodes or value.
 1# Select HTML elements with XPath in Python
 2
 3import aspose.html as ah
 4import aspose.html.dom.xpath as hxpath
 5
 6# Define the HTML content
 7html_code = """
 8<section>
 9    <p data-status="published">Getting Started</p>
10    <p data-status="draft">Draft tutorial</p>
11    <p data-status="published">API Reference</p>
12</section>
13"""
14
15# Evaluate XPath and print the matching paragraphs
16with ah.HTMLDocument(html_code, ".") as document:
17    result = document.evaluate(
18        "//p[@data-status = 'published']",
19        document,
20        None,
21        hxpath.XPathResultType.ANY,
22        None
23    )
24
25    node = result.iterate_next()
26    while node is not None:
27        print(node.text_content)
28        node = result.iterate_next()
Example-UseXPath.py hosted with ❤ by GitHub

The predicate inside square brackets removes the draft paragraph from the result. The output is:

1Getting Started
2API Reference

Narrow an XPath Query for Existing HTML

The sample xpath-image.htm deliberately mixes photo images with advertising images. The ads appear in the header, footer, separate rows inside <main>, and individual photo rows. This makes the file useful for showing how each part of an XPath expression changes the result.

For this sample, the expression is refined in four stages:

The final element query is:

1//main/div[position() mod 2 = 1]//img[@class = 'photo']

This predicate requires the complete class attribute to equal photo. If photo may be one of several classes, use a token-aware test:

1//main/div[position() mod 2 = 1]//img[contains(concat(' ', normalize-space(@class), ' '), ' photo ')]

Return Image src Attributes Directly

XPath can select the src attributes instead of returning image elements. Append /@src to the element query, iterate through the resulting attribute nodes, and read each value through node_value. This avoids an element cast in Python.

  1. Load the HTML source into an HTMLDocument.
  2. Write an XPath expression that returns the required attribute nodes.
  3. Evaluate the expression with the document or another suitable context node.
  4. Iterate through the result and read node_value from each attribute node.
 1# Extract image src attributes with XPath in Python
 2
 3import os
 4import aspose.html as ah
 5import aspose.html.dom.xpath as hxpath
 6
 7# Prepare the input path and XPath expression
 8data_dir = "data"
 9input_path = os.path.join(data_dir, "xpath-image.htm")
10expression = "//main/div[position() mod 2 = 1]//img[@class = 'photo']/@src"
11
12# Evaluate XPath and print the matching src values
13with ah.HTMLDocument(input_path) as document:
14    result = document.evaluate(
15        expression,
16        document,
17        None,
18        hxpath.XPathResultType.ANY,
19        None
20    )
21
22    attribute = result.iterate_next()
23    while attribute is not None:
24        print(attribute.node_value)
25        attribute = result.iterate_next()

For the supplied file, the expression returns 12 src attribute nodes. Some photos occur in more than one row, so repeated source values are expected. The example extracts their locations but does not download the image files.

Reuse XPath Patterns for Other Tasks

Adapt the same evaluation workflow by replacing only the expression:

Required resultXPath expression
Every image element//img
The first image in the document(//img)[1]
Elements with a particular ID//*[@id = 'content']
Images that define alternative text//img[@alt]
Paragraphs containing a phrase//p[contains(normalize-space(.), 'Release notes')]
Links inside the main content//main//a[@href]

Use // for a document-wide search. Use .// when the expression should search only descendants of the context node supplied to evaluate().

Common XPath Issues

IssueCause and recommended action
iterate_next() immediately returns NoneThe expression found no nodes. Test a broader expression before adding predicates.
An element with several classes is skipped@class = 'photo' compares the complete attribute value. Use the token-aware expression when photo can appear with other class names.
A relative query searches the wrong subtreeVerify the context node and begin the expression with .// when selection must stay within that node.
A text condition misses expected contentNested elements and whitespace affect string comparison. Use normalize-space(.) to compare normalized descendant text.
Code expects an element but XPath returns an attributeAn expression ending in /@name returns attribute nodes. Read node_value, or remove the attribute step when the element itself is required.

FAQ

When is XPath more useful than a CSS selector?

Choose XPath for text conditions, positional logic, ancestor-dependent queries, or direct attribute results. CSS selectors are usually shorter for matching by tag, class, ID, or ordinary attributes.

How do I select an element by attribute?

Put the condition in square brackets. For example, //img[@alt] selects images that have an alt attribute, and //input[@name = 'email'] selects inputs with a particular name value.

Can XPath search text in HTML?

Yes. For example, //p[contains(normalize-space(.), 'Release notes')] matches paragraphs whose normalized descendant text contains that phrase.

Does evaluate() change the selected nodes?

No. evaluate() only produces an XPath result. Retrieve the nodes and modify the DOM explicitly if the document must be changed.

Related Articles

Other Platforms