Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
Load HTML with
HTMLDocument, pass an XPath expression to
evaluate(), request XPathResultType.ANY, and retrieve matching nodes by repeatedly calling iterate_next().
XPath is useful when a selection depends on document structure, position, attributes, or text. Unlike a CSS selector, an XPath expression can also return attribute nodes or values rather than their owning elements.
The examples below first evaluate an expression against HTML created from a string. They then use XPath to isolate photo images in an existing file and return their src attributes directly.
The call evaluate(expression, context_node, resolver, type, result) evaluates an XPath expression relative to a context node. For an ordinary document-wide HTML query, pass the document as the context node, None as the namespace resolver, XPathResultType.ANY as the requested type, and None when there is no existing result object to reuse.
The following example selects only paragraphs whose data-status value is published:
HTMLDocument.evaluate() and request an appropriate result type. 1# Select HTML elements with XPath in Python
2
3import aspose.html as ah
4import aspose.html.dom.xpath as hxpath
5
6# Define the HTML content
7html_code = """
8<section>
9 <p data-status="published">Getting Started</p>
10 <p data-status="draft">Draft tutorial</p>
11 <p data-status="published">API Reference</p>
12</section>
13"""
14
15# Evaluate XPath and print the matching paragraphs
16with ah.HTMLDocument(html_code, ".") as document:
17 result = document.evaluate(
18 "//p[@data-status = 'published']",
19 document,
20 None,
21 hxpath.XPathResultType.ANY,
22 None
23 )
24
25 node = result.iterate_next()
26 while node is not None:
27 print(node.text_content)
28 node = result.iterate_next()The predicate inside square brackets removes the draft paragraph from the result. The output is:
1Getting Started
2API ReferenceThe sample
xpath-image.htm deliberately mixes photo images with advertising images. The ads appear in the header, footer, separate rows inside <main>, and individual photo rows. This makes the file useful for showing how each part of an XPath expression changes the result.
For this sample, the expression is refined in four stages:
//img selects every image in the document.//main//img excludes images from the header and footer.//main/div[position() mod 2 = 1]//img keeps images from the odd-positioned photo rows and excludes separate advertising rows.[@class = 'photo'] excludes inline banners and retains only photo images.The final element query is:
1//main/div[position() mod 2 = 1]//img[@class = 'photo']This predicate requires the complete class attribute to equal photo. If photo may be one of several classes, use a token-aware test:
1//main/div[position() mod 2 = 1]//img[contains(concat(' ', normalize-space(@class), ' '), ' photo ')]XPath can select the src attributes instead of returning image elements. Append /@src to the element query, iterate through the resulting attribute nodes, and read each value through node_value. This avoids an element cast in Python.
HTMLDocument.node_value from each attribute node. 1# Extract image src attributes with XPath in Python
2
3import os
4import aspose.html as ah
5import aspose.html.dom.xpath as hxpath
6
7# Prepare the input path and XPath expression
8data_dir = "data"
9input_path = os.path.join(data_dir, "xpath-image.htm")
10expression = "//main/div[position() mod 2 = 1]//img[@class = 'photo']/@src"
11
12# Evaluate XPath and print the matching src values
13with ah.HTMLDocument(input_path) as document:
14 result = document.evaluate(
15 expression,
16 document,
17 None,
18 hxpath.XPathResultType.ANY,
19 None
20 )
21
22 attribute = result.iterate_next()
23 while attribute is not None:
24 print(attribute.node_value)
25 attribute = result.iterate_next()For the supplied file, the expression returns 12 src attribute nodes. Some photos occur in more than one row, so repeated source values are expected. The example extracts their locations but does not download the image files.
Adapt the same evaluation workflow by replacing only the expression:
| Required result | XPath expression |
|---|---|
| Every image element | //img |
| The first image in the document | (//img)[1] |
| Elements with a particular ID | //*[@id = 'content'] |
| Images that define alternative text | //img[@alt] |
| Paragraphs containing a phrase | //p[contains(normalize-space(.), 'Release notes')] |
| Links inside the main content | //main//a[@href] |
Use // for a document-wide search. Use .// when the expression should search only descendants of the context node supplied to evaluate().
| Issue | Cause and recommended action |
|---|---|
iterate_next() immediately returns None | The expression found no nodes. Test a broader expression before adding predicates. |
| An element with several classes is skipped | @class = 'photo' compares the complete attribute value. Use the token-aware expression when photo can appear with other class names. |
| A relative query searches the wrong subtree | Verify the context node and begin the expression with .// when selection must stay within that node. |
| A text condition misses expected content | Nested elements and whitespace affect string comparison. Use normalize-space(.) to compare normalized descendant text. |
| Code expects an element but XPath returns an attribute | An expression ending in /@name returns attribute nodes. Read node_value, or remove the attribute step when the element itself is required. |
Choose XPath for text conditions, positional logic, ancestor-dependent queries, or direct attribute results. CSS selectors are usually shorter for matching by tag, class, ID, or ordinary attributes.
Put the condition in square brackets. For example, //img[@alt] selects images that have an alt attribute, and //input[@name = 'email'] selects inputs with a particular name value.
Yes. For example, //p[contains(normalize-space(.), 'Release notes')] matches paragraphs whose normalized descendant text contains that phrase.
No. evaluate() only produces an XPath result. Retrieve the nodes and modify the DOM explicitly if the document must be changed.
query_selector() and query_selector_all().Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.