HTML Navigation in Python

Aspose.HTML for Python via .NET represents an HTML document as a Document Object Model (DOM) tree. After loading an HTMLDocument, Python code can move between related nodes, inspect elements, evaluate XPath expressions, or select elements with CSS selectors.

Use direct DOM properties such as first_child and next_sibling when the node position is known. Use evaluate() for XPath expressions and query_selector_all() for matching by tag, class, ID, attribute, or element relationship.

HTML DOM Navigation APIs

The aspose.html.dom module provides the Document, Node, and Element types used to inspect a document tree. These APIs follow the node relationships defined by the WHATWG DOM Standard.

Node properties can return element, text, or comment nodes. Element-specific properties skip non-element nodes.

APIResult
first_child and last_childThe first or last child node, including text and comment nodes
next_sibling and previous_siblingThe adjacent sibling node
child_nodesAll child nodes
first_element_child and last_element_childThe first or last child element
childrenChild elements without text or comment nodes
get_element_by_id()The element with the specified unique ID
get_elements_by_tag_name()Elements with the specified tag name

For DOM modification methods, attributes, inner_html, and text_content, see Edit HTML Documents in Python.

The first example creates two <span> elements separated by a whitespace text node. It demonstrates that sibling navigation visits the text node as well as the elements.

To navigate adjacent HTML nodes:

  1. Load or create an HTMLDocument.
  2. Choose a starting node through the document body or another parent node.
  3. Use child and sibling properties to move through adjacent DOM nodes.
  4. Inspect each returned node and read its content or other required properties.
 1# Navigate the HTML DOM in Python
 2
 3import aspose.html as ah
 4
 5# Define the HTML content
 6html_code = "<span>Hello,</span> <span>World!</span>"
 7
 8# Navigate between sibling nodes
 9with ah.HTMLDocument(html_code, ".") as document:
10    element = document.body.first_child
11    print(element.text_content)
12
13    element = element.next_sibling
14    print(element.text_content)
15
16    element = element.next_sibling
17    print(element.text_content)

Inspect HTML Elements

Use element traversal properties when whitespace and other non-element nodes should be skipped. The following example loads html_file.html and follows the element structure defined by the Element Traversal specification.

To inspect HTML elements:

  1. Load the source HTML file into an HTMLDocument.
  2. Access the root element through document_element.
  3. Use element traversal properties to move between parent and child elements.
  4. Inspect the tag names, text, or other properties of the selected elements.
 1# Navigate and inspect an HTML document in Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Prepare the input path
 7data_dir = "data"
 8document_path = os.path.join(data_dir, "html_file.html")
 9
10# Navigate through the HTML element tree
11with ah.HTMLDocument(document_path) as document:
12    element = document.document_element
13    print(element.tag_name)
14
15    element = element.last_element_child
16    print(element.tag_name)
17
18    element = element.first_element_child
19    print(element.tag_name)
20    print(element.text_content)

Select HTML Nodes with XPath

XPath selects nodes by document structure, attributes, and relationships. Aspose.HTML exposes the DOM XPath workflow through document.evaluate(), which returns an XPathResult of the requested type.

The example evaluates //*[@class='happy']//span, which selects descendant <span> elements under elements whose class attribute is exactly happy.

To select HTML nodes with XPath:

  1. Load or create the HTML document that contains the target nodes.
  2. Define an XPath expression for the required structure or conditions.
  3. Pass the expression and an appropriate context node to evaluate().
  4. Request XPathResultType.ANY.
  5. Call iterate_next() until every matching node has been processed.
  6. Read the text_content of each selected node.
 1# Select HTML nodes with XPath in Python
 2
 3import aspose.html as ah
 4import aspose.html.dom.xpath as hxpath
 5
 6# Define the HTML content
 7html_code = """
 8    <div class='happy'>
 9        <div>
10            <span>Hello,</span>
11        </div>
12    </div>
13    <p class='happy'>
14        <span>World!</span>
15    </p>
16"""
17
18# Select <span> nodes inside elements with the happy class
19with ah.HTMLDocument(html_code, ".") as document:
20    result = document.evaluate(
21        "//*[@class='happy']//span",
22        document,
23        None,
24        hxpath.XPathResultType.ANY,
25        None
26    )
27
28    node = result.iterate_next()
29    while node is not None:
30        print(node.text_content)
31        node = result.iterate_next()

Select HTML Elements with CSS Selectors

CSS selectors provide concise matching by tag name, class, ID, attribute, and relationship. Use query_selector() to return the first matching element or query_selector_all() to return a NodeList containing every match.

The example uses #catalog article.product[data-status='available'] h2 to select <h2> elements from available products in the catalog. The expression combines an ID selector, an element and class selector, an attribute selector, and descendant relationships.

To select HTML elements with a CSS selector:

  1. Load or create the HTML document that contains the target elements.
  2. Define a CSS selector for the required tags, classes, attributes, or relationships.
  3. Pass the selector to query_selector_all().
  4. Iterate through the returned NodeList and process each matching node.
 1# Select HTML elements using a CSS selector in Python
 2
 3import aspose.html as ah
 4
 5# Define the HTML content
 6html_code = """
 7<section id="catalog">
 8    <article class="product" data-status="available">
 9        <h2>HTML Editor</h2>
10    </article>
11    <article class="product" data-status="unavailable">
12        <h2>PDF Converter</h2>
13    </article>
14    <article class="product" data-status="available">
15        <h2>Image Converter</h2>
16    </article>
17</section>
18"""
19
20# Select headings for available products
21selector = "#catalog article.product[data-status='available'] h2"
22
23with ah.HTMLDocument(html_code, ".") as document:
24    elements = document.query_selector_all(selector)
25
26    for heading in elements:
27        print(heading.text_content)

The selector excludes the unavailable product and produces the following output:

1HTML Editor
2Image Converter

Choose an HTML Navigation Method

MethodBest suited for
Direct node navigationMoving between known parents, children, or siblings while preserving text and comment nodes
Element traversalMoving between elements while skipping whitespace, text, and comment nodes
XPathSelecting nodes with structural paths, attribute tests, and relationship conditions
CSS selectorsMatching elements with familiar tag, class, ID, attribute, and combinator syntax

Common HTML Navigation Issues

IssueCause and recommended action
first_child returns a text nodeWhitespace between tags is represented by a DOM text node. Use first_element_child when only an element is required.
XPath returns no nodesThe expression does not match the loaded DOM. Inspect the available elements and attributes, then test a simpler expression.
query_selector_all() returns an empty NodeListThe selector does not match the document or uses the wrong class, ID, attribute, or relationship. Verify the loaded HTML and selector syntax.
Element-specific code receives another node typeDirect node navigation can return text or comment nodes. Check the node type or use element traversal properties.

FAQ

What is the difference between Node and Element navigation?

Node navigation includes element, text, and comment nodes. Element traversal properties return only elements, which is useful when formatting whitespace should not affect the traversal logic.

Should I use XPath or CSS selectors?

Use CSS selectors for familiar matching by tag, class, ID, or attribute. Use XPath when selection depends on document hierarchy, node relationships, or path expressions.

Does query_selector_all() return one element or multiple elements?

query_selector_all() returns a NodeList containing all matching elements. Use query_selector() when only the first matching element is required.

Other Platforms

Related Articles