Select HTML Elements with CSS Selectors in Python

Load HTML with HTMLDocument, pass a CSS selector to query_selector() when you need the first match, or use query_selector_all() to process every match. The returned DOM nodes can then be read or modified in Python.

CSS selectors provide a compact way to locate HTML content before extracting text, changing the DOM, or saving an edited document. A selector can describe an element name, class, ID, attribute, or relationship between elements.

The examples below cover three Python workflows: selecting one element from an HTML string, processing several matching nodes, and querying an existing HTML file.

Get the First Matching Element

Use query_selector() when one matching element is enough. It returns an Element for the first match and None when the selector does not match the loaded document.

This example uses main > p to find the first paragraph whose direct parent is <main>:

  1. Load the HTML source into an HTMLDocument.
  2. Define a CSS selector that identifies the required element.
  3. Call query_selector() and check whether it returns None.
  4. Read or modify the matched Element.
 1# Select the first matching HTML element with a CSS selector in Python
 2
 3import aspose.html as ah
 4
 5# Define the HTML content
 6html_code = """
 7<main>
 8    <p>First paragraph</p>
 9    <section><p>Nested paragraph</p></section>
10    <p>Second paragraph</p>
11</main>
12"""
13
14# Select the first direct paragraph child of <main>
15with ah.HTMLDocument(html_code, ".") as document:
16    paragraph = document.query_selector("main > p")
17
18    if paragraph is not None:
19        print(paragraph.text_content)

The output is First paragraph. The nested paragraph is not a direct child of <main>, and query_selector() stops after the first match.

Process All Elements That Match a Selector

The query_selector_all() method returns a NodeList. Iterate over it when the same operation must be applied to every match.

The compound selector li.task[data-status='new'] matches <li> elements that have the task class and whose data-status attribute equals new. The following code marks the text of those tasks and saves the updated HTML:

  1. Load or create the HTML document.
  2. Define a selector that combines the required matching conditions.
  3. Call query_selector_all() and iterate through the returned NodeList.
  4. Process the matching nodes and save the document when the changes must persist.
 1# Select all matching HTML elements with CSS selectors in Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Prepare the output path
 7output_dir = "output"
 8os.makedirs(output_dir, exist_ok=True)
 9output_path = os.path.join(output_dir, "selected-tasks.html")
10
11# Define the HTML content
12html_code = """
13<ul>
14    <li class="task" data-status="new">Review the document</li>
15    <li class="task" data-status="done">Publish the document</li>
16    <li class="task" data-status="new">Archive the source files</li>
17</ul>
18"""
19
20# Select and update new tasks
21with ah.HTMLDocument(html_code, ".") as document:
22    tasks = document.query_selector_all("li.task[data-status='new']")
23
24    for task in tasks:
25        task.text_content = f"{task.text_content} (new)"
26
27    document.save(output_path)

Only the first and third items receive the (new) suffix. The item with data-status="done" remains unchanged.

In the Python binding, iterating a NodeList returned by query_selector_all() exposes its items as DOM Node objects. Common node properties such as text_content are available directly. By contrast, query_selector() returns a single Element, which also exposes element-specific methods such as get_attribute() and set_attribute().

Select Content from an HTML File

Selectors work the same way after loading a local file. The next example reuses paragraphs.html and selects only paragraphs that are direct children of its <article> element:

  1. Load the source HTML file with HTMLDocument.
  2. Define a selector based on stable elements, classes, IDs, attributes, or relationships in the source markup.
  3. Call query_selector() or query_selector_all() according to the number of required matches.
  4. Read or process the selected content.
 1# Select content from an HTML file with CSS selectors in Python
 2
 3import os
 4import aspose.html as ah
 5
 6# Prepare the input path
 7data_dir = "data"
 8input_path = os.path.join(data_dir, "paragraphs.html")
 9
10# Select and print paragraphs inside <article>
11with ah.HTMLDocument(input_path) as document:
12    paragraphs = document.query_selector_all("article > p")
13
14    for paragraph in paragraphs:
15        print(paragraph.text_content)

The selector returns the three paragraphs inside <article>. It does not return the article heading or elements outside that container.

Adapt a CSS Selector to Your HTML

Start with the simplest condition that uniquely identifies the required content, then add another condition only when necessary:

A space changes the meaning of a selector. .notice.active requires both classes on the same element, whereas .notice .active looks for an .active descendant inside .notice. The & character is not an AND operator for selectors passed to these DOM methods.

Common CSS Selector Issues

IssueCause and recommended action
query_selector() returns NoneThe selector does not match the loaded DOM. Check its spelling, attribute values, and element relationships before accessing the result.
query_selector_all() produces no outputThe returned NodeList is empty. Test a simpler selector first, then add classes, attributes, or combinators.
A child selector misses nested content> matches only direct children. Use a space when descendants at any depth should match.
The output file is unchangedSelecting nodes does not change them. Modify the returned nodes and call save() with the intended output path.
A saved style is not visibleAnother CSS declaration may have greater specificity or use !important. Inspect the existing inline, internal, and external styles.

FAQ

How do I select an element by class or ID in Python?

Pass .className or #elementId to query_selector() or query_selector_all(). Use the first method for one match and the second for all matches.

How do I select an element by its name attribute?

Use [name='email'] for any element with that value, or narrow it to a particular element type with input[name='email'].

Can a CSS selector match visible text?

Standard CSS selectors do not provide a general text-content condition. Select candidate nodes and compare their text_content values in Python, or use XPath when the expression itself must test text.

Do CSS selectors modify the HTML document?

No. They only locate nodes. Change the returned node or element explicitly, then call save() if the updated DOM must be written to a file.

Related Articles

Other Platforms