Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.
Load HTML with HTMLDocument, pass a CSS selector to query_selector() when you need the first match, or use query_selector_all() to process every match. The returned DOM nodes can then be read or modified in Python.
CSS selectors provide a compact way to locate HTML content before extracting text, changing the DOM, or saving an edited document. A selector can describe an element name, class, ID, attribute, or relationship between elements.
The examples below cover three Python workflows: selecting one element from an HTML string, processing several matching nodes, and querying an existing HTML file.
Use query_selector() when one matching element is enough. It returns an Element for the first match and None when the selector does not match the loaded document.
This example uses main > p to find the first paragraph whose direct parent is <main>:
HTMLDocument.query_selector() and check whether it returns None.Element. 1# Select the first matching HTML element with a CSS selector in Python
2
3import aspose.html as ah
4
5# Define the HTML content
6html_code = """
7<main>
8 <p>First paragraph</p>
9 <section><p>Nested paragraph</p></section>
10 <p>Second paragraph</p>
11</main>
12"""
13
14# Select the first direct paragraph child of <main>
15with ah.HTMLDocument(html_code, ".") as document:
16 paragraph = document.query_selector("main > p")
17
18 if paragraph is not None:
19 print(paragraph.text_content)The output is First paragraph. The nested paragraph is not a direct child of <main>, and query_selector() stops after the first match.
The query_selector_all() method returns a NodeList. Iterate over it when the same operation must be applied to every match.
The compound selector li.task[data-status='new'] matches <li> elements that have the task class and whose data-status attribute equals new. The following code marks the text of those tasks and saves the updated HTML:
query_selector_all() and iterate through the returned NodeList. 1# Select all matching HTML elements with CSS selectors in Python
2
3import os
4import aspose.html as ah
5
6# Prepare the output path
7output_dir = "output"
8os.makedirs(output_dir, exist_ok=True)
9output_path = os.path.join(output_dir, "selected-tasks.html")
10
11# Define the HTML content
12html_code = """
13<ul>
14 <li class="task" data-status="new">Review the document</li>
15 <li class="task" data-status="done">Publish the document</li>
16 <li class="task" data-status="new">Archive the source files</li>
17</ul>
18"""
19
20# Select and update new tasks
21with ah.HTMLDocument(html_code, ".") as document:
22 tasks = document.query_selector_all("li.task[data-status='new']")
23
24 for task in tasks:
25 task.text_content = f"{task.text_content} (new)"
26
27 document.save(output_path)Only the first and third items receive the (new) suffix. The item with data-status="done" remains unchanged.
In the Python binding, iterating a NodeList returned by query_selector_all() exposes its items as DOM Node objects. Common node properties such as text_content are available directly. By contrast, query_selector() returns a single Element, which also exposes element-specific methods such as get_attribute() and set_attribute().
Selectors work the same way after loading a local file. The next example reuses
paragraphs.html and selects only paragraphs that are direct children of its <article> element:
HTMLDocument.query_selector() or query_selector_all() according to the number of required matches. 1# Select content from an HTML file with CSS selectors in Python
2
3import os
4import aspose.html as ah
5
6# Prepare the input path
7data_dir = "data"
8input_path = os.path.join(data_dir, "paragraphs.html")
9
10# Select and print paragraphs inside <article>
11with ah.HTMLDocument(input_path) as document:
12 paragraphs = document.query_selector_all("article > p")
13
14 for paragraph in paragraphs:
15 print(paragraph.text_content)The selector returns the three paragraphs inside <article>. It does not return the article heading or elements outside that container.
Start with the simplest condition that uniquely identifies the required content, then add another condition only when necessary:
p selects paragraphs by tag name..notice selects elements containing the notice class.#summary selects the element with id="summary".input[name='email'] combines a tag and attribute value.article > p selects direct children, while article p includes paragraphs at any descendant level.li.task[data-status='new'] combines a tag, class, and attribute as an AND condition.h1, h2, h3 combines alternative selectors as an OR condition.A space changes the meaning of a selector. .notice.active requires both classes on the same element, whereas .notice .active looks for an .active descendant inside .notice. The & character is not an AND operator for selectors passed to these DOM methods.
| Issue | Cause and recommended action |
|---|---|
query_selector() returns None | The selector does not match the loaded DOM. Check its spelling, attribute values, and element relationships before accessing the result. |
query_selector_all() produces no output | The returned NodeList is empty. Test a simpler selector first, then add classes, attributes, or combinators. |
| A child selector misses nested content | > matches only direct children. Use a space when descendants at any depth should match. |
| The output file is unchanged | Selecting nodes does not change them. Modify the returned nodes and call save() with the intended output path. |
| A saved style is not visible | Another CSS declaration may have greater specificity or use !important. Inspect the existing inline, internal, and external styles. |
Pass .className or #elementId to query_selector() or query_selector_all(). Use the first method for one match and the second for all matches.
Use [name='email'] for any element with that value, or narrow it to a particular element type with input[name='email'].
Standard CSS selectors do not provide a general text-content condition. Select candidate nodes and compare their text_content values in Python, or use XPath when the expression itself must test text.
No. They only locate nodes. Change the returned node or element explicitly, then call save() if the updated DOM must be written to a file.
querySelector() and querySelectorAll() with Aspose.HTML for Java.Analyzing your prompt, please hold on...
An error occurred while retrieving the results. Please refresh the page and try again.