NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

How to Find Elements With Selenium in Python

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Find Elements With Selenium in Python

Updated 2026-09-28

Replaced the earlier locator examples with a runnable local catalog and captured Chrome/Firefox results. Corrected singular/plural lookup, class-token matching, stale-reference recovery and shadow-root access; removed unsupported locator speed rankings.

Use driver.find_element(By.CSS_SELECTOR, selector) for the first matching element and driver.find_elements(By.CSS_SELECTOR, selector) for a list of all matches. If nothing matches, the singular method raises NoSuchElementException; the plural method returns an empty list.

For scraping, finding elements is only the first step. You also need the right container, populated fields and a way to recover when the page replaces nodes. The examples below extract a complete catalog and deliberately test selectors that return plausible but wrong results.

Find one element or all matching elements​

With a driver already open on the example catalog, scope the search and choose the required cardinality:

from selenium.webdriver.common.by import By

catalog = driver.find_element(By.ID, 'catalog')
first = catalog.find_element(By.CSS_SELECTOR, '.product.active')
rows = catalog.find_elements(By.CSS_SELECTOR, '.product.active')

In the captured fixture, first is the C-101 product and rows contains all three intended products. A successful singular lookup does not prove that the page contains only one match. See Selenium's element-finding documentation for the singular, plural and scoped search APIs.

The By strategy is the first argument. This reference table shows the locator forms; the expressions are examples to adapt to your page, not additional measured cases.

StrategyExample valueWhat it describes
By.IDcatalogAn element's ID
By.NAMEemailA name attribute
By.CSS_SELECTOR#catalog .product.activeA CSS selector
By.XPATH.//articleAn XPath expression relative to an element
By.CLASS_NAMEproductOne class name
By.TAG_NAMEinputAn element tag
By.LINK_TEXTView productLink text
By.PARTIAL_LINK_TEXTViewA portion of link text

These are the documented By strategies. Choose a locator that identifies the intended elements; this article does not benchmark one strategy against another.

Run a complete local example​

Download the example repository at the tested commit and open a terminal in its examples/selenium-python-find-element directory. Use Python 3.12, Bash and installed Chrome:

git clone https://github.com/ScrapingAnt/scrapingant-examples.git
cd scrapingant-examples
git checkout 74638f75381bdf2ff0652f1f07bd98f65186b40d
cd examples/selenium-python-find-element
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements-lock.txt
./run.sh

The command runs the tests, quickstart and Chrome cases. To include installed Firefox, run ./run.sh --all-browsers. Selenium Manager can obtain a compatible driver if you have not supplied one; initial dependency or driver setup may need Internet access. The README documents existing-browser and driver overrides.

The fixture itself is served on loopback and needs no API key. It contains three intended products, an outside decoy, an inactive product, optional badges, a same-origin frame and an open shadow root. Its delayed mode begins with present but empty product elements.

This is the complete quickstart.py entry point. The fixture server, browser setup, extractor and literal expected-record validator are included in the linked directory:

"""Complete no-key example: serve, navigate, wait, validate, print and clean up."""
import json
from selenium.webdriver.common.by import By
from browser_support import browser, fixture_server
from extraction import wait_records


def main():
with fixture_server() as origin, browser() as driver:
driver.get(origin + '/catalog.html?delayed=1')
driver.find_element(By.ID, 'begin').click()
print(json.dumps(wait_records(driver), ensure_ascii=False, indent=2))


if __name__ == '__main__':
main()

The context managers quit the driver and stop the local server when the block exits. wait_records returns only after the complete expected catalog is available. Here is the actual captured output:

[
{
"sku": "C-101",
"title": "Café & Cocoa",
"currency": "USD",
"price": "12.50",
"href": "/products/cafe",
"badge": "New"
},
{
"sku": "T-202",
"title": "Tea \"No. 2\"",
"currency": "EUR",
"price": "8.00",
"href": "/products/tea",
"badge": null
},
{
"sku": "N-303",
"title": "Notebook <A5>",
"currency": "GBP",
"price": "5.25",
"href": "/products/notebook",
"badge": "Sale"
}
]

New runs write to run_output/, leaving the committed captures intact. If the browser cannot start, check its installation and driver configuration. If binding a local port is prohibited in your environment, allow the loopback fixture server. A wait timeout means the expected record contract was not satisfied; inspect the selected rows and fields before extending the timeout.

Scope selectors before extracting fields​

The unscoped selector .product.active also matched the outside decoy. Searching inside catalog excluded it. The inactive row was excluded by the separate .active class selector.

The following project function is copied from the tested extractor. Each title, price and optional badge is queried relative to its own product row:

def project(rows):
result = []
for row in rows:
title = row.find_element(By.CSS_SELECTOR, '.title')
badges = row.find_elements(By.CSS_SELECTOR, '.badge')
result.append({
'sku': row.get_dom_attribute('data-sku'),
'title': title.text,
'currency': row.find_element(By.CSS_SELECTOR, '.currency').text,
'price': row.find_element(By.CSS_SELECTOR, '.price').text,
'href': title.get_dom_attribute('href'),
'badge': badges[0].text if badges else None,
})
return result

Call project(rows) after the scoped plural lookup above. It reads visible text for the title, currency and price, preserves the href DOM attribute, and returns None when the optional badge is absent. It does not silently convert missing required fields into a complete record.

For XPath, the tested scoped lookup uses .// and a class-token expression:

rows = catalog.find_elements(
By.XPATH,
".//article[contains(concat(' ', normalize-space(@class), ' '), ' active ')]",
)

It returned the same three intended records as scoped CSS. Keep the relative expression when your starting point is a container element.

Match class tokens, not substrings​

Use .product.active to require both classes. Passing product active to By.CLASS_NAME produced InvalidSelectorException in both captured browsers.

If you need to inspect a class string in Python, the tested helper splits it into tokens:

def has_class_token(classes, token):
return token in classes.split()

The negative control 'active' in row.get_dom_attribute('class') included the fixture's inactive row. Splitting into tokens excluded it. That distinction matters even when every returned record has plausible-looking text and prices.

Wait for populated records and recover from stale elements​

presence_of_all_elements_located returned three product shells before their fields were filled. The elements existed, but their title, currency, price and link values were empty. Waiting for presence alone therefore failed this fixture's extraction task.

The tested wait projects fresh matches on every poll and checks the resulting data:

from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException
from selenium.webdriver.support.ui import WebDriverWait
from oracle import exact_records

CATALOG = (By.CSS_SELECTOR, '#catalog .product.active')

def wait_records(driver, locator=CATALOG, timeout=5):
def complete(current):
try:
records = project(current.find_elements(*locator))
return records if exact_records(records) else False
except (NoSuchElementException, StaleElementReferenceException):
return False
return WebDriverWait(driver, timeout, poll_frequency=0.05).until(complete)

Here, exact_records checks the independent expected tuples, required keys, types and SKU uniqueness. It rejects missing, extra, duplicate, malformed or wrong-valued records. The literal three-record expectation belongs to this controlled fixture. For a real scraper, define a completeness contract appropriate to the page instead of copying that record count.

A WebElement is a reference to a particular node. In the replacement case, the fixture replaced its catalog children; reading the old handles raised StaleElementReferenceException. Calling wait_records(driver) again found the replacement nodes and returned the complete dataset. Selenium documents the need to re-locate stale elements.

The wait also handles replacement between the lookup and the field reads. A separate browser regression forced that race and recovered on the second locator search. Removing the stale-reference handler made that regression fail. Retry the locator and projection, rather than repeatedly reading an old handle.

Read visible text, DOM properties and HTML attributes deliberately​

These APIs answered different questions in the fixture:

ReadCaptured result
element.text on the visibility probeVisible
element.get_property('textContent') on that probeVisibleHidden
field.get_dom_attribute('value') after its value was editedinitial
field.get_property('value')edited
field.get_attribute('value')edited
link.get_dom_attribute('href')/products/cafe
link.get_property('href')An absolute URL on the loopback fixture origin

The probe's hidden span explains the text difference. The input's property changed while its HTML attribute retained the initial value. Selenium's get_attribute method tries a property first; use get_dom_attribute or get_property when that distinction is part of your output contract.

Search inside frames and open shadow roots​

For the fixture's frame, switch into its document, extract, and restore the parent context in finally. This excerpt uses the same wait_records helper:

from selenium.webdriver.support import expected_conditions as EC

WebDriverWait(driver, 5).until(
EC.frame_to_be_available_and_switch_to_it((By.ID, 'catalog-frame'))
)
try:
records = wait_records(driver)
finally:
driver.switch_to.default_content()

For the open shadow root, use Selenium's native root object:

root = driver.find_element(By.ID, 'shadow-host').shadow_root
records = project(root.find_elements(By.CSS_SELECTOR, '.product.active'))

Both approaches returned the fixture's three target records. The frame test was same-origin and the shadow root was open; these observations do not establish arbitrary cross-origin-frame or closed-shadow-root access. See Selenium's frame documentation and shadow-root search example.

What to check when extraction fails​

Symptom in the controlled casesLikely mistake shown by the caseNext action
Only C-101 is returnedSingular lookup used for a collectionUse scoped find_elements.
Outside or inactive decoy appearsScope or class matching is too broadQuery within the intended container and match class tokens.
Three rows exist but required values are emptyPresence mistaken for populated dataWait on the projected record contract.
Old handles raise after replacementNode references were retained across a DOM updateRe-find and re-project within a bounded wait.
InvalidSelectorException for product activeMultiple classes passed to By.CLASS_NAMEUse .product.active as CSS.
A missing singular lookup raisesNo element matched in that search contextCheck the locator and context; use the plural result if absence is expected.
Text or URLs differ from the intended outputAttribute/property or visible/raw-text meanings were mixedChoose the API that matches the field contract.

The saved matrix ran 12 extraction cases and five separate diagnostics, with three rounds in each of Chrome and Firefox. Of 72 extraction observations, 42 returned the complete catalog; the other 30 matched deliberately wrong results or the expected stale exception. The 30 diagnostics also matched their specified outcomes. These are controlled correctness checks, not a locator speed comparison or a production success rate.

When ScrapingAnt fits this task​

Keep the local Selenium workflow when it already supplies the browser interaction and data you need. This fixture needs no remote service. For already available static HTML, the adjacent task is selecting from parsed HTML.

ScrapingAnt can execute a custom JavaScript snippet in the loaded page. The snippet can extract DOM data and write JSON into an element that the client reads from the returned HTML. This moves the DOM query and projection into the remote request; it does not accept Selenium By locators or existing WebElement handles.

The separate tested API example extracts product, price, units and share from a public price-table fixture. Its full snippet selects table rows and validates their cells before creating a result envelope. This carrier-writing excerpt is copied from that snippet; result is the envelope and config.marker is a unique request marker supplied by the runner:

const carrier = document.createElement("script");
carrier.id = config.marker;
carrier.type = "application/json";
// Script data is raw text: escape less-than, not HTML entities.
carrier.textContent = JSON.stringify(result).replace(/</g, "\\u003c");
document.body.append(carrier);

Escaping < as a JSON Unicode escape keeps a value such as </script> from ending the raw-text carrier. The client must not HTML-entity-decode that JSON: a literal &amp; value is data. The recorded transport probe preserved those strings, quotes, newlines and Unicode through the returned HTML.

The Python client uses parse_html(html, marker) to check the carrier and decode its envelope, then validates the application fields against the expected records. The packet also has an exact-marker regex reader with a structural HTML check; its narrow format is not a general-purpose HTML extraction regex. Validate the result marker, JSON envelope and required data fields; an HTTP 200 response alone does not prove the extraction succeeded.

That distinction appeared in the live controls: returning a JavaScript object without writing a carrier yielded no marker, and an absent selector yielded an ok: false envelope. Both API responses were HTTP 200. Across six calls, the example captured three complete price-table extractions, one delayed-data result and those two controls. The six receipts totaled 60 credits. A request with JavaScript rendering through a datacenter proxy costs 10 API credits.

The complete request runner uses browser=true, return_page_source=false, standard Base64 for UTF-8 JavaScript, and query-parameter URL encoding. Its paid run requires explicit --live; the packet's default reproduction is local and free. If using wait_for_selector, target an element the page creates, not the result marker your snippet will create. Start with the JavaScript execution documentation.

Scope of these examples​

The local catalog and API price table are self-authored fixtures. They do not measure anti-bot behavior, pagination completeness, production uptime, closed shadow roots or arbitrary browser-state transfer. The Selenium capture used Python 3.12.10, Selenium 4.49.0, Chrome 154.0.8037.57 and Firefox 156.0 on Darwin 25.6.0 arm64. Keep the raw observations and validation alongside your scraper so that a selector returning data is not confused with a selector returning the right data.

For related workflows, see CSS selectors in BeautifulSoup and the Playwright web-scraping guide.

Examples tested on 2026-09-28 with Python 3.12.10 and Selenium 4.49.0. Code: Selenium example and captured outputs; separate DOM extraction API example.

This article was drafted with AI assistance. The linked evidence contains the executed code, captured outputs and known limitations.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster