NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

Puppeteer Find Elements: Select, Wait and Extract Records

· 14 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Puppeteer Find Elements: Select, Wait and Extract Records

Use page.$() to get the first matching element and page.$$() to get all matching elements. When you want data rather than handles, page.$eval() runs a callback on the first match, and page.$$eval() passes all matches to a callback that can return JSON records.

The selector still needs to identify the right records, and the fields must be populated before you read them. This guide demonstrates both problems with a local catalogue: a promotional card that matches an overly broad selector, and product cards that exist before their titles and prices arrive.

Find one element, all elements, or their data​

These are the outcomes recorded with Puppeteer 25.12.0:

MethodWhat you receiveIf the selector matches nothing
page.$(selector)First matching ElementHandlenull
page.$$(selector)Array of matching ElementHandlesEmpty array
page.$eval(selector, callback)Callback result for the first matchError naming the missing selector
page.$$eval(selector, callback)Callback result for the matched arrayCallback receives an empty array; our mapping returned []

These query methods do not establish that your application's data is ready. Puppeteer documents them as queries without waiting. For clicking and other interactions, the same documentation recommends the Locator API. For this extraction task, we explicitly wait for the required fields and then project their values.

If you only need the page's readable text, use the whole-page text guide. Flattening a page loses the product-to-price relationship that this example preserves.

Reproduce a selector that finds the wrong card​

Download the example repository at this exact snapshot, then open its examples/puppeteer-find-elements directory. The packet contains the HTTP fixture, independent expected records, complete scripts and a dependency lock. Use Node 22.12 or newer; the recorded run used Node 22.20.0.

git clone https://github.com/ScrapingAnt/scrapingant-examples.git
cd scrapingant-examples
git checkout 74638f75381bdf2ff0652f1f07bd98f65186b40d
cd examples/puppeteer-find-elements
npm ci
node quickstart.mjs

The fixture has three product cards inside #catalog and a promotional .card before that container. Before building the successful example, compare the global_first and global_all cases in the raw capture.

The first global .card matched this record:

[
{
"sku": "DECOY",
"title": "Promotional decoy",
"currency": "USD",
"price": "0.01",
"href": "/promotion"
}
]

Changing $ to $$ collected that decoy plus all three products. Changing the scope to #catalog > .card collected only the catalogue's products. A larger result set does not correct a selector that includes unrelated content.

Complete example: wait, project, validate​

This is quickstart.mjs in full. Run it from the downloaded example directory so its fixture, launch and validation helpers are available:

// Run after npm ci. Starts only this packet's self-authored local HTTP fixture.
import assert from 'node:assert/strict';
import {writeFile,mkdir} from 'node:fs/promises';
import {dirname} from 'node:path';
import {startFixture} from './fixture.mjs';
import {launchBrowser} from './browser.mjs';
import {validateRecords} from './oracle.mjs';
const fixture=await startFixture();let browser;
try{
browser=await launchBrowser();const page=await browser.newPage();
const response=await page.goto(fixture.origin+'/catalog',{waitUntil:'load'});assert.equal(response.status(),200);
await page.evaluate(()=>window.releaseCatalog()); // Fixture handshake; a real site loads its own data.
await page.waitForFunction(()=>{
const rows=[...document.querySelectorAll('#catalog > .card')];
return rows.length===3&&rows.every(row=>row.querySelector('.title')?.textContent.trim()&&/^\d+\.\d{2}$/.test(row.querySelector('.price')?.textContent.trim()));
},{timeout:5000});
const records=await page.$$eval('#catalog > .card',cards=>cards.map(card=>({
sku:card.dataset.sku,title:card.querySelector('.title').textContent.trim(),currency:card.querySelector('.price').dataset.currency,
price:card.querySelector('.price').textContent.trim(),href:card.querySelector('.title').getAttribute('href'),
})));
validateRecords(records);const result={records};console.log(JSON.stringify(result,null,2));
const output=process.argv[2]??'run_output/quickstart.json';await mkdir(dirname(output),{recursive:true});await writeFile(output,JSON.stringify(result,null,2)+'\n');
}finally{await browser?.close();await fixture.close();}

releaseCatalog() is a fixture-only handshake that makes the initially empty cards populate on a later animation frame. A real website loads its own data; replace that test handshake with the application's normal trigger. The important part to retain is a readiness condition for the fields you intend to extract.

Here, the condition requires exactly three cards, a nonempty title in each, and a price matching the fixture's decimal format. The three-card count and price format are specific to this catalogue. Adapt them to the target's actual contract instead of copying the count into a scraper for an unknown-length list.

The browser callback reads each title, currency, price and href from its own card. It returns plain objects, preserving the relationship between those fields. The independent validateRecords() helper checks the result against literal expected tuples; it rejects missing, duplicate, extra, reordered or changed records.

Captured output:

{
"records": [
{
"sku": "A101",
"title": "Ant Field Kit",
"currency": "USD",
"price": "19.95",
"href": "/p/a101"
},
{
"sku": "B202",
"title": "Sugar & Water Feeder",
"currency": "USD",
"price": "7.50",
"href": "/p/b202"
},
{
"sku": "C303",
"title": "Tunnel Kit — Mini",
"currency": "USD",
"price": "12.00",
"href": "/p/c303"
}
]
}

Prices remain strings in this example. There is no locale-sensitive money parser or currency conversion hidden in the projection. Required title/price/link nodes are treated as required; if one is missing, extraction fails instead of silently manufacturing a value.

Presence is different from populated data​

The matrix first called waitForSelector('#catalog .card') while the fixture's population gate was still closed. The card existed, so that wait completed. Its first captured record was:

{
"sku": "A101",
"title": "",
"currency": "USD",
"price": "",
"href": "/p/a101"
}

Waiting for a container, navigating successfully, or counting the expected cards does not prove their fields contain data. The populated-field condition in the quickstart returned all three expected records after the fixture acknowledged population. No arbitrary sleep was needed.

For a real page, choose a condition that distinguishes loading placeholders from finished records. A nonempty string alone can still be a loading message. Validate the extracted schema and values after the wait as well.

CSS, text and XPath selectors​

These selectors were exercised against the same populated fixture:

SelectorCaptured match
#catalog > .cardAll three catalogue cards
#catalog ::-p-text(Sugar & Water Feeder)The matching title; its closest card was B202
::-p-xpath(//section[@id="catalog"]/article)All three catalogue cards
#shadow-host >>> .cardThe SHADOW record inside an open shadow root

The ::-p-text(...), ::-p-xpath(...) and shadow traversal forms are Puppeteer's documented selector syntax. They are not selectors you can pass unchanged to browser document.querySelector().

For repeated records, a parent handle is another way to establish scope: the tested sequence used page.$('#catalog'), then catalog.$$('.card'), read each handle, and disposed of the handles. The scoped handle method and scoped $$eval projection returned the same three records in this fixture. Choose the form that makes the ownership of each field clear; no speed comparison was made.

The B202 link deliberately has no data-stock attribute. getAttribute('data-stock') returned null. Its raw href was /p/b202, while its DOM href property was an absolute URL with the same path. Decide whether your output contract wants the source attribute or the resolved URL; the example oracle expects relative hrefs.

For optional attributes, preserve absence deliberately. For required fields, fail validation rather than substituting a value that makes a broken selector look successful.

A more subtle failure appeared after replacing the first card. The existing handle still read Ant Field Kit, even though its isConnected value was false. A fresh selector read Ant Field Kit — revised. This was same-document node replacement: it demonstrates that reading from an old handle can return stale data, without proving that every operation on a detached handle succeeds or that handles survive navigation.

Frames and shadow roots have their own query boundaries​

An ordinary main-document query for #catalog-frame .card returned no matches. The tested frame path located the iframe, called its contentFrame(), and ran frame.$$eval('.card', projectCards) inside that document. It returned the FRAME record.

Similarly, document.querySelectorAll('#shadow-host .card') returned no matches, while Puppeteer's explicit #shadow-host >>> .card query reached the open shadow root. These checks used one same-origin iframe and one open root; they do not establish cross-origin access or closed-root behavior.

What to check when the output is wrong​

SymptomCheck demonstrated by the fixture
Only one record appears$ and $eval select the first match; use an all-match operation when the task requires all records
A promotion appears among productsNarrow the selector to the intended catalogue container
The right number of rows has empty fieldsWait for populated fields, then validate their values
A missing selector looks like an empty successful scrapeCheck the API's missing-match result and enforce the expected record contract
A title stays unchanged after the page updatesRe-query the current document instead of reusing a detached handle
A visible item is absent from main-document queriesEstablish whether it belongs to a frame or an open shadow root

The full matrix ran 13 extraction cases three times: 39 extraction observations, plus nine API/boundary diagnostics and 15 text diagnostics. All 63 checks matched their declared outcomes, including deliberately wrong selections and incomplete data. They are regression checks, not a production scraping success rate. The separate boundary suite passed 31 tests.

When to use ScrapingAnt for the DOM extraction​

You do not need ScrapingAnt to run this fixture, parse HTML you already have, or extract data in a Puppeteer browser you already manage. It becomes relevant when you want the demonstrated DOM projection to run inside the remotely loaded page.

ScrapingAnt can execute a custom JavaScript snippet in the loaded page. The snippet can extract DOM data and write JSON into an element that the client reads from the returned HTML. Use browser=true, return_page_source=false and the general HTML endpoint. Encode the UTF-8 JavaScript as standard Base64, then URL-encode the query parameters. The custom JavaScript documentation describes that interface.

The snippet uses native browser DOM APIs. Puppeteer's ::-p-text, ::-p-xpath and shadow selector extensions are not automatically available there. The shared live example selects #prices tbody tr on a separate public table fixture; its four fields are product, price, units and share, rather than the local card example's SKU-based schema.

Write the extraction result into the returned HTML​

This is the tested catalogue projection from snippet.js. It is an excerpt inside the runner's error-handling envelope, not a standalone script: state.py inserts the selector and a unique result-marker ID into config.

const rows = [...document.querySelectorAll(config.selector)];
if (!rows.length) throw new Error("ELEMENT_MISSING");
const records = rows.map(row => {
const cells = [...row.querySelectorAll("td")].map(cell => cell.textContent.trim());
if (cells.length !== 4 || cells.some(value => !value)) throw new Error("FIELDS_MISSING");
return {product: cells[0], price: cells[1], units: cells[2], share: cells[3]};
});

The runner puts those records in a versioned result envelope and writes that envelope using this code:

const carrier = document.createElement("script");
carrier.id = config.marker;
carrier.type = "application/json";
// Script data is raw text: escape less-than, not HTML entities.
carrier.textContent = JSON.stringify(result).replace(/</g, "\\u003c");
document.body.append(carrier);

Escaping < as a JSON Unicode escape keeps data such as a closing-script string inside the carrier when the returned HTML is parsed. Do not HTML-escape the JSON or entity-decode its contents: the synthetic transport probe deliberately contains literal &amp;, quotes, a newline and Unicode. Python and Node recovered that probe exactly from all three catalogue responses. The probe is a separate transport test, not a field in the catalogue.

Returning a JavaScript object alone did not create the result marker in the live control. The general endpoint returned HTML, so the client must read and validate the explicitly written carrier.

Read one carrier, then validate the records​

Validate the result marker, JSON envelope and required data fields; an HTTP 200 response alone does not prove the extraction succeeded. The Node parseMarker helper first uses pinned parse5 to check that exactly one matching script element has the expected type. Its exact-format framing and JSON parsing then include:

const frames = [...html.matchAll(new RegExp(`<script id="${marker}" type="application/json">([^<]*)</script>`, 'g'))];
if (frames.length !== 1) throw new Error('MARKER_FORMAT');
let result;
try { result = JSON.parse(frames[0][1]); } catch { throw new Error('JSON_INVALID'); }

Use the complete helper, which also restricts the marker ID and validates the envelope. The regex is deliberately limited to the writer's exact opening tag; changed formatting fails rather than being treated as arbitrary HTML parsing. A valid ok: false envelope is an extraction failure. After ok: true, check the required records against your own contract. A unique marker prevents accidental collisions; it does not authenticate a hostile page's data.

Install the shared example before parsing a saved response​

Open examples/javascript-dom-extraction in the shared evidence snapshot. Its local reproduction uses Python 3.12 and Node 22:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
python -m playwright install chromium
npm ci
./run.sh

From examples/javascript-dom-extraction after that installation, this saved-response command reads the expected marker from the recorded receipt and runs the tested Node consumer. It makes no API request:

node --input-type=module -e 'import fs from "node:fs"; import {parseMarker} from "./parse-marker.mjs"; const run=JSON.parse(fs.readFileSync("expected_output/live/live.json","utf8")); console.log(JSON.stringify(parseMarker(fs.readFileSync("expected_output/live/catalog_1.html","utf8"),run.calls[0].marker),null,2));'

The captured data.records value was:

[
{
"product": "Ant Farm Deluxe",
"price": "$1,234.50",
"units": "12,345",
"share": "45.5%"
},
{
"product": "Ant Farm Mini",
"price": "$99.00",
"units": "1,020",
"share": "3.8%"
},
{
"product": "Magnifier",
"price": "$12.25",
"units": "987,654",
"share": "50.7%"
}
]

Repeat the paid experiment​

The default run never calls ScrapingAnt. To inspect the live plan, run the first command below. Only the second command sends requests; it requires SCRAPINGANT_API_KEY in the environment and sends exactly six requests with no automatic retries:

python live_probe.py
python live_probe.py --live

The recorded browser/datacenter run received 10 credits per response, 60 credits total. A request with JavaScript rendering through a datacenter proxy costs 10 API credits; the packet also records each actual credit receipt. Inspect the plan before intentionally repeating it.

The six live outcomes were three exact catalogue extractions, one delayed-data result, one ok: false / ELEMENT_MISSING result, and one return-only response with no marker. All had HTTP 200; only four carried successful data envelopes. These controlled observations are not a production extraction success rate.

For the delayed case, wait_for_selector targeted the page's own #loaded element, not the result marker created later by the snippet. Keep readiness and result validation separate when adapting the projection.

Limits and further reading​

The local results cover one Chrome build on one host, with three repetitions. The fixture has a known catalogue, a controlled population trigger and explicit expected records. It does not test pagination completeness, production site reliability, Firefox, cross-origin frames or closed shadow roots. The scripts close their browser/server resources, and reruns write separate output files.

📚Related Reading

How to Get All Text from a Webpage with Puppeteer

Extract whole-page text with Puppeteer. Compare innerText, textContent, DOM selection and HTML-to-text using tested hidden and dynamic-content examples.

📚Related Reading

How to Find Elements With Selenium in Python

Find one or many elements with Selenium Python. Extract complete records with scoped selectors, explicit waits, stale recovery, frames and shadow roots.

📚Related Reading

Mastering CSS Selectors in BeautifulSoup for Efficient Web Scraping

This research report delves into the intricacies of optimizing CSS selectors for BeautifulSoup, exploring best practices and advanced techniques that can significantly enhance the efficiency and resilience of web scraping projects.

Examples tested on 2026-09-28 with Node 22.20.0, Puppeteer 25.12.0 and Chrome 154.0.8037.57. Code: pinned example and captured outputs.

This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is responsible for the code, measurements and corrections.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster