NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

How to Get All Text from a Webpage with Puppeteer

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Get All Text from a Webpage with Puppeteer

Correction (2026-09-28)

The earlier article described DOM Range selection as Ctrl+A/copy-paste and claimed it could recover text from sites using display optimizations. The example performed neither keyboard input nor clipboard copying, and the general recovery claim was unsupported. This refresh compares actual text returned by five methods on a controlled fixture. Code and captured outputs.

For the readable text of an already rendered page, start with page.$eval('body', body => body.innerText). Use textContent when you explicitly want descendant text that can include hidden content and script/style source. If the HTML already contains the data, a converter such as html-to-text can work without a browser.

Those methods produce different outputs. Below, the browser sees three dynamically populated products, while the raw HTTP response contains empty product placeholders. The examples show how that difference affects text extraction, and when to keep structured records instead of flattening the page.

Choose the text you actually need​

NeedStarting pointImportant boundary
Readable text from the rendered bodybody.innerTextWait for the content you need to render
Descendant text including hidden/source contentbody.textContentScript and style text can enter the result
Text from a browser selectionSelection.toString() over a DOM RangeThis is not a keyboard or clipboard operation
Text from already retrieved HTMLhtml-to-textConversion does not execute the page's JavaScript
Product titles with their prices and URLsA scoped JSON projectionPreserve record relationships before flattening

The distinction between innerText and textContent is also documented by MDN. The output comparisons below are from the supplied fixture, rather than a promise about every website's layout or loading behavior.

Run the complete text comparison​

Download the example repository at the recorded snapshot and open examples/puppeteer-find-elements. The packet contains the fixture and both articles' executable examples. Use Node 22.12 or newer; the recorded version was 22.20.0.

git clone https://github.com/ScrapingAnt/scrapingant-examples.git
cd scrapingant-examples
git checkout 74638f75381bdf2ff0652f1f07bd98f65186b40d
cd examples/puppeteer-find-elements
npm ci
node text-quickstart.mjs

The dependency lock pins Puppeteer 25.12.0 and html-to-text 9.0.5. This complete script starts the local HTTP fixture, waits for its population handshake, collects the five outputs and closes its resources:

import assert from 'node:assert/strict';
import {writeFile,mkdir} from 'node:fs/promises';
import {dirname} from 'node:path';
import {convert} from 'html-to-text';
import {startFixture} from './fixture.mjs';
import {launchBrowser} from './browser.mjs';
import {TEXT_EXPECTED,sentinels} from './oracle.mjs';
const fixture=await startFixture();let browser;
try{
browser=await launchBrowser();const page=await browser.newPage();
await page.goto(fixture.origin+'/catalog',{waitUntil:'load'});
await page.evaluate(()=>window.releaseCatalog());
await page.waitForFunction(()=>document.querySelector('#catalog')?.dataset.ready==='yes',{timeout:5000});
const rawResponse=await fetch(fixture.origin+'/catalog');assert.equal(rawResponse.status,200);
const rawHTML=await rawResponse.text();
const texts={
body_inner_text:await page.$eval('body',body=>body.innerText),
body_text_content:await page.$eval('body',body=>body.textContent),
body_selection:await page.evaluate(()=>{
const selection=window.getSelection();const range=document.createRange();range.selectNodeContents(document.body);
selection.removeAllRanges();selection.addRange(range);const text=selection.toString();selection.removeAllRanges();return text;
}),
raw_html_to_text:convert(rawHTML,{wordwrap:false}),
rendered_html_to_text:convert(await page.content(),{wordwrap:false}),
};
const result=Object.fromEntries(Object.entries(texts).map(([method,text])=>{const markers=sentinels(text);assert.deepEqual(markers,TEXT_EXPECTED[method]);return [method,{text,sentinels:markers}];}));
console.log(JSON.stringify(result,null,2));const output=process.argv[2]??'run_output/text-quickstart.json';await mkdir(dirname(output),{recursive:true});await writeFile(output,JSON.stringify(result,null,2)+'\n');
}finally{await browser?.close();await fixture.close();}

releaseCatalog() belongs to the controlled fixture. It separates a page containing empty product cards from a page whose product fields are populated. On a real site, wait for the application-specific content you need instead of copying this test function. The fixture's data-ready marker is set only after its card population finishes.

Get the rendered body's text with innerText​

The body_inner_text.text value in the captured quickstart output was:

VISIBLE_SENTINEL
Promotional decoy0.01
Ant Field Kit 19.95
Sugar & Water Feeder 7.50
Tunnel Kit — Mini 12.00

This includes the promotional text as well as the products, because the query asks for the whole body. It excludes the fixture's hidden, script and style sentinels. The title-price spacing comes from the page's rendered text; do not treat those line breaks as a record schema.

If the task concerns one part of the page, scope the query to that container. If you need product titles paired with prices and links, read each card separately and return objects. The Puppeteer element-finding guide demonstrates that workflow using the same records.

Understand what textContent adds​

The fixture places HIDDEN_SENTINEL in a hidden element and SCRIPT_SENTINEL/STYLE_SENTINEL inside body script/style nodes. Its body_text_content output contained all three, in addition to VISIBLE_SENTINEL and the catalogue title. The saved diagnostic flags were:

{
"visible": true,
"hidden": true,
"script": true,
"style": true,
"catalog": true
}

The full value also includes the fixture's JavaScript source, including product-title strings inside its data array. Finding a desired word anywhere in body.textContent therefore does not prove that its product has rendered. Scope the extraction to the intended DOM records and check the fields there.

Use textContent deliberately for the text contained in the chosen nodes. It is not a drop-in replacement for the rendered body text when hidden or executable source content is unwanted.

DOM Range selection is a separate browser operation​

The complete script creates a range over the body's contents, installs it in the browser's Selection, reads selection.toString(), then clears it. Range.selectNodeContents() defines the range; no Ctrl+A keypress or clipboard access occurs.

In this fixture, Selection included the same declared visible/catalogue sentinels as innerText and excluded the hidden/script/style sentinels. Its exact saved string had a trailing newline:

"VISIBLE_SENTINEL\nPromotional decoy0.01\nAnt Field Kit 19.95\nSugar & Water Feeder 7.50\nTunnel Kit — Mini 12.00\n"

That observation does not establish that Selection recovers text unavailable to innerText on Bing or other sites. It also does not make content that has never entered the DOM available. The example changes and clears the page's selection; preserve an existing selection if your application needs it.

Convert raw HTML without rendering JavaScript​

The script makes a separate HTTP request to the local catalogue and passes that response body to convert(rawHTML, {wordwrap:false}). This is the same document source before browser-side population. Its captured text was:

VISIBLE_SENTINEL

HIDDEN_SENTINEL
Promotional decoy [/promotion]0.01
/p/a101
/p/b202
/p/c303

The converter found the placeholder links but none of the populated product titles. It did not execute the fixture's script. It also included the hidden sentinel: HTML conversion in this test did not apply the browser's rendered-visibility rules.

Use this approach when the retrieved HTML already contains the content you need. In this fixture, the converter’s raw-HTML result omitted the populated titles. Reading structured data embedded in scripts would be a separate extraction approach.

Convert rendered HTML when you need converter formatting​

Passing await page.content() to the same converter after population produced:

VISIBLE_SENTINEL

HIDDEN_SENTINEL
Promotional decoy [/promotion]0.01
Ant Field Kit [/p/a101] 19.95
Sugar & Water Feeder [/p/b202] 7.50
Tunnel Kit — Mini [/p/c303] 12.00

Now the product titles and prices appear, and the converter preserves the source link paths in its output. The hidden sentinel still appears. Choose this path when you need the converter's formatting and already have a rendered page; it does not remove the need to decide which elements belong in the output.

The complete fixture comparison was:

MethodVisible textHidden sentinelScript sentinelStyle sentinelCatalogue-title sentinel
body.innerTextyesnononoyes
body.textContentyesyesyesyesyes
DOM Range + Selectionyesnononoyes
html-to-text on raw HTTP HTMLyesyesnonono
html-to-text on rendered HTMLyesyesnonoyes

Each method was captured in three Chrome rounds, giving 15 text observations. The checks compare declared sentinels; they do not promise identical whitespace across browsers or platforms. The complete raw capture retains every text value.

Do not use whole-page text as a product schema​

The visible body includes a promotional decoy and three catalogue records. Flattening them gives you a text stream, not a validated association between each SKU, title, currency, price and href. The companion quickstart.mjs scopes selection to #catalog > .card and returns one object per card.

The same packet also checked missing-match behavior: $ returned null, $$ returned an empty array, $eval threw an Error naming the missing selector, and the tested $$eval mapping returned an empty array. Decide whether an empty result is valid for your task; otherwise enforce the expected record contract instead of saving an apparently successful empty scrape.

When to use ScrapingAnt for text or record extraction​

You do not need ScrapingAnt for this local comparison, for HTML you already have, or for text extraction in a Puppeteer browser you already manage. Static HTML-to-text conversion also needs no browser rendering when the source already contains the required data.

ScrapingAnt can execute a custom JavaScript snippet in the loaded page. The snippet can extract DOM data and write JSON into an element that the client reads from the returned HTML. This provides a way to run a DOM projection remotely, then consume its structured result in Node. Use native DOM queries in the snippet; Puppeteer-specific selector syntax is not automatically available inside the remote page. The JavaScript execution documentation describes the interface.

The tested shared example extracted a catalogue table and one delayed text field. It did not run a remote whole-page innerText, textContent or Selection comparison. Adapting its projection to document.body.innerText would be a separate application to test against the target page. The local method comparison above remains the evidence for those whole-page text choices.

The demonstrated transport is useful for either a chosen text field or related records: build a result envelope in the page, write it to one script[type="application/json"] marker, then parse the marker from the returned HTML. The tested writer escapes < in JSON.stringify(result) as a JSON Unicode escape so a closing-script string remains data. Do not entity-decode script JSON. The companion element guide shows the actual writer, bounded marker parser and Node consumer.

Validate the result marker, JSON envelope and required data fields; an HTTP 200 response alone does not prove the extraction succeeded. In the recorded live controls, simply returning a JavaScript object produced no marker, while a missing selector produced a valid envelope with ok: false and ELEMENT_MISSING. Both API responses had HTTP 200. The expected data is whatever your projection explicitly returns and validates, not whatever text happens to appear in a successful HTTP response.

To reproduce the shared packet locally, open examples/javascript-dom-extraction in the linked snapshot and use Python 3.12 with Node 22:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
python -m playwright install chromium
npm ci
./run.sh

That default run makes no paid requests and includes parsing saved live responses. For a new live run, set SCRAPINGANT_API_KEY in your environment. The first command below prints the plan; the explicit --live command sends six requests without automatic retries:

python live_probe.py
python live_probe.py --live

The captured browser/datacenter run cost 60 credits: six receipts of 10 credits each. A request with JavaScript rendering through a datacenter proxy costs 10 API credits. Three calls returned all expected catalogue records, one returned the delayed text result, one reported the deliberate missing-selector error, and one omitted the marker in the return-only control. These are transport and extraction observations on owned fixtures, not a sitewide reliability claim.

The delayed request waited for the page's own #loaded element. A result marker created by the snippet is not an appropriate pre-snippet wait target. Keep the waiting condition tied to page content and validate the result separately.

These results come from one self-authored fixture and Chrome build, with an explicit population handshake. They do not measure extraction speed, production reliability, Bing behavior, browser clipboard access, lazy-loading completeness or a general solution to display optimizations. The local Puppeteer scripts make no paid requests; the shared live probe is an explicit paid opt-in. The packet's 31 boundary tests validate record contracts and capture integrity separately from the 15 text observations.

Examples tested on 2026-09-28 with Node 22.20.0, Puppeteer 25.12.0, Chrome 154.0.8037.57 and html-to-text 9.0.5. Code: pinned examples and captured outputs.

This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is responsible for the code, measurements and corrections.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster