NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

Playwright Local Storage: Set Before Load and Validate Data

· 10 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Playwright localStorage initialization and extracted data

Updated 2026-09-27

Replaced unsupported performance percentages and broken advanced recipes with runnable Chromium and Firefox experiments. The examples now distinguish startup state, current stored values and the records actually extracted.

To change localStorage on an already loaded origin, use page.evaluate(). If the application reads that value during startup, install an origin-guarded context.add_init_script() before navigation, or create the context with the required storage_state.

That timing matters when scraping. In our catalog fixture, setting the region to eu after the table had loaded changed the stored value but left four US-priced rows on screen. A reload produced the intended EUR records. A successful storage write alone did not validate the data.

Read, set and clear localStorage​

These are API patterns for an existing page or context; the complete runnable example follows. Evaluate storage code on the intended HTTP(S) origin. Values passed to the Python API belong in its one optional argument; package several values into an array or object. See evaluation arguments.

TaskAPI pattern
Read a stringpage.evaluate("key => localStorage.getItem(key)", "region")
Set a stringpage.evaluate("([key, value]) => localStorage.setItem(key, value)", ["region", "eu"])
Remove one keypage.evaluate("key => localStorage.removeItem(key)", "region")
Clear this originpage.evaluate("() => localStorage.clear()")
Seed before startupcontext.add_init_script(script=init_script(origin)), using the supplied helper below
Save context statesnapshot = context.storage_state()
Restore into a fresh contextbrowser.new_context(storage_state=snapshot)

localStorage stores strings. Serialize structured values with JSON.stringify() in the page and parse them when reading; the packet's JSON diagnostic preserves quotes, a backslash, a newline and nested values. Passing data as an evaluation argument avoids treating those characters as JavaScript source. Clearing storage affects the current origin, so use a disposable fixture context for these examples.

A complete example: seed before the catalog loads​

Download or check out the entire example directory, including its fixture and helpers. Use Python 3.12; the captured environment used Python 3.12.10 and Playwright 1.63.0. To check out the exact tested snapshot and run it:

git clone https://github.com/ScrapingAnt/scrapingant-examples.git
cd scrapingant-examples
git checkout 5dcd4058ba609fd86d7ffaad98deb0d8d2df6714
cd examples/playwright-local-storage
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -r requirements-lock.txt
python -m playwright install chromium
./run.sh

On Linux, install browser system dependencies with python -m playwright install --with-deps chromium. Installation downloads packages and browsers; the experiment itself uses a local server and needs no account or API key. ./run.sh runs the tests, quickstart and three Chromium rounds. Output goes to ignored run_output/, preserving the published captures.

This is the complete quickstart.py:

"""Complete no-key example; all HTTP traffic goes to an ephemeral local fixture."""
import argparse
import json
from pathlib import Path
from playwright.sync_api import sync_playwright
from catalog_fixture import serve_fixture
from state import init_script, score_records
from browser_matrix import environment, fresh_context, navigate_catalog, now


def run():
with serve_fixture() as origin, sync_playwright() as playwright:
browser = playwright.chromium.launch(headless=True)
try:
with fresh_context(browser, {'primary': origin}) as context:
context.add_init_script(script=init_script(origin))
page = context.new_page()
observed = navigate_catalog(page, origin)
result = {'captured_at': now(), 'environment': environment(browser),
**observed, **score_records(observed['records'])}
if not result['exact_match']:
raise AssertionError('Catalog differs from four literal expected records')
return result
finally:
browser.close()


if __name__ == '__main__':
parser = argparse.ArgumentParser()
parser.add_argument('--output', type=Path)
args = parser.parse_args()
result = run()
text = json.dumps(result, indent=2) + '\n'
if args.output:
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(text)
print(text, end='')

The helper in state.py checks the exact origin before writing and seeds region only when the key is missing. It is deliberately restricted to this loopback fixture. For a real target, define and validate its intended origin rather than removing the guard. Playwright runs an initialization script before the page's own scripts, including on subsequent navigations.

Selected fields from the actual quickstart output:

{
"page_status": 200,
"request_status": 200,
"response_region": "eu",
"storage_region": "eu",
"row_count": 4,
"matching_records": 4,
"exact_match": true
}

The output also contains the four DOM records, the API records and actual request URLs. The literal expected tuples are ATLAS-01 / EUR / 12.50, BIRCH-02 / EUR / 24.00, CORAL-03 / EUR / 8.75 and DELTA-04 / EUR / 41.20. The scorer does not import the fixture's catalog. It compares all three fields and rejects missing, extra or duplicate substitutes.

What changed the extracted data?​

The fixture page reads a region once during startup, in this order: URL parameter, localStorage, then its US default. Page JavaScript explicitly puts that value into /api/catalog?region=... and renders the returned records. This is application code, not an automatic browser conversion of storage into request parameters.

The primary matrix contains 72 extraction observations and 24 diagnostics: 12 extraction cases and four diagnostics, repeated three times in Chromium and Firefox. Every extraction returned HTTP 200 and four rows. Only 36 matched the four intended EU tuples. Each row below summarizes six observations with the same result.

CaseStored regionRendered recordsEU matches
Missing-only initializer before startupeuEUR catalog4/4
Write eu after the table loadseuExisting USD catalog0/4
Reload after that writeeuEUR catalog4/4
Fresh isolated contextMissingUSD catalog0/4
Restore the source storage_stateeuEUR catalog4/4
New page in the source contexteuEUR catalog4/4
Restore state, visit a different portMissingUSD catalog0/4
Initializer guarded for the other originMissingUSD catalog0/4
Missing-only initializer after changing preference to ususUSD catalog0/4
Unconditional initializer after the same changeeuEUR catalog4/4
Fresh context, URL /?region=euMissingEUR catalog4/4
Fresh context, URL without regionMissingUSD catalog0/4

Raw captures and the recomputed comparison include the DOM and API records. All 96 designed checks passed, meaning both intended successes and wrong-state controls behaved as specified. It does not mean every extraction selected the desired dataset, and it is not a production scraping success rate.

The initializer rows show why even “4/4 matches” needs context. If a user deliberately changes the preference to us, preserving that change is correct behavior. The unconditional initializer gets the original EUR data back by overwriting the new preference on reload. Choose whether a seed is a default or a forced value; do not hide that decision in a preload script.

Save the right origin, then restore deliberately​

A new page in the same browser context shared localStorage in this test. A fresh context did not inherit it. Explicit storage_state restoration populated the new context's original origin, but changing only the server port left the visited origin empty. LocalStorage belongs to an origin, including its scheme, host and port.

The state-restoration case keeps the fixture origin running while creating the new context. It tests same-browser context restoration, not a browser restart, a cross-browser login transfer or restoration after a changed deployment origin. Keep exported production state private; a snapshot can contain credentials. The quickstart uses only synthetic state and creates no state file.

The separate storage diagnostic also saved a synthetic sessionStorage value. It was absent in a new page without an opener and in a context restored from the snapshot, while localStorage was present. Playwright documents separate handling for sessionStorage. Do not infer that the context snapshot reproduces every kind of application state.

A storage event is not a page refresh​

The sibling-page diagnostic recorded one storage event in another same-origin page and none in the writing page. Both pages could read the changed value. This agrees with the storage-event contract.

Whether that event updates a catalog depends on the application. Our catalog reads its preference at startup and does not subscribe to changes. The late-write case therefore keeps the old rows until reload. For your scraper, wait for the application's observable result—such as the expected currency and record identities—after the action that actually triggers a refresh. Waiting only for the key to equal eu would accept the stale table here.

When ScrapingAnt fits this workflow​

The URL-selected case suggests a useful decision before maintaining a stateful browser workflow: does this target already expose the desired view through a reproducible URL? In the fixture, /?region=eu rendered the same EUR records in a fresh context with no stored region. The website explicitly implemented that route; you cannot assume adding a parameter works elsewhere.

With browser=true, ScrapingAnt renders the target page with JavaScript and returns its HTML. With browser=true, wait_for_selector tells ScrapingAnt to wait for the specified DOM element to appear. For a target whose required view works from its URL, that provides a way to request rendered HTML without implementing the retrieval step in your own Playwright script. Keep the same SKU, currency and price validation: the presence of a table does not establish the right data. See ScrapingAnt browser rendering.

ScrapingAnt does not support localStorage operations. Keep browser automation when your task requires setting, exporting or restoring that state. Its cookie input is a separate option only when the target actually uses cookies for the required state: ScrapingAnt accepts cookies as name/value pairs through the cookies parameter, with multiple pairs separated by semicolons. A localStorage object is not a cookie jar.

A request with JavaScript rendering through a datacenter proxy costs 10 API credits. That is the documented request cost, not a charge from this experiment. These fixtures were never sent to ScrapingAnt, and no API calls were made. ScrapingAnt is unnecessary for this local example or for retrieval you already handle directly; its relevant role is rendered-page retrieval once custom storage control is no longer a requirement.

Reproduce the second browser and interpret the limits​

To add the captured Firefox path:

python -m playwright install firefox
./run.sh --all-browsers

The packet also runs 15 boundary tests, separately from the browser matrix. They check oracle multiplicity, fixture inputs, initializer scope and rejection of incomplete or contradictory captures. The three repeats use one deterministic application on one host. They establish mechanisms, not timing gains, website-wide reliability, authentication validity or storage-quota behavior. HTTPS transitions, third-party storage partitioning and other storage systems are outside these tests.

For adjacent tasks, see Playwright cookies and browser/API sharing, Selenium localStorage and Puppeteer localStorage.

Examples tested on 2026-09-27 with Python 3.12.10, Playwright 1.63.0, Chromium 153.0.8010.12 and Firefox 155.0 on Darwin 25.6.0 arm64. Code and evidence.

This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is responsible for the article, measurements and corrections.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster