NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

71 posts tagged with "python"

View All Tags

Selenium Local Storage in Python: Set, Restore, Extract

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Selenium Local Storage in Python: Set, Restore, Extract

Updated 2026-09-27

Replaced broken advanced examples and unsupported performance claims with a runnable catalog experiment in Chrome and Firefox. The code, captures and result definitions distinguish changing storage from changing the data you extract.

Use Selenium Python's execute_script to access localStorage in the current document. Navigate to the intended origin first, pass keys and values as arguments, and include return when you need a value back in Python. If the application reads a preference at startup, set it before entering that application page or trigger its documented refresh afterward.

The last step matters for extraction. In our fixture, writing eu after the table loaded changed localStorage but left all four prices in USD. A wait that checked only the stored key would have accepted the wrong dataset. The examples below check storage, the application's request and the resulting records.

Playwright Local Storage: Set Before Load and Validate Data

· 10 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Playwright localStorage initialization and extracted data

Updated 2026-09-27

Replaced unsupported performance percentages and broken advanced recipes with runnable Chromium and Firefox experiments. The examples now distinguish startup state, current stored values and the records actually extracted.

To change localStorage on an already loaded origin, use page.evaluate(). If the application reads that value during startup, install an origin-guarded context.add_init_script() before navigation, or create the context with the required storage_state.

That timing matters when scraping. In our catalog fixture, setting the region to eu after the table had loaded changed the stored value but left four US-priced rows on screen. A reload produced the intended EUR records. A successful storage write alone did not validate the data.

Avoid Detection with Puppeteer Stealth

· 7 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Avoid Detection with Puppeteer Stealth

Puppeteer is a powerful Node.js library that provides a high-level API for controlling browsers through the DevTools Protocol. It is commonly used for testing, web scraping, and automating repetitive browser tasks. However, Puppeteer's default settings can trigger bot detection systems, especially in headless mode.

Playwright Stealth: 5 Libraries Tested Against Bot Detectors

· 13 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Playwright Stealth: 5 Libraries Tested Against Bot Detectors

Updated 2026-09-22

Replaced the legacy Playwright 1.40 recipe and historical screenshots with five library integrations, plain-browser controls and dated BrowserScan/Sannysoft observations. A diagnostic result does not establish that a scraper is undetectable or will reach a protected target. The tested code, raw results and screenshots preserve both successful captures and excluded harness errors.

Which Playwright stealth library should you use? In this comparison, Patchright returned BrowserScan's Normal verdict when headed, while its headless user agent was flagged. Node's playwright-extra with the stealth plugin and the Firefox-based Camoufox returned Normal in both modes. Python playwright-stealth had no failed Sannysoft checks but still received BrowserScan's Navigator flag.

That difference is the useful result: changing a browser property, passing a diagnostic page and successfully retrieving your target are separate tests. Start with a reproducible control and keep the browser version, operating system and launch mode attached to every observation.

Playwright Cookies: Save State and Share Sessions with API Requests

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Playwright Cookies: Save State and Share Sessions with API Requests

Updated 2026-09-27

Replaced mixed sync/async recipes, manual Set-Cookie parsing and untested state assumptions with runnable Chromium and Firefox experiments. The new examples test context restoration, browser/API cookie sharing, missing local storage and duplicate requests against extracted catalog records. The original URL, banner and publication date are preserved.

To set cookies with Playwright Python, call context.add_cookies() with a list of cookie dictionaries. Supply url, or a suitable domain and path, and set the cookies before navigating when the first request needs them. Use context.storage_state() when the workflow also depends on supported browser storage beyond cookies.

Restoring the cookie does not necessarily restore the dataset. In our catalog, cookies-only restoration kept member pricing but selected USD instead of EUR. A browser-associated API request made the same mistake when it omitted a region parameter that the page normally reads from local storage.

This guide shows the working state handoff and the failures around it. The fixture is self-authored and synthetic; it demonstrates specific mechanisms, not the probability that a real website will accept a restored login.

Open Source Web Scraping Libraries to Bypass Anti-Bot Systems

· 7 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Open Source Web Scraping Libraries to Bypass Anti-Bot Systems

Approximately one in five websites targeted for scraping employ advanced anti-bot systems that can easily result in access being blocked. These systems, such as Cloudflare, DataDome, and PerimeterX, are designed to detect and block automated access, making it increasingly difficult for traditional scraping tools to function effectively.

To address these challenges, a variety of open-source libraries have emerged, each offering unique features and techniques to bypass these anti-bot mechanisms.

JavaScript vs Python for Web Scraping - Which Is Best?

· 8 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

JavaScript vs Python for Web Scraping: Which Is Best?

In the rapidly evolving landscape of web technologies, web scraping has emerged as a crucial tool for data extraction and analysis. As of 2024, two programming languages, JavaScript and Python, stand out as popular choices for developers engaging in web scraping tasks. Each language offers unique strengths and capabilities, making the decision between them a significant consideration for developers at all levels.

Playwright vs. Puppeteer in 2024 - Which Should You Choose?

· 9 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Playwright vs. Puppeteer in 2024: Which Should You Choose?

In the ever-evolving landscape of web automation and testing, two tools have consistently stood out: Playwright and Puppeteer. As of 2024, both have matured significantly, offering robust features for developers and testers alike. Both tools, developed by teams at Microsoft and Google respectively, offer robust solutions for automating browser tasks, but they cater to slightly different needs and preferences.

Playwright vs. Selenium - A Comprehensive Comparison for 2024

· 7 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Playwright vs. Selenium - A Comprehensive Comparison for 2024

In the rapidly evolving landscape of web automation and testing, two open-source frameworks have emerged as leading tools: Playwright and Selenium. Both frameworks offer unique features and capabilities, making the choice between them a nuanced decision that depends on specific project requirements and team expertise.

Top Python HTTP Clients for Web Scraping

· 10 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Top Python HTTP Clients for Web Scraping

In the ever-evolving landscape of web scraping, Python remains the language of choice for developers due to its simplicity, readability, and a robust ecosystem of libraries. Python offers a diverse array of HTTP clients that cater to various web scraping needs, from simple data extraction to complex, high-concurrency tasks.

This guide delves into the top Python HTTP clients, exploring their features, pros, cons, and providing code examples to get started.