NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

9 posts tagged with "selenium"

View All Tags

Changing User Agent in Selenium for Effective Web Scraping

· 6 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Changing User Agent in Selenium for Effective Web Scraping

As of October 2024, with web technologies advancing rapidly, the need for sophisticated techniques to interact with websites programmatically has never been more pressing. This comprehensive guide focuses on changing user agents in Python Selenium, a powerful tool for web automation that has gained significant traction in recent years.

User agents, the strings that identify browsers and their capabilities to web servers, play a vital role in how websites interact with clients. By manipulating these identifiers, developers can enhance the anonymity and effectiveness of their web scraping scripts, avoid detection, and simulate various browsing environments. According to recent statistics, Chrome dominates the browser market with approximately 63% share (StatCounter), making it a prime target for user agent spoofing in Selenium scripts.

The importance of user agent manipulation is underscored by the increasing sophistication of bot detection mechanisms. This guide will explore various methods to change user agents in Python Selenium, from basic techniques using ChromeOptions to more advanced approaches leveraging the Chrome DevTools Protocol (CDP) and third-party libraries.

As we delve into these techniques, we'll also discuss the importance of user agent rotation and verification, crucial steps in maintaining the stealth and reliability of web automation scripts. With JavaScript being used by 98.3% of all websites as of October 2024 (W3Techs), understanding how to interact with modern, dynamic web pages through user agent manipulation is more important than ever for developers and data scientists alike.

Selenium Cookies in Python: Set, Save and Validate Scraped Data

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Selenium Cookies in Python: Set, Save and Validate Scraped Data

Updated 2026-09-27

Replaced mixed-language API examples, cookie insertion on about:blank and untested session recipes with runnable extraction tests. The evidence includes Chrome and Firefox cookie restoration, incorrect-data cases, and a separate ScrapingAnt test that receives and resends synthetic cookies. The original URL and publication date are preserved.

To set a cookie in Selenium Python, first navigate to its target origin, then call driver.add_cookie() with a name/value dictionary. Navigate or refresh afterward so the next request uses it. To reuse a session, save the relevant cookies, restore them on the same origin in the new browser, and check the resulting data.

That last step matters. In our controlled catalog, every extraction returned HTTP 200 and four product rows. Restoring only the session cookie still selected member prices—but in USD when the intended dataset was EUR. A successful navigation, an authenticated-looking page and the expected row count all missed the error.

This guide shows the cookie operations, a complete runnable save/restore example, and the measured failures behind that distinction. The catalog and accounts are synthetic; the results demonstrate mechanisms, not production-site success rates.

Selenium Local Storage in Python: Set, Restore, Extract

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Selenium Local Storage in Python: Set, Restore, Extract

Updated 2026-09-27

Replaced broken advanced examples and unsupported performance claims with a runnable catalog experiment in Chrome and Firefox. The code, captures and result definitions distinguish changing storage from changing the data you extract.

Use Selenium Python's execute_script to access localStorage in the current document. Navigate to the intended origin first, pass keys and values as arguments, and include return when you need a value back in Python. If the application reads a preference at startup, set it before entering that application page or trigger its documented refresh afterward.

The last step matters for extraction. In our fixture, writing eu after the table loaded changed localStorage but left all four prices in USD. A wait that checked only the stored key would have accepted the wrong dataset. The examples below check storage, the application's request and the resulting records.

Playwright vs. Selenium - A Comprehensive Comparison for 2024

· 7 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Playwright vs. Selenium - A Comprehensive Comparison for 2024

In the rapidly evolving landscape of web automation and testing, two open-source frameworks have emerged as leading tools: Playwright and Selenium. Both frameworks offer unique features and capabilities, making the choice between them a nuanced decision that depends on specific project requirements and team expertise.

How to use Selenium Wire in 2024

· 11 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

How to use Selenium Wire in 2024

Web scraping has become an essential technique for extracting data from websites, especially in an era where data-driven decision-making is paramount. Among the myriad of tools available for web scraping, Selenium stands out due to its ability to interact with web pages like a real user.

However, when it comes to accessing and manipulating network traffic, Selenium's capabilities are limited. This is where Selenium Wire comes into play, offering a powerful extension to the standard Selenium library.

This blog delves into various aspects of Selenium Wire, covering its installation, configuration, and features. It includes details on capturing and modifying HTTP requests, proxy configuration, and advanced request blocking techniques to enhance performance. Additionally, it delves into advanced techniques for request blocking, optimization of performance, and troubleshooting common issues.

Download Files with Selenium in Python: Chrome and Firefox

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Download Files with Selenium in Python: Chrome and Firefox

Updated 2026-09-21

Re-tested Chrome and Firefox downloads with Selenium 4.49.0 in headed and headless modes. Replaced legacy APIs and fixed sleeps with an explicit file-verification helper, and removed the broken JavaScript and Safe Browsing workaround. The runnable evidence packet includes the fixtures, captured output and failure cases.

To download a file with Selenium in Python, configure the browser's download directory before starting the session, click the link, and wait for the expected file to pass a content check before closing the browser. Waiting for the link to be clickable and waiting for the file to finish are separate steps.

Below are tested Chrome and Firefox configurations, including headless runs. The examples download a CSV, a delayed binary file and a PDF from a local fixture server. They verify the expected length and SHA-256 digest rather than trusting that a filename appeared.

How to Find Elements With Selenium in Python

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Find Elements With Selenium in Python

Updated 2026-09-28

Replaced the earlier locator examples with a runnable local catalog and captured Chrome/Firefox results. Corrected singular/plural lookup, class-token matching, stale-reference recovery and shadow-root access; removed unsupported locator speed rankings.

Use driver.find_element(By.CSS_SELECTOR, selector) for the first matching element and driver.find_elements(By.CSS_SELECTOR, selector) for a list of all matches. If nothing matches, the singular method raises NoSuchElementException; the plural method returns an empty list.

For scraping, finding elements is only the first step. You also need the right container, populated fields and a way to recover when the page replaces nodes. The examples below extract a complete catalog and deliberately test selectors that return plausible but wrong results.

Puppeteer vs. Selenium - Which Is Better? + Bonus

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Puppeteer Vs. Selenium: Which Is Better?

With the increasing use of the internet worldwide, it is being implemented in all aspects of our daily lives. So using it efficiently and effectively becomes crucial and could be the difference between competitors and businesses. This is where the concept of Web Automation comes in. Today I shall teach you one of the most debated topics of web automation, Puppeteer vs. Selenium.

Let's begin!

Scrape a Dynamic Website with Python

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Scrape a Dynamic Website with Python

Updated 2026-09-14

The Selenium and Playwright examples were re-run with Selenium 4.49, Playwright 1.62 and Python 3.12, and their output below is copied from that run. The Selenium example was rewritten because webdriver.Chrome() no longer accepts the executable_path argument used in the 2021 version. The Pyppeteer section was removed because the library is unmaintained. A second fixture was added for content that arrives after a delay. The ScrapingAnt API example was updated to the same fixtures but was not executed for this update; no API output is shown as measured. Code, fixtures and captured output: scrapingant-examples/examples/scrape-dynamic-website-with-python.

You open a page in the browser and the data is there. You fetch the same URL with requests, parse it with BeautifulSoup, and the field is empty or shows placeholder text. That is the symptom of a dynamic website: the HTML the server sends is not the HTML you see, because JavaScript changes the page after it loads. This tutorial shows how to confirm that diagnosis on a small fixture and how to get the rendered content with Selenium, Playwright, or the ScrapingAnt API, including the case where the content shows up only after a delay.