NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

69 posts tagged with "python"

View All Tags

Find Elements with Playwright in Python: Extract Validated Records

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Find Elements with Playwright in Python: Extract Validated Records

Use page.locator("#catalog > .product") to find repeated elements in Python Playwright. Narrow to an intended record before reading a single value, or call all() to iterate the matches after the list has finished loading. To build records in the browser, use evaluate_all() with a JavaScript projection.

Finding a match is only part of extraction. In the examples below, .first returned a sponsored product, while an early count() found a loading placeholder. Both operations worked; both selected the wrong data. The runnable example waits for the catalog's ready state and checks every SKU, title, currency, price and link against independently specified records.

How to Scrape Tripadvisor Data Using ScrapingAnt's Web Scraping API in Python

· 6 min read
Tanweer Ali
Tanweer Ali
I'm a technical writer with a background in software development who loves writing about data and automation software.

How to Scrape Tripadvisor Data Using ScrapingAnt's Web Scraping API in Python

Tripadvisor is without a doubt one of the biggest travel platforms out there travelers will consult to find out about the next hot summer destination.

It's a goldmine for user reviews and ratings of hotels, restaurants and vacation rentals.

In this short tutorial we will be scraping the names, reviews and standard prices of hotels in Python using ScrapingAnts Web Scraping API.

How to scrape dynamic websites with Scrapy Splash

· 8 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to scrape dynamic websites with Scrapy Splash

Handling dynamic websites with JavaScript-rendered content presents a significant challenge for traditional scraping tools. Scrapy Splash emerges as a powerful solution by combining the robust crawling capabilities of Scrapy with the JavaScript rendering prowess of the Splash headless browser. This comprehensive guide explores the integration and optimization of Scrapy Splash for effective dynamic website scraping.

Scrapy Splash has become an essential tool for developers and data scientists who need to extract data from JavaScript-heavy websites. The middleware (scrapy-plugins/scrapy-splash) seamlessly bridges Scrapy's asynchronous architecture with Splash's rendering engine, enabling the handling of complex web applications. This integration provides a robust foundation for handling modern web applications while maintaining high performance and reliability.

The system's architecture is specifically designed to handle the challenges of dynamic content rendering while ensuring efficient resource utilization.

Python Requests Cookies: Sessions, Scope and Scraped Data

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Python Requests Cookies: Sessions, Scope and Scraped Data

Updated 2026-09-27

Replaced untested persistence and recovery recipes with a runnable Requests cookie experiment. The new examples test redirect cookies, path collisions, prepared requests and session restoration against extracted records. Corrected unsupported cookie keyword arguments and the claim that mounting an HTTPS adapter forces HTTPS. The original URL, publication date and video are preserved.

Use a requests.Session() when later requests need cookies received earlier. Keep its cookie jar intact: a cookie is identified by more than its name, and a plain dictionary cannot represent two region cookies with different paths. For a single call, the cookies= parameter can supply cookies without making them persistent session state.

We tested those distinctions against a small catalog. A dictionary backup retained the member session but changed the selected currency from EUR to USD. The request returned HTTP 200 and all four expected product IDs. Only checking the currency and prices caught the wrong dataset.

This guide gives you a complete save/restore example and the failures behind it. The catalog, identities and prices are synthetic; the measurements demonstrate mechanisms, not success rates on real websites.

Web Scraping with VPN and Python

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping with VPN and Python

Web scraping with VPN integration has become an essential practice in modern data collection strategies, combining the need for efficient data gathering with robust privacy and security measures. As organizations increasingly rely on web-based data for business intelligence and research, the implementation of VPN-enabled scraping solutions has evolved into a sophisticated technical domain. According to ScrapingAnt's implementation guide, the integration of VPNs with web scraping not only provides enhanced anonymity but also enables more reliable and sustainable data collection operations. The combination of Python's powerful scraping libraries with VPN technology creates a robust framework for handling large-scale data extraction while maintaining privacy and avoiding IP-based restrictions. Proper VPN implementation in web scraping projects has become crucial for maintaining consistent access to target websites while ensuring compliance with various access policies and restrictions. This research explores the technical implementations, best practices, and advanced techniques necessary for successfully combining VPN services with Python-based web scraping operations.

Web Scraping with Tor and Python

· 6 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping with Tor and Python

Web scraping has become an essential tool for gathering information at scale. However, with increasing concerns about privacy and data collection restrictions, anonymous web scraping through the Tor network has emerged as a crucial methodology. This comprehensive guide explores the technical implementation and optimization of web scraping using Tor and Python, providing developers with the knowledge to build robust, anonymous data collection systems.

The integration of Tor with Python-based web scraping tools offers a powerful solution for maintaining anonymity while collecting data. Proper implementation of anonymous scraping techniques can significantly enhance privacy protection while maintaining efficient data collection capabilities. The combination of Tor's anonymity features with Python's versatile scraping libraries creates a framework that addresses both security concerns and performance requirements in modern web scraping applications.

Proxy Rotation Implementation in Playwright

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Proxy Rotation Implementation in Playwright

This comprehensive guide explores the intricate details of proxy rotation implementation, drawing from extensive research and industry best practices. Proper proxy rotation can significantly reduce detection rates and improve scraping success rates by up to 85%. The implementation of proxy rotation in Playwright involves multiple sophisticated approaches, from dynamic pool management to geolocation-based rotation strategies. The key to successful proxy rotation lies in maintaining a balance between performance, reliability, and anonymity. This research delves into various implementation methods, best practices, and optimization techniques that enable developers to create robust proxy rotation systems within the Playwright framework. The guide addresses critical aspects such as authentication, monitoring, load balancing, and error handling, providing practical solutions for common challenges faced in proxy rotation implementation.

Best Web Scraping Detection Avoidance Libraries for Python

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Best Web Scraping Detection Avoidance Libraries for Python

As websites implement sophisticated anti-bot systems, developers require robust tools to maintain efficient and reliable data collection processes. According to ScrapeOps' analysis, approximately 20% of websites now employ advanced anti-bot systems, making detection avoidance a critical consideration for web scraping projects. This research examines the five most effective Python libraries for web scraping detection avoidance, analyzing their features, performance metrics, and implementation complexities. These tools range from sophisticated proxy management systems to advanced browser automation solutions, each offering unique approaches to circumvent detection mechanisms. The analysis encompasses both traditional request-based methods and modern browser-based solutions, providing a comprehensive overview of the current state of detection avoidance technology in Python-based web scraping.

How to Change User Agent in HTTPX

· 5 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Change User Agent in HTTPX

HTTPX, a modern HTTP client for Python, offers robust capabilities for handling user agents, which play a vital role in how web requests are identified and processed. This comprehensive guide explores the various methods and best practices for implementing and managing user agents in HTTPX applications. User agents, which identify the client software making requests to web servers, are essential for maintaining transparency and avoiding potential blocking mechanisms. The proper implementation of user agents can significantly impact the success rate of web requests, particularly in scenarios involving web scraping or high-volume API interactions. This research delves into various implementation strategies, from basic configuration to advanced rotation techniques, providing developers with the knowledge needed to effectively manage user agents in their HTTPX applications.