NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

251 posts tagged with "data extraction"

View All Tags

Understanding the High Cost of Residential Proxies

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Understanding the High Cost of Residential Proxies

In the rapidly evolving landscape of internet technologies, residential proxies have emerged as a critical tool for businesses and researchers seeking to access geo-restricted content, conduct market research, and perform large-scale web scraping operations. However, the high cost associated with these services has become a significant point of discussion within the industry. This comprehensive report delves into the multifaceted factors contributing to the elevated prices of residential proxies and examines the complex market dynamics shaping this sector.

At the heart of the cost issue lies the scarcity of residential IP addresses. As the internet continues its exponential growth, the pool of available IPv4 addresses has become increasingly depleted (Harvard Business School). This scarcity has given rise to a second-hand market for IP addresses, driving up costs and creating new challenges for proxy providers (VMBlog).

Beyond the issue of scarcity, the operational complexities involved in maintaining a vast and distributed network of residential IPs contribute significantly to the high costs. Unlike datacenter proxies, residential proxies rely on a decentralized infrastructure that spans multiple geographic locations and involves real residential internet connections. This decentralized nature introduces additional challenges in terms of stability, management, and performance optimization (Infatica).

Ethical considerations and regulatory compliance also play a crucial role in the cost structure of residential proxy services. Reputable providers must navigate a complex landscape of legal requirements, including data protection laws like GDPR, while ensuring that their IP sources are ethically obtained with proper user consent (Geekflare).

This report will explore these factors in detail, providing insights into the technical aspects of residential proxy networks, the strategies employed by premium providers to differentiate their services, and the innovative solutions being developed to address the challenges in this field. We will also examine pricing models, performance metrics, and real-world use cases to provide a comprehensive understanding of the residential proxy market.

To illustrate the practical implementation of residential proxies, we will include code samples in popular programming languages such as Python and JavaScript, demonstrating how these tools can be effectively utilized in various scenarios. By the conclusion of this report, readers will have gained a thorough understanding of the factors driving the high costs of residential proxies and the complex market dynamics that shape this essential component of modern internet infrastructure.

Top Python HTTP Clients for Web Scraping

· 10 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Top Python HTTP Clients for Web Scraping

In the ever-evolving landscape of web scraping, Python remains the language of choice for developers due to its simplicity, readability, and a robust ecosystem of libraries. Python offers a diverse array of HTTP clients that cater to various web scraping needs, from simple data extraction to complex, high-concurrency tasks.

This guide delves into the top Python HTTP clients, exploring their features, pros, cons, and providing code examples to get started.

Web Scraping with Playwright Series Part 4 - Avoid Getting Blocked

· 12 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Web Scraping with Playwright Series Part 4 - Avoid Getting Blocked

In Part 3, we focused on analyzing and cleaning the extracted data to address potential issues like missing values, inconsistencies, and outliers. To make it easier for future decision-making, we saved the cleaned data in various formats, such as CSV, databases, and S3 buckets.

In Part 4, we'll delve into strategies for bypassing common web scraping hurdles. We'll explore techniques such as using proxies, rotating user agents, and leveraging web scraping APIs to keep your scraping tasks running smoothly.

Without further ado, let’s get started!

Web Scraping with Playwright Series Part 3 - Storing Data

· 23 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Web Scraping with Playwright Series Part 3 - Storing Data

In Part 2, we talked about creating a web scraper with Playwright to extract data from the Nike website, which has dynamically loaded content.

In Part 3, we will focus on carefully analyzing the extracted data and ensuring it's properly cleaned to deal with potential issues like missing values, inconsistencies, and outliers. The cleaned data will then be stored in different formats such as CSV, databases, and S3 buckets to make it easier for future decision-making.

Without further ado, let’s get started!

Web Scraping with Playwright Series Part 2 - Building a Scraper

· 19 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Web Scraping with Playwright Series Part 2 - Building a Scraper

Correction (2026-09-28)

Fixed an async function declaration and request.resource_type property access, awaited routing operations, and removed blanket XHR blocking that could prevent collecting product data. An unchanged scroll height is a stopping heuristic, not proof of complete extraction. This correction does not claim a fresh live Nike test.

In Part 1, you learned about the basics of Playwright, environment setup, browser launching, and taking screenshots.

In Part 2, you’ll learn how to build a scraper from scratch. We'll cover how to locate and extract data, manage dynamically loaded content, utilize Playwright's network event feature, and improve the scraper's performance by blocking unnecessary resources.

Without further ado, let’s get started!

Web Scraping with Playwright Series Part 1 - Getting Started

· 7 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Web Scraping with Playwright Series Part 1 - Getting Started

Introducing the 4-Part Series on Web Scraping with Playwright! This comprehensive series will delve into web scraping using Playwright, a powerful and versatile tool for automating browser interactions.

By the end of this series, you'll have a solid understanding of web scraping with Playwright. You'll be able to build robust scrapers that can handle dynamic content, efficiently store data, and navigate through anti-scraping mechanisms.

In Part 1, you'll learn about the basics of Playwright, why it's useful, how to set up the environment, how to launch the browser using Playwright, and how to take screenshots.

Requests vs HTTPX in 2026: Measured Differences and Migration Gotchas

· 26 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Requests vs HTTPX in 2026: Measured Differences and Migration Gotchas

Updated 2026-09-17

Every timing on this page now comes from one local server run by the evidence packet on requests 2.34.2 and httpx 0.28.1. The earlier version's benchmark table, which quoted third-party posts and unsourced figures, is gone. New: connection counting, threads vs asyncio, HTTP/2 on the same server, what "retries" means in each library, the errors a migration hits, and what the httpx 1.0 pre-release does to 0.28 code.

requests and httpx make the same request with almost the same line of code. The differences that matter are elsewhere: whether a connection is reused, how you run 200 requests at once, which protocol you speak, what happens by default on a redirect or a slow server, and what "retry" means. This page measures each of those on one server, then lists what breaks when you move a requests codebase to httpx 0.28.

The short version, from the tables below (Apple Silicon laptop, loopback, Python 3.12.11; the ratios are the claim, not the seconds):

Caserequests 2.34.2httpx 0.28.1
300 sequential GETs, one Session / Client (plain HTTP)0.150 s, 1 connection0.126 s, 1 connection
300 sequential GETs, no Session / Client (plain HTTP)0.217 s, 300 connections3.090 s, 300 connections (an SSL context is built per call)
200 GETs with a 50 ms server delay, 20 threads0.543 s (32 connections; 20 with pool_maxsize=20)0.545 s, 20 connections
Same 200 GETs, AsyncClient + asyncio.gather, Semaphore(20)not available0.558 s, 20 connections
Same 200 GETs over TLS, no semaphore (httpx caps in-flight work at 100), HTTP/1.1 vs HTTP/2HTTP/1.1 only0.476 s, 200 connections vs 0.171 s, 1 connection
50 MiB download, streamed, peak RSS35 MiB44 MiB
Retries against a 503urllib3.Retry(status_forcelist=[503]): 4 requests sentHTTPTransport(retries=3): 1 request (retries cover connection failures only)

Pick requests when the code is synchronous, the targets speak HTTP/1.1, and you want the larger ecosystem (plugins, mocks, examples). Pick httpx when you are inside asyncio (or an async framework), when you want HTTP/2, or when you want timeouts on by default. On a single reused connection the two are within a third of each other, and the one case where requests came out ahead was the loopback streaming test; nothing here says one is "faster".

BeautifulSoup Cheat Sheet (bs4 4.15): Tested Snippets, Parsers, Traps

· 20 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

BeautifulSoup Cheat Sheet (bs4 4.15): Tested Snippets, Parsers, Traps

Updated 2026-09-17

Rewritten as a tested cheat sheet. Every snippet below was run by the example package against one fixture page on beautifulsoup4 4.15.0, and the output is copied from that run. The 2024 version had no captured output (only inline # Output: comments), used the text= argument that has warned since 4.11, passed from_encoding to a str (ignored), and called pd.read_html with a string that pandas 3 now treats as a filename; all corrected. New: parsers compared on broken markup, a CPU-time table with min/median/max columns and a note on how much such timings move under load, the 4.13–4.15 deprecations captured with a type-checked example, and the errors you will actually see.

pip install beautifulsoup4 lxml # html5lib, requests and pandas only for the sections that use them
from bs4 import BeautifulSoup
with open("fixtures/page.html", encoding="utf-8") as f:
PAGE = f.read()

soup = BeautifulSoup(PAGE, "html.parser") # from a str
print(soup.title.string, "|", soup.h1.get_text(" ", strip=True))
Ant Supply — catalogue | Products (3)

That is the whole idea: parse, find, read. The sheet below is the rest of it, every line run against the same fixture page (a small catalogue with a nav, three product cards, a table, a comment and a script).

The best Python HTTP clients

· 8 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

The best Python HTTP clients

Python has emerged as a dominant language due to its simplicity and versatility. One crucial aspect of web development and scraping is making HTTP requests, and Python offers a rich ecosystem of libraries tailored for this purpose.

This report delves into the best Python HTTP clients, exploring their unique features and use cases. From the ubiquitous Requests library, known for its simplicity and ease of use, to the modern and asynchronous HTTPX, which supports the latest protocols like HTTP/2 and WebSockets, there is a tool for every need. Additionally, libraries like aiohttp offer versatile async capabilities, making them ideal for real-time data scraping tasks.

For those requiring low-level control, urllib3 stands out with its robust and flexible features. On the other hand, Uplink provides a declarative approach to API interactions, while GRequests combines the simplicity of Requests with the power of Gevent's asynchronous capabilities. This report also highlights best practices for making HTTP requests and provides a comprehensive guide to efficient web scraping using HTTPX and ScrapingAnt. By understanding the strengths and weaknesses of each library, developers can make informed decisions and choose the best tool for their web scraping and development tasks.

Python Requests: Ignore SSL Certificate Errors, and the Right Way

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Python Requests: Ignore SSL Certificate Errors, and the Right Way

Updated 2026-09-16

Rewritten as a tested reference. Every call shown below was executed by the example package against a local HTTPS server with a self-signed certificate; the output blocks are what it printed (see "How the outputs were captured"). Four recipes from the 2024 version were wrong and are shown failing here: passing an ssl.SSLContext as verify=, REQUESTS_CA_BUNDLE="", "verify=False disables verification globally", and the "certificate pinning" example built on the same invalid verify= call.

You called requests.get() and got this:

requests.exceptions.SSLError: HTTPSConnectionPool(host='localhost', port=8443): Max retries exceeded with url: / (Caused by SSLError(SSLCertVerificationError(1, '[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: self-signed certificate (_ssl.c:1010)')))

The two-line way to make it go away:

import requests, urllib3

urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
response = requests.get("https://localhost:8443/", verify=False)

And the one-line way to fix it instead of hiding it, when you have the server's certificate or your company's CA file:

response = requests.get("https://localhost:8443/", verify="certs/localhost.pem")

The rest of this page shows what each option does, with the real output, including the recipes you will find elsewhere that do nothing or crash in current requests.