NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

250 posts tagged with "data extraction"

View All Tags

How to avoid IP rate limits

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to avoid IP rate limiting?

Web scraping specialists are dealing with using proxy servers to overcome various anti-bot defenses every day. One of those protections is IP rate limiting, a primary anti-scraping mechanism.

Let's learn more about this protection method and the most effective ways of bypassing it.

Best Free Proxy Scraping Tools

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Best open source proxy scrapers

Using a quality proxy server is the key to a successful web scraper. A variety of IPs along with their quality make it possible to collect data from various web sites without worrying about being blocked.

Still, many websites provide free proxy lists, so can the process of getting IP addresses from them be automated? Are free proxies good enough for web scraping? Let's check it out.

How to Get All Text from a Webpage with Puppeteer

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Get All Text from a Webpage with Puppeteer

Correction (2026-09-28)

The earlier article described DOM Range selection as Ctrl+A/copy-paste and claimed it could recover text from sites using display optimizations. The example performed neither keyboard input nor clipboard copying, and the general recovery claim was unsupported. This refresh compares actual text returned by five methods on a controlled fixture. Code and captured outputs.

For the readable text of an already rendered page, start with page.$eval('body', body => body.innerText). Use textContent when you explicitly want descendant text that can include hidden content and script/style source. If the HTML already contains the data, a converter such as html-to-text can work without a browser.

Those methods produce different outputs. Below, the browser sees three dynamically populated products, while the raw HTTP response contains empty product placeholders. The examples show how that difference affects text extraction, and when to keep structured records instead of flattening the page.

How to download images with NodeJS?

· 5 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with NodeJS?

Working with images in NodeJS extends your web scraping capabilities, from downloading the image with an URL to retrieving photo attributes like EXIF. How to achieve the image download and obtain the data?

This article is a part of the series on image downloading with different programming languages. Check out the other articles in the series:

How to parse HTML in .NET

· 8 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to parse HTML in .NET

HTML parsing is a vital part of web scraping, as it allows convert web page content to meaningful and structured data. Still, as HTML is a tree-structured format, it requires a proper tool for parsing, as it can't be property traversed using Regex.

This article will reveal the most popular .NET libraries for HTML parsing with their strong and weak parts.

Web Scraping with Java

· 16 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping with Java

Java is one of the most popular and high demanded programming languages nowadays. It allows creating highly-scalable and reliable services as well as multi-threaded data extraction solutions. Let's check out the main concepts of web scraping with Java and review the most popular libraries to setup your data extraction flow.

How to download a file with Playwright?

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download a file with Playwright?

In this article, we will share several ideas on how to download files with Playwright. Automating file downloads can sometimes be confusing. You need to handle a download location, download multiple files simultaneously, support streaming, and even more. Unfortunately, not all the cases are well documented. Let's go through several examples and take a deep dive into Playwright's APIs used for file download.

This guide is a part of the series on web scraping and file downloading with different web drivers and programming languages. Check out the other articles in the series: