NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

265 posts tagged with "web scraping"

View All Tags

Data to the Rescue. The Role of Data Collection in the Russia-Ukraine War

· 10 min read
Rakia Ben Sassi
Senior Software Engineer and Content Creator

Data to the Rescue. The Role of Data Collection in the Russia-Ukraine War

As I have started writing this article, I didn’t expect it to end this way. Weeks after creating my first draft, Russian forces entered its western neighbor’s border and war raged in Ukraine.

Many questions have been raised. People around the world kept their eyes glued to their screens, waiting for more news about the invasion and looking for answers. I was no exception. I’ve seen the steady stream of content, talking about the different sides of the crisis, its ramifications, and its ripple effect.

How Data Collection Can Improve HR Processes

· 11 min read
ScrapingAnt Team
ScrapingAnt

Data Collection for e-Commerce

HR or Human Resources is a department as important as any other in a business or a corporation. It helps manage the workforce so that workers are happy and a healthy environment is created, which helps achieve the organization's targets.

In a world where we have employed computers, AI, and the internet to better everything, why should HR lag behind? After all, if the employees are loyal and happy where they work, they are more likely to give their all while doing their jobs, all of which ultimately leads to growth. In order to do that, many big companies have made use of a newer approach, using public web data in order to improve human resource processes.

Rule eCommerce with Data Collection

· 8 min read
ScrapingAnt Team
ScrapingAnt

Data Collection for e-Commerce

A flood of information and data runs on the web. In this global age, people use the internet to achieve almost everything. Everything they click on, things they search for, and the websites they spend the most time on translates to user behavior and tells us about what they like to see. Such information is invaluable and is waiting right there in front of us.

Collection of such data, ensuring proper data processing with the help of data collection can help your eCommerce business grow in ways unimaginable.

How to avoid IP rate limits

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to avoid IP rate limiting?

Web scraping specialists are dealing with using proxy servers to overcome various anti-bot defenses every day. One of those protections is IP rate limiting, a primary anti-scraping mechanism.

Let's learn more about this protection method and the most effective ways of bypassing it.

Best Free Proxy Scraping Tools

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Best open source proxy scrapers

Using a quality proxy server is the key to a successful web scraper. A variety of IPs along with their quality make it possible to collect data from various web sites without worrying about being blocked.

Still, many websites provide free proxy lists, so can the process of getting IP addresses from them be automated? Are free proxies good enough for web scraping? Let's check it out.

How to Get All Text from a Webpage with Puppeteer

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Get All Text from a Webpage with Puppeteer

Correction (2026-09-28)

The earlier article described DOM Range selection as Ctrl+A/copy-paste and claimed it could recover text from sites using display optimizations. The example performed neither keyboard input nor clipboard copying, and the general recovery claim was unsupported. This refresh compares actual text returned by five methods on a controlled fixture. Code and captured outputs.

For the readable text of an already rendered page, start with page.$eval('body', body => body.innerText). Use textContent when you explicitly want descendant text that can include hidden content and script/style source. If the HTML already contains the data, a converter such as html-to-text can work without a browser.

Those methods produce different outputs. Below, the browser sees three dynamically populated products, while the raw HTTP response contains empty product placeholders. The examples show how that difference affects text extraction, and when to keep structured records instead of flattening the page.

How to download images with NodeJS?

· 5 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with NodeJS?

Working with images in NodeJS extends your web scraping capabilities, from downloading the image with an URL to retrieving photo attributes like EXIF. How to achieve the image download and obtain the data?

This article is a part of the series on image downloading with different programming languages. Check out the other articles in the series: