NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

250 posts tagged with "data extraction"

View All Tags

Using Wget with Proxies

· 6 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Using Wget with Proxies

In today's interconnected digital landscape, wget stands as a powerful command-line utility for retrieving content from web servers. When combined with proxy capabilities, it becomes an even more versatile tool for secure and efficient web content retrieval.

This comprehensive guide explores the implementation, configuration, and optimization of wget when working with proxies. As organizations increasingly rely on proxy servers for enhanced security and access control (GNU Wget Manual), understanding the proper configuration and usage of wget with proxies has become crucial for system administrators and developers alike.

The integration of wget with proxy servers enables features such as anonymous browsing, geographic restriction bypass, and improved security measures. This research delves into various aspects of wget proxy implementation, from basic configuration to advanced authentication mechanisms, while also addressing critical performance optimization and troubleshooting strategies.

How to download images with wget

· 6 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with wget

wget stands as a powerful and versatile tool, particularly for retrieving images from websites. This comprehensive guide explores the intricacies of using wget for image downloads, a critical skill for system administrators, web developers, and digital content managers. Originally developed as part of the GNU Project (GNU Wget Manual), wget has evolved into an essential utility that combines robust functionality with flexible implementation options.

The tool's capability to handle recursive downloads, pattern matching, and authentication mechanisms makes it particularly valuable for bulk image retrieval tasks (Robots.net).

As websites become increasingly complex and security measures more sophisticated, understanding wget's advanced features and technical considerations becomes crucial for efficient and secure image downloading operations.

How to Send POST Requests With Wget: Form Data, JSON, Files, Redirects

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Send POST Requests With Wget: Form Data, JSON, Files, Redirects

Updated 2026-09-16

Rewritten as a tested reference. Every command below was run by the example package against a local echo server that reports the method, headers and body it received, and the output is copied from that run (wget 1.25.0). The 2024 version's redirect table, copied from the manual, was wrong for this wget: 301 and 302 turn the POST into a GET, only 307 and 308 keep it. New: --method with --body-data/--body-file, the missing-error-body trap and --content-on-error, --keep-session-cookies, --retry-on-http-error, percent-encoding, exit codes.

The three commands most people are looking for:

wget -qO- --post-data 'user=foo&lang=en' http://127.0.0.1:8000/echo # a form
wget -qO- --header='Content-Type: application/json' --post-data='{"key":"value"}' http://127.0.0.1:8000/echo # JSON
wget -qO- --header='Content-Type: application/json' --post-file=fixtures/data.json http://127.0.0.1:8000/echo # a body from a file

--post-data makes the request a POST, sets Content-Type: application/x-www-form-urlencoded unless you set your own with --header, and sends the string as the body. -qO- writes the response to stdout and silences the progress output. Everything else on this page is what happens around that: the response you did not get on an error, the redirect that silently changed your method, the login cookie that was not saved, the retry that did not retry.

Python Requests Cookies: Sessions, Scope and Scraped Data

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Python Requests Cookies: Sessions, Scope and Scraped Data

Updated 2026-09-27

Replaced untested persistence and recovery recipes with a runnable Requests cookie experiment. The new examples test redirect cookies, path collisions, prepared requests and session restoration against extracted records. Corrected unsupported cookie keyword arguments and the claim that mounting an HTTPS adapter forces HTTPS. The original URL, publication date and video are preserved.

Use a requests.Session() when later requests need cookies received earlier. Keep its cookie jar intact: a cookie is identified by more than its name, and a plain dictionary cannot represent two region cookies with different paths. For a single call, the cookies= parameter can supply cookies without making them persistent session state.

We tested those distinctions against a small catalog. A dictionary backup retained the member session but changed the selected currency from EUR to USD. The request returned HTTP 200 and all four expected product IDs. Only checking the currency and prices caught the wrong dataset.

This guide gives you a complete save/restore example and the failures behind it. The catalog, identities and prices are synthetic; the measurements demonstrate mechanisms, not success rates on real websites.

How to Send POST Requests With cURL

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Send POST Requests With cURL

In today's interconnected digital landscape, making HTTP POST requests has become a fundamental skill for developers and system administrators. cURL, a powerful command-line tool for transferring data, stands as one of the most versatile and widely-used utilities for making these requests.

According to recent statistics, JSON has emerged as the preferred format for over 70% of web APIs, making it crucial to understand how to effectively use cURL for POST operations.

This comprehensive guide explores the intricacies of sending POST requests with cURL, from basic syntax to advanced authentication methods. Whether you're testing APIs, uploading files, scraping the web or integrating with web services, understanding cURL's POST capabilities is essential for modern web development and system administration.

How to scrape a dynamic website with Puppeteer-Sharp

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to scrape a dynamic website with Puppeteer-Sharp

Scraping dynamic websites with Puppeteer-Sharp can be challenging for many developers. Puppeteer-Sharp, a .NET port of the Puppeteer library, enables effective browser automation in C#.

This article provides step-by-step guidance on using Puppeteer-Sharp to simplify data extraction from complex web pages. Enhance your web scraping skills now.

Web Scraping with VPN and Python

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping with VPN and Python

Web scraping with VPN integration has become an essential practice in modern data collection strategies, combining the need for efficient data gathering with robust privacy and security measures. As organizations increasingly rely on web-based data for business intelligence and research, the implementation of VPN-enabled scraping solutions has evolved into a sophisticated technical domain. According to ScrapingAnt's implementation guide, the integration of VPNs with web scraping not only provides enhanced anonymity but also enables more reliable and sustainable data collection operations. The combination of Python's powerful scraping libraries with VPN technology creates a robust framework for handling large-scale data extraction while maintaining privacy and avoiding IP-based restrictions. Proper VPN implementation in web scraping projects has become crucial for maintaining consistent access to target websites while ensuring compliance with various access policies and restrictions. This research explores the technical implementations, best practices, and advanced techniques necessary for successfully combining VPN services with Python-based web scraping operations.

Web Scraping with Tor and Python

· 6 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping with Tor and Python

Web scraping has become an essential tool for gathering information at scale. However, with increasing concerns about privacy and data collection restrictions, anonymous web scraping through the Tor network has emerged as a crucial methodology. This comprehensive guide explores the technical implementation and optimization of web scraping using Tor and Python, providing developers with the knowledge to build robust, anonymous data collection systems.

The integration of Tor with Python-based web scraping tools offers a powerful solution for maintaining anonymity while collecting data. Proper implementation of anonymous scraping techniques can significantly enhance privacy protection while maintaining efficient data collection capabilities. The combination of Tor's anonymity features with Python's versatile scraping libraries creates a framework that addresses both security concerns and performance requirements in modern web scraping applications.

How to Build a Web Scraper Using Playwright C#

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Build a Web Scraper Using Playwright C#

Web scraping has become an essential tool in modern data extraction and automation workflows. Playwright, Microsoft's powerful browser automation framework, has emerged as a leading solution for robust web scraping implementations in C#. This comprehensive guide explores the implementation of web scraping using Playwright, offering developers a thorough understanding of its capabilities and best practices.

Playwright stands out in the automation landscape by offering multi-browser support and superior performance compared to traditional tools like Selenium and Puppeteer (Playwright Documentation). According to recent benchmarks, Playwright demonstrates up to 40% faster execution times compared to Selenium, while providing more reliable wait mechanisms and better cross-browser compatibility.

The framework's modern architecture and sophisticated API make it particularly well-suited for handling dynamic content, complex JavaScript-heavy applications, and single-page applications (SPAs). With support for multiple browser engines including Chromium, Firefox, and WebKit, Playwright offers unparalleled flexibility in web scraping scenarios (Microsoft .NET Blog).

This guide will walk through the essential components of implementing web scraping with Playwright in C#, from initial setup to advanced techniques and performance optimization strategies. Whether you're building a simple data extraction tool or a complex web automation system, this comprehensive implementation guide will provide the knowledge and best practices necessary for successful deployment.

Web Scraping with Playwright Java: A Runnable Maven Example

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping with Playwright Java: A Runnable Maven Example

Correction (2026-09-30)

The earlier version contained disconnected Maven fragments, a subclass accessing a private field, an undefined robots helper, and an incorrect retry helper. Those were source-inspection findings, not compiler results from this experiment. This refresh replaces them with a tested project and removes unsupported hardware and performance claims.

A Java scraper needs to know when a page's data is complete, what a valid record looks like, and whether an empty result means success or failure. This walkthrough runs Playwright against an owned catalog whose products appear after JavaScript executes. It exports validated JSON and checks failure outcomes without contacting a customer site.

The complete Maven project includes the entry point, fixture server, scraper, build configuration and tests. Its captured run passed all six controlled scenarios. That is evidence about this fixture, not a success rate for arbitrary websites.