NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

267 posts tagged with "web scraping"

View All Tags

How to scrape dynamic websites with Scrapy Splash

· 8 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to scrape dynamic websites with Scrapy Splash

Handling dynamic websites with JavaScript-rendered content presents a significant challenge for traditional scraping tools. Scrapy Splash emerges as a powerful solution by combining the robust crawling capabilities of Scrapy with the JavaScript rendering prowess of the Splash headless browser. This comprehensive guide explores the integration and optimization of Scrapy Splash for effective dynamic website scraping.

Scrapy Splash has become an essential tool for developers and data scientists who need to extract data from JavaScript-heavy websites. The middleware (scrapy-plugins/scrapy-splash) seamlessly bridges Scrapy's asynchronous architecture with Splash's rendering engine, enabling the handling of complex web applications. This integration provides a robust foundation for handling modern web applications while maintaining high performance and reliability.

The system's architecture is specifically designed to handle the challenges of dynamic content rendering while ensuring efficient resource utilization.

Using Cookies with Wget

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Using Cookies with Wget

GNU Wget stands as a powerful command-line utility that has become increasingly essential for managing web interactions. This comprehensive guide explores the intricate aspects of using cookies with Wget, a crucial feature for maintaining session states and handling authenticated requests.

Cookie management in Wget has evolved significantly, offering robust mechanisms for both basic and advanced implementations (GNU Wget Manual). The ability to handle cookies effectively is particularly vital when dealing with modern web applications that rely heavily on session management and user authentication.

Recent developments in browser integration capabilities have further enhanced Wget's cookie handling capabilities, allowing seamless interaction with existing browser sessions. This research delves into the various aspects of cookie implementation in Wget, from basic session management to advanced security considerations, providing a thorough understanding of both theoretical concepts and practical applications.

Using Cookies with cURL

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Using Cookies with cURL

Managing cookies effectively is crucial for maintaining state and handling user sessions. cURL, a powerful command-line tool for transferring data, provides robust cookie handling capabilities that have become essential for developers and system administrators.

This comprehensive guide explores the intricacies of using cookies with cURL, from basic operations to advanced security implementations. According to (curl.se), cURL adopts the Netscape cookie file format, providing a standardized approach to cookie management that ensures compatibility across different platforms and use cases.

The tool's cookie handling capabilities have evolved significantly, incorporating security features and compliance with modern web standards (everything.curl.dev). As web applications become increasingly complex, understanding how to effectively manage cookies with cURL has become paramount for secure and efficient data transfer operations.

Using Wget with Proxies

· 6 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Using Wget with Proxies

In today's interconnected digital landscape, wget stands as a powerful command-line utility for retrieving content from web servers. When combined with proxy capabilities, it becomes an even more versatile tool for secure and efficient web content retrieval.

This comprehensive guide explores the implementation, configuration, and optimization of wget when working with proxies. As organizations increasingly rely on proxy servers for enhanced security and access control (GNU Wget Manual), understanding the proper configuration and usage of wget with proxies has become crucial for system administrators and developers alike.

The integration of wget with proxy servers enables features such as anonymous browsing, geographic restriction bypass, and improved security measures. This research delves into various aspects of wget proxy implementation, from basic configuration to advanced authentication mechanisms, while also addressing critical performance optimization and troubleshooting strategies.

How to download images with wget

· 6 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with wget

wget stands as a powerful and versatile tool, particularly for retrieving images from websites. This comprehensive guide explores the intricacies of using wget for image downloads, a critical skill for system administrators, web developers, and digital content managers. Originally developed as part of the GNU Project (GNU Wget Manual), wget has evolved into an essential utility that combines robust functionality with flexible implementation options.

The tool's capability to handle recursive downloads, pattern matching, and authentication mechanisms makes it particularly valuable for bulk image retrieval tasks (Robots.net).

As websites become increasingly complex and security measures more sophisticated, understanding wget's advanced features and technical considerations becomes crucial for efficient and secure image downloading operations.

How to Send POST Requests With Wget: Form Data, JSON, Files, Redirects

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Send POST Requests With Wget: Form Data, JSON, Files, Redirects

Updated 2026-09-16

Rewritten as a tested reference. Every command below was run by the example package against a local echo server that reports the method, headers and body it received, and the output is copied from that run (wget 1.25.0). The 2024 version's redirect table, copied from the manual, was wrong for this wget: 301 and 302 turn the POST into a GET, only 307 and 308 keep it. New: --method with --body-data/--body-file, the missing-error-body trap and --content-on-error, --keep-session-cookies, --retry-on-http-error, percent-encoding, exit codes.

The three commands most people are looking for:

wget -qO- --post-data 'user=foo&lang=en' http://127.0.0.1:8000/echo # a form
wget -qO- --header='Content-Type: application/json' --post-data='{"key":"value"}' http://127.0.0.1:8000/echo # JSON
wget -qO- --header='Content-Type: application/json' --post-file=fixtures/data.json http://127.0.0.1:8000/echo # a body from a file

--post-data makes the request a POST, sets Content-Type: application/x-www-form-urlencoded unless you set your own with --header, and sends the string as the body. -qO- writes the response to stdout and silences the progress output. Everything else on this page is what happens around that: the response you did not get on an error, the redirect that silently changed your method, the login cookie that was not saved, the retry that did not retry.

Python Requests Cookies: Sessions, Scope and Scraped Data

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Python Requests Cookies: Sessions, Scope and Scraped Data

Updated 2026-09-27

Replaced untested persistence and recovery recipes with a runnable Requests cookie experiment. The new examples test redirect cookies, path collisions, prepared requests and session restoration against extracted records. Corrected unsupported cookie keyword arguments and the claim that mounting an HTTPS adapter forces HTTPS. The original URL, publication date and video are preserved.

Use a requests.Session() when later requests need cookies received earlier. Keep its cookie jar intact: a cookie is identified by more than its name, and a plain dictionary cannot represent two region cookies with different paths. For a single call, the cookies= parameter can supply cookies without making them persistent session state.

We tested those distinctions against a small catalog. A dictionary backup retained the member session but changed the selected currency from EUR to USD. The request returned HTTP 200 and all four expected product IDs. Only checking the currency and prices caught the wrong dataset.

This guide gives you a complete save/restore example and the failures behind it. The catalog, identities and prices are synthetic; the measurements demonstrate mechanisms, not success rates on real websites.

How to Send POST Requests With cURL

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Send POST Requests With cURL

In today's interconnected digital landscape, making HTTP POST requests has become a fundamental skill for developers and system administrators. cURL, a powerful command-line tool for transferring data, stands as one of the most versatile and widely-used utilities for making these requests.

According to recent statistics, JSON has emerged as the preferred format for over 70% of web APIs, making it crucial to understand how to effectively use cURL for POST operations.

This comprehensive guide explores the intricacies of sending POST requests with cURL, from basic syntax to advanced authentication methods. Whether you're testing APIs, uploading files, scraping the web or integrating with web services, understanding cURL's POST capabilities is essential for modern web development and system administration.