NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

Web Scraping Blog — Page 16

How to Ignore SSL Certificate With cURL

· 21 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Ignore SSL Certificate With cURL

In today's digital landscape, securing internet communications is paramount, and SSL/TLS certificates play a crucial role in this process. SSL (Secure Sockets Layer) and its successor TLS (Transport Layer Security) are cryptographic protocols designed to ensure data privacy, authentication, and trust between web servers and browsers. SSL/TLS certificates, issued by Certificate Authorities (CAs), authenticate a website's identity and enable encrypted connections. This authentication process is similar to issuing passports, wherein the CA verifies the entity's identity before issuing the certificate.

However, there are scenarios, especially during development and testing, where developers might need to bypass these SSL checks. This is where cURL, a command-line tool for transferring data using various protocols, comes into play. cURL provides options to handle SSL certificate validation, allowing developers to ignore SSL checks temporarily. While this practice can be invaluable in non-production environments, it also comes with significant security risks. Ignoring SSL certificate checks can expose systems to man-in-the-middle attacks, phishing, and data integrity compromises. Therefore, it's essential to understand both the methods and the implications of bypassing SSL checks with cURL.

Download Files with Selenium in Python: Chrome and Firefox

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Download Files with Selenium in Python: Chrome and Firefox

Updated 2026-09-21

Re-tested Chrome and Firefox downloads with Selenium 4.49.0 in headed and headless modes. Replaced legacy APIs and fixed sleeps with an explicit file-verification helper, and removed the broken JavaScript and Safe Browsing workaround. The runnable evidence packet includes the fixtures, captured output and failure cases.

To download a file with Selenium in Python, configure the browser's download directory before starting the session, click the link, and wait for the expected file to pass a content check before closing the browser. Waiting for the link to be clickable and waiting for the file to finish are separate steps.

Below are tested Chrome and Firefox configurations, including headless runs. The examples download a CSV, a delayed binary file and a PDF from a local fixture server. They verify the expected length and SHA-256 digest rather than trusting that a filename appeared.

Wget vs cURL for Downloading Files in Linux

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Wget vs cURL for Downloading Files in Linux

In the realm of Linux-based environments, downloading files from the internet is a common task that can be accomplished using a variety of tools. Among these tools, Wget and cURL stand out as the most popular and widely used. Both tools offer robust capabilities for downloading files, but they cater to slightly different use cases and have unique strengths and weaknesses. Understanding these differences is crucial for selecting the right tool for specific tasks, whether you are downloading a single file, mirroring a website, or interacting with complex APIs.

Wget, short for 'World Wide Web get', is designed primarily for downloading files and mirroring websites. Its straightforward syntax and default behaviors make it user-friendly for quick, one-off downloads. For example, the command wget [URL] will download the file from the specified URL and save it to the current directory. Wget excels in tasks like recursive downloads and website mirroring, making it a preferred choice for archiving websites or downloading entire directories of files.

cURL, short for 'Client URL', is a versatile tool that supports a wide array of protocols beyond HTTP and HTTPS. It can be used for various network operations, including FTP, SCP, SFTP, and more. cURL requires additional options for saving files, such as curl -O [URL], but offers extensive customization options for HTTP headers, methods, and data. This makes cURL particularly useful for API interactions and complex web requests.

This comprehensive guide aims to provide a detailed comparison of Wget and cURL, covering their basic file download capabilities, protocol support, recursive download features, resume mechanisms, and advanced HTTP request handling. By the end of this guide, you will have a clear understanding of which tool is best suited for your specific needs.

How to Find Elements With Selenium in Python

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Find Elements With Selenium in Python

Updated 2026-09-28

Replaced the earlier locator examples with a runnable local catalog and captured Chrome/Firefox results. Corrected singular/plural lookup, class-token matching, stale-reference recovery and shadow-root access; removed unsupported locator speed rankings.

Use driver.find_element(By.CSS_SELECTOR, selector) for the first matching element and driver.find_elements(By.CSS_SELECTOR, selector) for a list of all matches. If nothing matches, the singular method raises NoSuchElementException; the plural method returns an empty list.

For scraping, finding elements is only the first step. You also need the right container, populated fields and a way to recover when the page replaces nodes. The examples below extract a complete catalog and deliberately test selectors that return plausible but wrong results.

How to submit a form with Puppeteer?

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to submit a form with Puppeteer?

Puppeteer, a Node.js library developed by Google, offers a high-level API to control headless Chrome or Chromium browsers, making it an indispensable tool for web scraping, automated testing, and form submission automation. In today's digital landscape, automating form submissions is crucial for a variety of applications, ranging from data collection to user interaction testing. Puppeteer provides a robust solution for these tasks, allowing developers to programmatically interact with web pages as if they were using a regular browser. This guide delves into the setup and advanced techniques for using Puppeteer to automate form submissions, ensuring reliable and efficient automation processes. By following the outlined steps, users can install and configure Puppeteer, create basic scripts, handle dynamic form elements, manage complex inputs, and integrate with testing frameworks like Jest. Additionally, this guide explores effective strategies for bypassing CAPTCHAs and anti-bot measures, which are common obstacles in web automation.

Looking for a Playwright guide? Check out: How to submit a form with Playwright?

How to download images with cURL?

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with cURL?

The ability to efficiently download images from the internet is not just a convenience but a necessity for developers, system administrators, and many other professionals. cURL, a robust command-line tool, provides a versatile and powerful solution for this task. Whether you are looking to perform basic image downloads, handle complex redirects, or manage multiple simultaneous transfers, cURL has the capabilities to meet your needs. This comprehensive guide delves into both fundamental and advanced image downloading techniques using cURL, offering insights into handling redirects, managing authentication, optimizing large image transfers, and ensuring secure file storage. By mastering these techniques, users can significantly enhance their image retrieval processes, making them faster, more secure, and more efficient. The following sections will provide detailed explanations, code samples, and best practices drawn from authoritative sources to help you leverage cURL to its fullest potential.

This article is a part of the series on image downloading with different programming languages. Check out the other articles in the series:

How to download images with Java?

· 16 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with Java?

In the current digital age, the ability to download and process images efficiently is an essential skill for Java developers. Whether it's for a simple application or a complex system, understanding the various methods available for image downloading can significantly enhance performance and functionality. This comprehensive guide explores five key methods for downloading images in Java, utilizing built-in libraries, third-party libraries, and advanced techniques (Oracle Java Documentation). Each method is detailed with step-by-step explanations and code samples, making it suitable for both beginners and experienced developers. Additionally, we delve into performance optimization, reliability, memory management, security considerations, and the best libraries for efficient image downloading. By understanding these concepts, developers can create robust and efficient image downloading solutions tailored to their specific needs.

This article is a part of the series on image downloading with different programming languages. Check out the other articles in the series:

How to Download Images in C# with HttpClient

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with C#?

Updated 2026-10-08

Replaced repeated snippets and unsupported timing comparisons with a complete HttpClient program and owned PNG/JPEG failure tests. Corrected legacy API guidance and added body-size limits, body cancellation and temporary-file cleanup. Runnable evidence.

To download an image from a URL in C#, use HttpClient to request the bytes and save them to a file. For a bounded download, inspect the status and content type, stream the body with a byte limit and cancellation token, and publish the destination only after the transfer passes your checks.

The program below accepts PNG and JPEG responses. It checks their declared media type and leading signature bytes, preserves an existing destination and removes temporary bytes when a transfer fails. Those checks are useful format filters; they do not fully decode an image or prove that arbitrary content is safe.

How to download images with Go?

· 14 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with Go?

Downloading images programmatically is a vital task in many applications, ranging from web scrapers to automated backups. The Go programming language, with its powerful standard library and rich ecosystem of third-party packages, offers efficient tools for accomplishing this task. This guide explores how to use Go's net/http package for downloading images, handling different formats, and implementing best practices for error handling and concurrency. Additionally, it delves into enhanced image downloading with third-party packages, providing detailed explanations and step-by-step instructions for leveraging popular Go libraries like go-getter and grab to improve efficiency. These libraries, combined with image processing packages such as imaging and bild, enable developers to create robust and high-performance image downloading systems. By integrating AI-powered tools like Gigapixel AI and AVCLabs Photo Enhancer API, you can further enhance image quality and processing capabilities. This comprehensive guide covers everything from basic image downloading to advanced techniques, ensuring that your applications are both efficient and secure.

This article is a part of the series on image downloading with different programming languages. Check out the other articles in the series:

How to Configure Proxies in Laravel and Symfony for PHP Clients

· 16 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Configure Proxies in Laravel and Symfony for PHP Clients

Proxy configurations are a fundamental aspect of web development, serving multiple essential purposes such as enhancing security, optimizing performance, and overcoming network restrictions. Both Laravel and Symfony, two of the most popular PHP frameworks, offer robust methods for integrating proxy settings into their HTTP clients. Understanding how to set up proxies in these frameworks is crucial for developers aiming to build secure and efficient web applications. This report delves into the step-by-step processes for configuring proxies in Laravel and Symfony, providing detailed explanations and practical code samples. By following the guidelines and best practices outlined here, developers can ensure their applications are both resilient and performant. Laravel's HTTP client, built on Guzzle, offers various ways to configure proxies, including global settings via environment variables and route-specific settings using middleware (Laravel HTTP Client Documentation). Similarly, Symfony's HTTP client, which leverages PHP's native cURL extension, provides flexible proxy configurations that can be tailored to different environments and authentication requirements (Symfony HTTP Client Documentation).