NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

71 posts tagged with "python"

View All Tags

Guide to Scraping and Storing Data to MongoDB Using Python

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Guide to Scraping and Storing Data to MongoDB Using Python

Data is a critical asset, and the ability to efficiently extract and store it is a valuable skill. Web scraping, the process of extracting data from websites, is a fundamental technique for data scientists, analysts, and developers. Python, with its powerful libraries such as BeautifulSoup and Scrapy, provides a robust environment for web scraping. MongoDB, a NoSQL database, complements this process by offering a flexible and scalable solution for storing the scraped data. This comprehensive guide will walk you through the steps of scraping web data using Python and storing it in MongoDB, leveraging the capabilities of BeautifulSoup, Scrapy, and PyMongo. Understanding these tools is not only essential for data extraction but also for efficiently managing and analyzing large datasets. This guide is designed to be SEO-friendly and includes detailed explanations and code samples to help you seamlessly integrate web scraping and data storage into your projects. (source, source, source, source, source)

Guide to Cleaning Scraped Data and Storing it in PostgreSQL Using Python

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Guide to Cleaning Scraped Data and Storing it in PostgreSQL Using Python

In today's data-driven world, the ability to efficiently clean and store data is paramount for any data scientist or developer. Scraped data, often messy and inconsistent, requires meticulous cleaning before it can be effectively used for analysis or storage. Python, with its robust libraries such as Pandas, NumPy, and BeautifulSoup4, offers a powerful toolkit for data cleaning. PostgreSQL, a highly efficient open-source database, is an ideal choice for storing this cleaned data. This research report provides a comprehensive guide on setting up a Python environment for data cleaning, connecting to a PostgreSQL database, and ensuring data integrity through various cleaning techniques. With detailed code samples and explanations, this guide is designed to be both practical and SEO-friendly, helping readers navigate the complexities of data preprocessing and storage with ease (Python Official Website, Anaconda, GeeksforGeeks).

Crawlee for Python Tutorial with Examples

· 8 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Crawlee for Python Tutorial with Examples

Web scraping has become an essential tool for data extraction in various industries, from market analysis to academic research. One of the most effective libraries for Python available today is Crawlee, which provides a robust framework for both simple and complex web scraping tasks. Crawlee supports various scraping scenarios, including dealing with static web pages using BeautifulSoup and handling JavaScript-rendered content with Playwright. In this tutorial, we will delve into how to set up and effectively use Crawlee for Python, providing clear examples and best practices to ensure efficient and scalable web scraping operations. This comprehensive guide aims to equip you with the knowledge to build your own web scrapers, whether you are just getting started or looking to implement advanced features. For more detailed documentation, you can visit the Crawlee Documentation and the Crawlee PyPI.

How to Read HTML Tables With Pandas read_html() (pandas 3, Tested)

· 19 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Read HTML Tables With Pandas read_html() (pandas 3, Tested)

Updated 2026-09-17

Re-tested on pandas 3.0.5, where pd.read_html("<table>…") with a raw string no longer works. Every option on this page now runs against one fixture page with eleven tables in the evidence packet, with the DataFrame output shown. The old page's df.append, chunksize= and usecols= recipes, which do not run on pandas 2.0 or later, are shown failing and replaced.

pandas.read_html() finds every <table> in a page and returns a list of DataFrames. It takes a file path, a URL, or a file-like object (StringIO around HTML you already have; a raw string no longer works on pandas 3). The three-line version, on a file:

import pandas as pd
PAGE = "fixtures/tables.html"
tables = pd.read_html(PAGE)
print("file:", len(tables), "tables; shapes:", [df.shape for df in tables])
print(tables[0])
file: 12 tables; shapes: [(3, 3), (3, 4), (2, 5), (3, 3), (2, 3), (2, 3), (2, 2), (2, 2), (5, 2), (2, 2), (3, 2), (3, 3)]
Name Age City
0 Ada 36 London
1 Grace 45 Arlington
2 Linus 28 Helsinki

The same call on HTML you already hold, and on a URL:

pd.read_html(StringIO(html)): 12
pd.read_html(open(PAGE, 'rb')): 12
pd.read_html(URL): 12

The rest of this page is what happens after that line: how to get the one table you want, why headers, numbers and links come out the way they do, which parser is running, why a Wikipedia URL returns 403, and which idioms from older tutorials (including the 2024 version of this page) raise on pandas 3.

How to Parse XML in Python: ElementTree, Namespaces, Large Files, lxml

· 20 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Parse XML in Python: ElementTree, Namespaces, Large Files, lxml

Updated 2026-09-16

Rewritten as a tested reference. Every example was executed by the example package against small fixture files and one generated 23.8 MB file; the output blocks are copied from that run. The 2024 version's performance numbers were not measured and are replaced by a measured table; its xml.etree.cElementTree recommendation (a deprecated alias since Python 3.3; xml.etree.ElementTree already uses the C accelerator), its codecs.open encoding advice (wrong; the equivalent open(..., encoding="utf-8") is shown failing below) and a stray "Meta Description" section are gone; namespaces, large-file streaming, writing, xmltodict, pandas.read_xml and the current security picture are new.

The standard library does this without installing anything. Given a file books.xml with a <catalog> of <book> elements:

import xml.etree.ElementTree as ET

root = ET.parse("fixtures/books.xml").getroot()
for book in root.findall("book"):
print(book.get("id"), book.find("title").text, book.findtext("price", default="n/a"))
bk101 XML Developer's Guide 44.95
bk102 Midnight Rain 5.95
bk103 Maeve Ascendant n/a

That covers most "read an XML file" tasks. The rest of this page is the things that go wrong next, each with the real output: elements you cannot find because of a namespace, a file too big for parse(), encoding and syntax errors, and when to reach for lxml, xmltodict or pandas instead.

Download Files with Selenium in Python: Chrome and Firefox

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Download Files with Selenium in Python: Chrome and Firefox

Updated 2026-09-21

Re-tested Chrome and Firefox downloads with Selenium 4.49.0 in headed and headless modes. Replaced legacy APIs and fixed sleeps with an explicit file-verification helper, and removed the broken JavaScript and Safe Browsing workaround. The runnable evidence packet includes the fixtures, captured output and failure cases.

To download a file with Selenium in Python, configure the browser's download directory before starting the session, click the link, and wait for the expected file to pass a content check before closing the browser. Waiting for the link to be clickable and waiting for the file to finish are separate steps.

Below are tested Chrome and Firefox configurations, including headless runs. The examples download a CSV, a delayed binary file and a PDF from a local fixture server. They verify the expected length and SHA-256 digest rather than trusting that a filename appeared.

How to Find Elements With Selenium in Python

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Find Elements With Selenium in Python

Updated 2026-09-28

Replaced the earlier locator examples with a runnable local catalog and captured Chrome/Firefox results. Corrected singular/plural lookup, class-token matching, stale-reference recovery and shadow-root access; removed unsupported locator speed rankings.

Use driver.find_element(By.CSS_SELECTOR, selector) for the first matching element and driver.find_elements(By.CSS_SELECTOR, selector) for a list of all matches. If nothing matches, the singular method raises NoSuchElementException; the plural method returns an empty list.

For scraping, finding elements is only the first step. You also need the right container, populated fields and a way to recover when the page replaces nodes. The examples below extract a complete catalog and deliberately test selectors that return plausible but wrong results.

How to download images with Python?

· 18 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to download images with Python?

Downloading images using Python is an essential skill for various applications, including web scraping, data analysis, and machine learning. This comprehensive guide explores the top Python libraries for image downloading, advanced techniques, and best practices for ethical and efficient image scraping. Whether you're a beginner or an experienced developer, understanding the nuances of these tools and techniques can significantly enhance your projects. Popular libraries like Requests, Urllib3, Wget, PyCURL, and Aiohttp each offer unique features suited for different scenarios. For instance, Requests is known for its simplicity and user-friendly API, making it a favorite among developers for straightforward tasks. On the other hand, advanced users may prefer Urllib3 for its robust connection pooling and SSL verification capabilities. Additionally, leveraging asynchronous libraries like Aiohttp can optimize large-scale, concurrent downloads, which is crucial for high-performance scraping tasks. Beyond the basics, advanced techniques such as using Selenium for dynamic content, handling complex image sources, and implementing parallel downloads can further refine your scraping strategy. Ethical considerations, including compliance with copyright laws and website terms of service, are also paramount to ensure responsible scraping practices. This guide aims to provide a holistic view of Python image downloading, equipping you with the knowledge to handle various challenges effectively.

This article is a part of the series on image downloading with different programming languages. Check out the other articles in the series:

Handling Scrapy Failure URLs - A Comprehensive Guide

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Handling Scrapy Failure URLs - A Comprehensive Guide

Web scraping is an increasingly essential tool in data collection and analysis, enabling businesses and researchers to gather vast amounts of information from the web efficiently. Among the numerous frameworks available for web scraping, Scrapy stands out due to its robustness and flexibility. However, the process of web scraping is not without its challenges, especially when dealing with failures that can halt or disrupt scraping tasks. From network failures to HTTP errors and parsing issues, understanding how to handle these failures is crucial for maintaining the reliability and efficiency of your scraping projects. This guide delves into the common types of failures encountered in Scrapy and provides practical solutions to manage them effectively, ensuring that your scraping tasks remain smooth and uninterrupted. For those looking to deepen their web scraping skills, this comprehensive guide will equip you with the knowledge to handle failures adeptly, backed by detailed explanations and code examples. For more detailed information, you can visit the Scrapy documentation.

How to Create a Proxy Server in Python Using Proxy.py

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Create a Proxy Server in Python Using Proxy.py

Updated 2026-09-22

Re-tested HTTP forwarding, HTTPS CONNECT, authentication and plugins with proxy.py 2.4.10. Replaced the broken rotation recipe with a tested upstream pool. HTTPS forwarding does not require TLS interception, and local proxy ports do not establish different public exit IPs. Commands, fixtures and captured logs.

To create a local Python proxy server, start proxy.py on an explicit loopback address and point your client at it. Then verify both the response and the proxy log: receiving a page is not enough to prove which route the client used.

This guide builds that path with local HTTP and HTTPS origins, cURL and Requests clients, authentication, a header plugin and an upstream pool. You can run the complete lab without an external target, a ScrapingAnt account or an API key.