NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

245 posts tagged with "data extraction"

View All Tags

How to read from MongoDB to Pandas

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to read from MongoDB to Pandas

The ability to efficiently read and manipulate data is crucial for effective data analysis and application development. MongoDB, a leading NoSQL database, is renowned for its flexibility and scalability, making it a popular choice for modern applications. However, to leverage the full potential of MongoDB data for analysis, it is essential to seamlessly integrate it with powerful data manipulation tools like Pandas in Python.

This comprehensive guide delves into the various methods of reading data from MongoDB into Pandas DataFrames, providing a detailed roadmap for developers and data analysts. We will explore the use of PyMongo, the official MongoDB driver for Python, which allows for straightforward interactions with MongoDB. Additionally, we will discuss PyMongoArrow, a tool designed for efficient data transfer between MongoDB and Pandas, offering significant performance improvements. For handling large datasets, we will cover chunking techniques and the use of MongoDB's Aggregation Framework to preprocess data before loading it into Pandas.

Guide to Scraping and Storing Data to MongoDB Using Python

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Guide to Scraping and Storing Data to MongoDB Using Python

Data is a critical asset, and the ability to efficiently extract and store it is a valuable skill. Web scraping, the process of extracting data from websites, is a fundamental technique for data scientists, analysts, and developers. Python, with its powerful libraries such as BeautifulSoup and Scrapy, provides a robust environment for web scraping. MongoDB, a NoSQL database, complements this process by offering a flexible and scalable solution for storing the scraped data. This comprehensive guide will walk you through the steps of scraping web data using Python and storing it in MongoDB, leveraging the capabilities of BeautifulSoup, Scrapy, and PyMongo. Understanding these tools is not only essential for data extraction but also for efficiently managing and analyzing large datasets. This guide is designed to be SEO-friendly and includes detailed explanations and code samples to help you seamlessly integrate web scraping and data storage into your projects. (source, source, source, source, source)

Guide to Cleaning Scraped Data and Storing it in PostgreSQL Using Python

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Guide to Cleaning Scraped Data and Storing it in PostgreSQL Using Python

In today's data-driven world, the ability to efficiently clean and store data is paramount for any data scientist or developer. Scraped data, often messy and inconsistent, requires meticulous cleaning before it can be effectively used for analysis or storage. Python, with its robust libraries such as Pandas, NumPy, and BeautifulSoup4, offers a powerful toolkit for data cleaning. PostgreSQL, a highly efficient open-source database, is an ideal choice for storing this cleaned data. This research report provides a comprehensive guide on setting up a Python environment for data cleaning, connecting to a PostgreSQL database, and ensuring data integrity through various cleaning techniques. With detailed code samples and explanations, this guide is designed to be both practical and SEO-friendly, helping readers navigate the complexities of data preprocessing and storage with ease (Python Official Website, Anaconda, GeeksforGeeks).

Crawlee for Python Tutorial with Examples

· 8 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Crawlee for Python Tutorial with Examples

Web scraping has become an essential tool for data extraction in various industries, from market analysis to academic research. One of the most effective libraries for Python available today is Crawlee, which provides a robust framework for both simple and complex web scraping tasks. Crawlee supports various scraping scenarios, including dealing with static web pages using BeautifulSoup and handling JavaScript-rendered content with Playwright. In this tutorial, we will delve into how to set up and effectively use Crawlee for Python, providing clear examples and best practices to ensure efficient and scalable web scraping operations. This comprehensive guide aims to equip you with the knowledge to build your own web scrapers, whether you are just getting started or looking to implement advanced features. For more detailed documentation, you can visit the Crawlee Documentation and the Crawlee PyPI.

How to Read HTML Tables With Pandas read_html() (pandas 3, Tested)

· 19 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Read HTML Tables With Pandas read_html() (pandas 3, Tested)

Updated 2026-09-17

Re-tested on pandas 3.0.5, where pd.read_html("<table>…") with a raw string no longer works. Every option on this page now runs against one fixture page with eleven tables in the evidence packet, with the DataFrame output shown. The old page's df.append, chunksize= and usecols= recipes, which do not run on pandas 2.0 or later, are shown failing and replaced.

pandas.read_html() finds every <table> in a page and returns a list of DataFrames. It takes a file path, a URL, or a file-like object (StringIO around HTML you already have; a raw string no longer works on pandas 3). The three-line version, on a file:

import pandas as pd
PAGE = "fixtures/tables.html"
tables = pd.read_html(PAGE)
print("file:", len(tables), "tables; shapes:", [df.shape for df in tables])
print(tables[0])
file: 12 tables; shapes: [(3, 3), (3, 4), (2, 5), (3, 3), (2, 3), (2, 3), (2, 2), (2, 2), (5, 2), (2, 2), (3, 2), (3, 3)]
Name Age City
0 Ada 36 London
1 Grace 45 Arlington
2 Linus 28 Helsinki

The same call on HTML you already hold, and on a URL:

pd.read_html(StringIO(html)): 12
pd.read_html(open(PAGE, 'rb')): 12
pd.read_html(URL): 12

The rest of this page is what happens after that line: how to get the one table you want, why headers, numbers and links come out the way they do, which parser is running, why a Wikipedia URL returns 403, and which idioms from older tutorials (including the 2024 version of this page) raise on pandas 3.

How to Parse XML in C++

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Parse XML in C++

Parsing XML in C++ is a critical skill for developers who need to handle structured data efficiently and accurately. XML, or eXtensible Markup Language, is a versatile format for data representation and interchange, widely used in web services, configuration files, and data exchange protocols. Parsing XML involves reading XML documents and converting them into a usable format for further processing. C++ developers have a variety of XML parsing libraries at their disposal, each with its own strengths and trade-offs. This guide will explore popular XML parsing libraries for C++, including Xerces-C++, RapidXML, PugiXML, TinyXML, and libxml++, and provide insights into different parsing techniques such as top-down and bottom-up parsing. Understanding these tools and techniques is essential for building robust and efficient applications that require XML data processing. For more information on XML parsing, you can refer to Apache Xerces-C++, RapidXML, PugiXML, TinyXML, and libxml++.

Web Scraping with Haskell - A Comprehensive Tutorial

· 14 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping with Haskell - A Comprehensive Tutorial

Web scraping has become an essential tool for data extraction from websites, enabling developers to gather information for various applications such as market research, competitive analysis, and content aggregation. Haskell, a statically-typed, functional programming language, offers a robust ecosystem for web scraping through its strong type system, concurrency capabilities, and extensive libraries. This guide aims to provide a comprehensive overview of web scraping with Haskell, covering everything from setting up the development environment to leveraging advanced techniques for efficient and reliable scraping.

How to Parse HTML in C++

· 15 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Parse HTML in C++

HTML parsing is a fundamental process in web development and data extraction. It involves breaking down HTML documents into their constituent elements, allowing for easy manipulation and analysis of the structure and content. In the context of C++, HTML parsing can be particularly advantageous due to the language's high performance and low-level control. However, the process also presents challenges, such as handling nested elements, malformed HTML, and varying HTML versions.

This comprehensive guide aims to provide an in-depth exploration of HTML parsing in C++. It covers essential concepts such as tokenization, tree construction, and DOM (Document Object Model) representation, along with practical code examples. We will delve into various parsing techniques, discuss performance considerations, and highlight best practices for robust error handling. Furthermore, we will review some of the most popular HTML parsing libraries available for C++, including Gumbo Parser, libxml++, Boost.Beast, MyHTML, and TinyXML-2, to help developers choose the best tool for their specific needs.

How to Parse XML in Python: ElementTree, Namespaces, Large Files, lxml

· 20 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Parse XML in Python: ElementTree, Namespaces, Large Files, lxml

Updated 2026-09-16

Rewritten as a tested reference. Every example was executed by the example package against small fixture files and one generated 23.8 MB file; the output blocks are copied from that run. The 2024 version's performance numbers were not measured and are replaced by a measured table; its xml.etree.cElementTree recommendation (a deprecated alias since Python 3.3; xml.etree.ElementTree already uses the C accelerator), its codecs.open encoding advice (wrong; the equivalent open(..., encoding="utf-8") is shown failing below) and a stray "Meta Description" section are gone; namespaces, large-file streaming, writing, xmltodict, pandas.read_xml and the current security picture are new.

The standard library does this without installing anything. Given a file books.xml with a <catalog> of <book> elements:

import xml.etree.ElementTree as ET

root = ET.parse("fixtures/books.xml").getroot()
for book in root.findall("book"):
print(book.get("id"), book.find("title").text, book.findtext("price", default="n/a"))
bk101 XML Developer's Guide 44.95
bk102 Midnight Rain 5.95
bk103 Maeve Ascendant n/a

That covers most "read an XML file" tasks. The rest of this page is the things that go wrong next, each with the real output: elements you cannot find because of a namespace, a file too big for parse(), encoding and syntax errors, and when to reach for lxml, xmltodict or pandas instead.

How to Ignore SSL Certificate With cURL

· 21 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Ignore SSL Certificate With cURL

In today's digital landscape, securing internet communications is paramount, and SSL/TLS certificates play a crucial role in this process. SSL (Secure Sockets Layer) and its successor TLS (Transport Layer Security) are cryptographic protocols designed to ensure data privacy, authentication, and trust between web servers and browsers. SSL/TLS certificates, issued by Certificate Authorities (CAs), authenticate a website's identity and enable encrypted connections. This authentication process is similar to issuing passports, wherein the CA verifies the entity's identity before issuing the certificate.

However, there are scenarios, especially during development and testing, where developers might need to bypass these SSL checks. This is where cURL, a command-line tool for transferring data using various protocols, comes into play. cURL provides options to handle SSL certificate validation, allowing developers to ignore SSL checks temporarily. While this practice can be invaluable in non-production environments, it also comes with significant security risks. Ignoring SSL certificate checks can expose systems to man-in-the-middle attacks, phishing, and data integrity compromises. Therefore, it's essential to understand both the methods and the implications of bypassing SSL checks with cURL.