NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

259 posts tagged with "web scraping"

View All Tags

Web Scraping HTML Tables with JavaScript

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping HTML Tables with JavaScript

This article delves into the world of web scraping HTML tables using JavaScript, exploring both basic techniques and advanced practices to help developers efficiently collect and process tabular data from web pages.

JavaScript, with its robust ecosystem of libraries and tools, offers powerful capabilities for web scraping. By leveraging popular libraries such as Axios for HTTP requests and Cheerio for HTML parsing, developers can create efficient and reliable scrapers (Axios documentation, Cheerio documentation). Additionally, tools like Puppeteer and Playwright enable the handling of dynamic content, making it possible to scrape even the most complex, JavaScript-rendered tables (Puppeteer documentation).

In this comprehensive guide, we'll walk through the process of setting up a scraping environment, implementing basic scraping techniques, and exploring advanced methods for handling dynamic content and complex table structures. We'll also discuss crucial ethical considerations to ensure responsible and lawful scraping practices. By the end of this article, you'll have a solid foundation in web scraping HTML tables with JavaScript, equipped with the knowledge to tackle a wide range of scraping challenges.

Web Scraping HTML Tables with Python

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping HTML Tables with Python

Web scraping, particularly the extraction of data from HTML tables, offers a powerful means to gather information efficiently and at scale. As of 2024, Python remains a dominant language in this domain, offering a rich ecosystem of libraries and tools tailored for web scraping tasks.

This comprehensive guide delves into the intricacies of web scraping HTML tables using Python, providing both novice and experienced programmers with the knowledge and techniques needed to navigate this essential data collection method. We'll explore a variety of tools and libraries, each with its unique strengths and applications, enabling you to choose the most suitable approach for your specific scraping needs.

From the versatile BeautifulSoup library, known for its ease of use in parsing HTML documents (Beautiful Soup Documentation), to the powerful Pandas library that streamlines table extraction directly into DataFrame objects (Pandas Documentation), we'll cover the fundamental tools that form the backbone of many web scraping projects. For more complex scenarios involving dynamic content, we'll examine how Selenium can interact with web pages to access JavaScript-rendered tables (Selenium Documentation), and for large-scale projects, we'll introduce Scrapy, a comprehensive framework for building robust web crawlers (Scrapy Documentation).

Through a step-by-step approach, complete with code samples and detailed explanations, this guide aims to equip you with the skills to effectively extract, process, and analyze tabular data from the web. Whether you're looking to gather market research, monitor competitor pricing, or compile datasets for machine learning projects, mastering the art of web scraping HTML tables will undoubtedly enhance your data collection capabilities and open new avenues for insight and innovation.

ScrapeGraphAI Tutorial - Scraping Websites with LLMs

· 13 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

ScrapeGraphAI Tutorial - Scraping Websites with LLMs

Part 1 of this series discussed setting up and running local models with Ollama to extract data from complex local documents such as HTML and JSON. This part will focus on using API-based models for more efficient web scraping.

ScrapeGraphAI Tutorial - Getting Started with LLMs Web Scraping

· 9 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

ScrapeGraphAI Tutorial - Getting Started with LLMs Web Scraping

Imagine if you could describe the data you need in simple English, and AI takes care of the entire extraction and processing, whether from websites or local documents like PDFs, JSON, Markdown, and more. Even better, what if AI could summarize the data into an audio file or find the most relevant Google search results for your query—all at no cost or for just a few cents? This powerful functionality is provided by ScrapeGraphAI, an open-source AI-based Python scraper!

Selenium Cookies in Python: Set, Save and Validate Scraped Data

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Selenium Cookies in Python: Set, Save and Validate Scraped Data

Updated 2026-09-27

Replaced mixed-language API examples, cookie insertion on about:blank and untested session recipes with runnable extraction tests. The evidence includes Chrome and Firefox cookie restoration, incorrect-data cases, and a separate ScrapingAnt test that receives and resends synthetic cookies. The original URL and publication date are preserved.

To set a cookie in Selenium Python, first navigate to its target origin, then call driver.add_cookie() with a name/value dictionary. Navigate or refresh afterward so the next request uses it. To reuse a session, save the relevant cookies, restore them on the same origin in the new browser, and check the resulting data.

That last step matters. In our controlled catalog, every extraction returned HTTP 200 and four product rows. Restoring only the session cookie still selected member prices—but in USD when the intended dataset was EUR. A successful navigation, an authenticated-looking page and the expected row count all missed the error.

This guide shows the cookie operations, a complete runnable save/restore example, and the measured failures behind that distinction. The catalog and accounts are synthetic; the results demonstrate mechanisms, not production-site success rates.

Selenium Local Storage in Python: Set, Restore, Extract

· 11 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Selenium Local Storage in Python: Set, Restore, Extract

Updated 2026-09-27

Replaced broken advanced examples and unsupported performance claims with a runnable catalog experiment in Chrome and Firefox. The code, captures and result definitions distinguish changing storage from changing the data you extract.

Use Selenium Python's execute_script to access localStorage in the current document. Navigate to the intended origin first, pass keys and values as arguments, and include return when you need a value back in Python. If the application reads a preference at startup, set it before entering that application page or trigger its documented refresh afterward.

The last step matters for extraction. In our fixture, writing eu after the table loaded changed localStorage but left all four prices in USD. A wait that checked only the stored key would have accepted the wrong dataset. The examples below check storage, the application's request and the resulting records.

Playwright Local Storage: Set Before Load and Validate Data

· 10 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Playwright localStorage initialization and extracted data

Updated 2026-09-27

Replaced unsupported performance percentages and broken advanced recipes with runnable Chromium and Firefox experiments. The examples now distinguish startup state, current stored values and the records actually extracted.

To change localStorage on an already loaded origin, use page.evaluate(). If the application reads that value during startup, install an origin-guarded context.add_init_script() before navigation, or create the context with the required storage_state.

That timing matters when scraping. In our catalog fixture, setting the region to eu after the table had loaded changed the stored value but left four US-priced rows on screen. A reload produced the intended EUR records. A successful storage write alone did not validate the data.

Puppeteer Local Storage: Set, Save and Restore Before Load

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Puppeteer Local Storage: Set, Save and Restore Before Load

Updated 2026-09-27

Replaced broken advanced examples and unsupported performance claims with a tested Chrome packet covering startup timing, state restoration, origin errors and extracted records. See the captured results.

Use page.evaluate() to read or change localStorage in an already loaded Puppeteer page. To supply a value before the application's scripts run, install an origin-guarded page.evaluateOnNewDocument() callback before navigation. Pass values as arguments; a function executing in the page cannot implicitly use your Node.js variables or helpers.

Then check the data you meant to extract. Our catalog returned four rows with a captured HTTP 200 status in every extraction observation, yet 15 of the 36 observations contained the deliberately selected wrong dataset. A successful write to localStorage was not enough to update the page.

How to Set Cookies in Puppeteer

· 12 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Set Cookies in Puppeteer

In the realm of web automation and testing, Puppeteer has emerged as a powerful tool for developers and QA engineers. One crucial aspect of web interactions is the management of cookies, which play a vital role in maintaining user sessions, personalizing experiences, and handling authentication. This comprehensive guide delves into the intricacies of setting cookies in Puppeteer using JavaScript, exploring various methods and best practices to enhance your web automation projects.

Cookies are small pieces of data stored by websites on a user's browser, serving as a memory for web applications. In Puppeteer, manipulating these cookies programmatically allows for sophisticated automation scenarios, from maintaining login states to testing complex user flows. As web applications become increasingly complex, the ability to effectively manage cookies in automated environments has become a critical skill for developers.

This article will explore the fundamental methods for setting cookies in Puppeteer, including the versatile page.setCookie() function and the context-wide context.addCookies() method. We'll also delve into advanced techniques for cookie persistence, handling secure and HttpOnly cookies, and managing cookie expiration and deletion. Additionally, we'll cover best practices and advanced techniques that will elevate your cookie management skills, ensuring your Puppeteer scripts are robust, secure, and efficient.

By mastering these techniques, developers can create more reliable and sophisticated web automation solutions, capable of handling complex authentication flows, maintaining long-running sessions, and accurately simulating user interactions across various web applications. Whether you're building automated testing suites, web scrapers, or complex browser-based tools, understanding the nuances of cookie management in Puppeteer is essential for success in modern web development landscapes.

As we explore these topics, we'll provide detailed code samples and explanations, ensuring that both beginners and experienced developers can enhance their Puppeteer skills and create more powerful, efficient, and secure web automation solutions.

Looking for Playwright? Check out our guide on How to Set Cookies in Playwright.

Avoid Detection with Puppeteer Stealth

· 7 min read
Satyam Tripathi
Satyam is a junior data engineer and seasoned blogger. He has created several top-ranked tutorials on different topics like web scraping, automation, and scraping tools. He is always open to working with new technologies in the market and sharing his knowledge.

Avoid Detection with Puppeteer Stealth

Puppeteer is a powerful Node.js library that provides a high-level API for controlling browsers through the DevTools Protocol. It is commonly used for testing, web scraping, and automating repetitive browser tasks. However, Puppeteer's default settings can trigger bot detection systems, especially in headless mode.