NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

2 posts tagged with ".NET"

View All Tags

How to Parse HTML in C# with AngleSharp and Html Agility Pack

· 8 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to parse HTML in .NET

Updated 2026-10-08

Replaced incomplete snippets and broad speed claims with a runnable CSS/XPath extraction example, captured JSON and explicit failure controls. The evidence packet records the tested versions and limits.

To parse HTML in C#, use an HTML parser to build a document, select the elements you need, then convert their text and attributes into validated records. AngleSharp provides CSS selectors and a DOM-style API; Html Agility Pack provides XPath selection. The right starting point is usually the selector language your application already uses.

This tutorial extracts the same product records with both libraries. It covers relative URLs, encoded text, selectors that match nothing and records with missing fields. The example parses HTML already supplied to the application; fetching a page and executing its JavaScript are separate steps.

HTML Parsing Libraries - C#

· 5 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

HTML Parsing Libraries - C#

Web sites are written using HTML, which means that each web page is a structured document. Sometimes the goal is to obtain some data from them and preserve the structure while we’re at it. Websites don’t always provide their data in comfortable formats such as CSV or JSON, so only the way to deal with it is to parse the HTML page.