How to Parse HTML in C# with AngleSharp and Html Agility Pack

Replaced incomplete snippets and broad speed claims with a runnable CSS/XPath extraction example, captured JSON and explicit failure controls. The evidence packet records the tested versions and limits.
To parse HTML in C#, use an HTML parser to build a document, select the elements you need, then convert their text and attributes into validated records. AngleSharp provides CSS selectors and a DOM-style API; Html Agility Pack provides XPath selection. The right starting point is usually the selector language your application already uses.
This tutorial extracts the same product records with both libraries. It covers relative URLs, encoded text, selectors that match nothing and records with missing fields. The example parses HTML already supplied to the application; fetching a page and executing its JavaScript are separate steps.
Choose a parser for the task
| Task | Starting point in this example |
|---|---|
Select elements with CSS, such as article.product a.name | AngleSharp |
| Select elements with XPath expressions | Html Agility Pack |
| Parse HTML that is already on disk or in a response string | Either implementation below |
| Obtain fields created by JavaScript after page load | Retrieve rendered HTML first, then parse it |
AngleSharp's upstream documentation describes its HTML parser and DOM/CSS selector APIs. Html Agility Pack's repository documents its HTML DOM and XPath support. These examples test AngleSharp 1.8.4, HtmlAgilityPack 1.13.0 and .NET SDK 10.0.401; they do not compare parsing speed.
Install and run the complete example
Install the .NET SDK, plus Git and Bash. The packet's global.json requests SDK 10.0.401 and permits a servicing patch in that feature band. Its project and NuGet lock file pin the parser packages. NuGet restore needs network access; the extraction and tests use local files and no credentials.
git clone https://github.com/ScrapingAnt/scrapingant-examples.git
cd scrapingant-examples
git checkout db55e3c5aacbcebba99e2be2b62364f9e3635022
cd examples/parse-html-dot-net
./run.sh
The script builds the project, runs the assertions, executes both parsers and compares their JSON with the saved capture. To inspect one implementation directly:
dotnet run --configuration Release
dotnet run --configuration Release -- --hap
The HTML fixture and expected records
The owned fixture deliberately includes an unrelated productivity class and encoded text. It has no external dependencies:
<!doctype html><html><body>
<main id="catalog">
<article class="featured product" data-id="tea"><a class="name" href="/items/tea?pack=1&size=2">Tea <em>&</em> biscuits</a></article>
<article class="productivity" data-id="decoy"><a class="name" href="/wrong">Wrong record</a></article>
<article class="product" data-id="cafe"><a class="name" href="../items/cafe">Café "Noir"</a></article>
<article class="product sale" data-id="literal"><a class="name" href="https://cdn.example.test/items/literal">Literal &lt;tag&gt;</a></article>
</main></body></html>
Use https://shop.example.test/catalog/index.html as the page URL when resolving links. It is a fixture identifier, not a fetched website. Both implementations print this captured JSON:
[
{
"Id": "tea",
"Title": "Tea \u0026 biscuits",
"Url": "https://shop.example.test/items/tea?pack=1\u0026size=2"
},
{
"Id": "cafe",
"Title": "Caf\u00E9 \u0022Noir\u0022",
"Url": "https://shop.example.test/items/cafe"
},
{
"Id": "literal",
"Title": "Literal \u0026lt;tag\u0026gt;",
"Url": "https://cdn.example.test/items/literal"
}
]
The JSON serializer uses Unicode escape sequences for some characters; parsing the JSON recovers the original characters. The third title deliberately contains the literal text <tag>. Decoding that text again would change the data.
Extract records with CSS or XPath
The following is the tested Extractors.cs file. Both methods return the same Product[] shape. The Create helper rejects incomplete records and resolves only HTTP or HTTPS links.
using System.Net;
using AngleSharp.Html.Parser;
using HtmlAgilityPack;
public sealed record Product(string Id, string Title, string Url);
public static class Extractors
{
public static Product[] WithCss(string html, Uri pageUrl)
{
using var document = new HtmlParser().ParseDocument(html);
return document.QuerySelectorAll("#catalog article.product")
.Select(card =>
{
var link = card.QuerySelector("a.name");
// DOM TextContent/GetAttribute already contain decoded entities.
return Create(card.GetAttribute("data-id"), link?.TextContent,
link?.GetAttribute("href"), pageUrl);
}).ToArray();
}
public static Product[] WithXPath(string html, Uri pageUrl)
{
var document = new HtmlDocument();
document.LoadHtml(html);
var cards = document.DocumentNode.SelectNodes(
"//*[@id='catalog']//article[contains(concat(' ', normalize-space(@class), ' '), ' product ')]");
if (cards is null) return [];
return cards.Select(card =>
{
var link = card.SelectSingleNode(
".//a[contains(concat(' ', normalize-space(@class), ' '), ' name ')]");
// Decode once for HAP, including attributes; never decode DOM text twice.
return Create(WebUtility.HtmlDecode(card.GetAttributeValue("data-id", null)),
WebUtility.HtmlDecode(link?.InnerText),
WebUtility.HtmlDecode(link?.GetAttributeValue("href", null)), pageUrl);
}).ToArray();
}
private static Product Create(string? id, string? title, string? href, Uri pageUrl)
{
if (string.IsNullOrWhiteSpace(id) || string.IsNullOrWhiteSpace(title) ||
string.IsNullOrWhiteSpace(href))
throw new InvalidDataException("A product requires data-id, title and href.");
if (!Uri.TryCreate(pageUrl, href, out var url) ||
(url.Scheme != Uri.UriSchemeHttp && url.Scheme != Uri.UriSchemeHttps))
throw new InvalidDataException("Product href must resolve to HTTP or HTTPS.");
return new Product(id.Trim(), title.Trim(), url.AbsoluteUri);
}
}
For AngleSharp, QuerySelectorAll("#catalog article.product") selects the cards. QuerySelector("a.name") then searches inside each card. TextContent includes nested text; DOM text and attributes are already entity-decoded, so there is no second decoding pass.
For Html Agility Pack, the XPath expression matches a class token, not an arbitrary substring. contains(@class, 'product') would also match the fixture's productivity decoy. SelectNodes can return null when no nodes match, which this method converts to an empty array. The HAP path decodes text and attributes once before projection.
The normal extraction path in Program.cs reads the file, chooses the method and serializes the records; the repository also contains a separate test entry point:
using System.Text.Json;
var html = await File.ReadAllTextAsync("fixtures/catalog.html");
var pageUrl = new Uri("https://shop.example.test/catalog/index.html");
var records = args.Contains("--hap")
? Extractors.WithXPath(html, pageUrl)
: Extractors.WithCss(html, pageUrl);
Console.WriteLine(JsonSerializer.Serialize(records, new JsonSerializerOptions { WriteIndented = true }));
Replace the selectors and required fields with your own document's contract. Keep record validation close to extraction so that a changed selector does not quietly produce partial data.
Handle missing data and URLs deliberately
| Condition | Tested behavior | What to do in your application |
|---|---|---|
| No cards match | Empty array | Decide whether an empty catalog is valid or indicates the wrong input/selector |
| A selected card lacks ID, title or href | InvalidDataException | Investigate the document or update the schema/selector |
| Link is relative | Resolve against the supplied page URL | Use the final response URL if redirects changed the page location |
| Link resolves to a non-HTTP scheme | InvalidDataException | Keep a deliberate allowlist for links your next step will fetch |
| A script contains markup it would insert into the page | No inserted record | Fetch rendered HTML when the required field is absent from the original response |
The dated run passed 15 assertions, including literal record equality, the no-match and invalid-record paths, script non-execution and a wrong-order oracle control. That is evidence for this fixture and these conditions, not a percentage success rate on arbitrary websites.
This sample uses the explicit page URL and does not apply an HTML <base href> element. If your input uses one, resolve and validate its effective base before applying relative links. It also does not claim that the parsers repair every malformed document identically.
Raw HTML, rendered HTML and the next step
Before changing parsers, check whether the required field exists in the HTML string you actually received. A CSS or XPath selector cannot recover an element that JavaScript has not yet inserted. For the raw-response versus rendered-DOM distinction, see scraping dynamic websites with Python; the language differs, but the input boundary is the same.
ScrapingAnt is not needed when you already have the HTML. If obtaining rendered HTML is the missing step, consult the ScrapingAnt request/response documentation before building that retrieval path. No ScrapingAnt request was executed for this parser packet.
For a table-shaped task in Python, see pandas read_html. For XML input, use an XML-specific approach, such as the separate C++ XML parsing guide, rather than assuming HTML parsing rules apply.
What happened to the old benchmark?
The original public benchmark project remains available as historical context. Its old timings are not a current library comparison or a prediction for another machine, input or extraction task. This refresh chooses selectors and validates output before considering speed. Measure your own representative workload if parsing time is a material bottleneck.
Examples tested on 2026-10-08 with .NET SDK 10.0.401, AngleSharp 1.8.4 and HtmlAgilityPack 1.13.0. Code and captured outputs.
This article was prepared with AI assistance from executed examples and independently reviewed. Oleg Kulyk is responsible for corrections.