NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

How to Parse XML in C++: Tested Examples and Library Choices

· 13 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

How to Parse XML in C++

Updated 2026-10-07

Replaced placeholder code with compiled C++17 examples and captured failure cases, corrected the TinyXML-2 naming and documentation links, and removed unsupported speed comparisons and unrelated compiler theory. Download the tested source, fixtures, and outputs.

To parse XML in C++, use a library to read the document, check the parse result, and then validate the fields your application requires. A successful parse alone does not tell you whether a record has a title, whether an ID is an integer, or whether the document is safe to process.

This guide starts with a complete pugixml file parser that extracts book IDs and titles. A smaller TinyXML-2 example shows how to parse a string. The library comparison explains when to move to Xerces-C++, libxml2, or libxml++ for namespaces, validation, or streaming. No scraping API is needed to parse XML you already have.

Choose a C++ XML library for the document you have​

Start with the XML features and input size you need. This is a selection guide, not a performance ranking; benchmark your own documents if throughput or memory use determines the choice.

LibraryUseful whenConstraints to account for
pugixmlYou want a DOM tree, straightforward traversal, and XPath queries.No DTD or XML Schema validation; some XML well-formedness checks are omitted. Namespace queries use literal qualified names.
TinyXML-2You need a small DOM API for a configuration file or simple export.UTF-8 input; no DTD processing. TinyXML-2 is a separate library from the older TinyXML-1.
Xerces-C++You need namespace processing, DTD/XSD validation, or SAX callbacks.Configure external resource access explicitly; library initialization and object lifetime need care.
libxml2 / libxml++You need namespace-aware processing or a reader/SAX interface; libxml++ offers C++ wrappers.libxml2 has a C API. libxml++ has separate API series and dependencies: use documentation and package names for the installed series.
RapidXMLAn existing project needs its header-based, in-place DOM parser.Keep the mutable source buffer alive. Default parsing does not check matching closing tag names; review its conformance limits.

For the small, namespace-free fixture below, pugixml is a practical starting point. Use a standards-focused parser when strict XML acceptance or namespace identity is part of your contract.

Install dependencies and create the fixture​

You need a C++17 compiler and either pkg-config or CMake. These package installation commands are setup instructions; they were not executed for this test, which used existing macOS installations. Distribution package versions may differ from the tested versions listed at the end.

Ubuntu/Debian:

sudo apt-get update
sudo apt-get install g++ cmake pkg-config libpugixml-dev libtinyxml2-dev

macOS with Homebrew:

brew install cmake pkg-config pugixml tinyxml2

Check the library versions your build will use:

pkg-config --modversion pugixml tinyxml2

The test environment reported:

1.16
10.0.0

Save this as catalog.xml. The XML declaration precedes the root element, so code that takes the document's first child indiscriminately can select a declaration rather than an element.

<?xml version="1.0" encoding="UTF-8"?>
<catalog>
<book id="101"><title>C++ &amp; XML</title></book>
<book id="102"><title><![CDATA[Parsing <feeds>]]></title></book>
</catalog>

The first title contains an escaped ampersand; the second uses CDATA. Both become ordinary text when extracted. The sample contract requires at least one book, a positive integer id, and exactly one nonblank, text-only title per book.

Parse an XML file with pugixml​

Save the following complete program as parse_catalog.cpp:

#include <pugixml.hpp>
#include <charconv>
#include <fstream>
#include <iostream>
#include <stdexcept>
#include <string>
#include <vector>

struct Book {
int id;
std::string title;
};

int main(int argc, char* argv[]) {
if (argc != 2) {
std::cerr << "Usage: parse_catalog FILE.xml\n";
return 1;
}
try {
// Read at most 1 MiB + 1 byte, including for non-regular files.
constexpr std::size_t max_bytes = 1024 * 1024;
std::ifstream input(argv[1], std::ios::binary);
if (!input) throw std::runtime_error("Cannot open input file");
std::string xml(max_bytes + 1, '\0');
input.read(xml.data(), static_cast<std::streamsize>(xml.size()));
if (input.bad()) throw std::runtime_error("Cannot read input file");
const auto count = static_cast<std::size_t>(input.gcount());
if (count > max_bytes) throw std::runtime_error("Input exceeds 1 MiB limit");
xml.resize(count);

pugi::xml_document doc;
const auto result = doc.load_buffer(
xml.data(), xml.size(), pugi::parse_default | pugi::parse_doctype);
if (!result) {
std::cerr << "XML parse error: " << result.description()
<< " (offset " << result.offset << ")\n";
return 1;
}
for (auto node : doc.children()) {
if (node.type() == pugi::node_doctype)
throw std::runtime_error("DOCTYPE is not allowed");
}
const auto root = doc.document_element();
if (std::string(root.name()) != "catalog")
throw std::runtime_error("Expected <catalog> root");
for (auto node = root.next_sibling(); node; node = node.next_sibling()) {
if (node.type() == pugi::node_element)
throw std::runtime_error("Expected one root element");
}

std::vector<Book> books;
for (auto book : root.children("book")) {
const std::string id_text = book.attribute("id").value();
int id = 0;
const auto conversion = std::from_chars(
id_text.data(), id_text.data() + id_text.size(), id);
if (conversion.ec != std::errc{} ||
conversion.ptr != id_text.data() + id_text.size() || id <= 0)
throw std::runtime_error("book/@id must be a positive integer");

const auto title_node = book.child("title");
if (!title_node || title_node.next_sibling("title"))
throw std::runtime_error("book requires exactly one <title>");
std::string title;
for (auto node : title_node.children()) {
if (node.type() != pugi::node_pcdata &&
node.type() != pugi::node_cdata)
throw std::runtime_error("title must contain text only");
title += node.value();
}
if (title.find_first_not_of(" \t\r\n") == std::string::npos)
throw std::runtime_error("book/title must not be blank");
books.push_back({id, title});
}
if (books.empty()) throw std::runtime_error("catalog contains no books");
// Emit only after every selected record passes validation.
for (const auto& book : books)
std::cout << book.id << " | " << book.title << '\n';
} catch (const std::exception& error) {
std::cerr << "Input error: " << error.what() << '\n';
return 1;
}
}

Compile and run it from the directory containing the program and fixture:

c++ -std=c++17 -Wall -Wextra -Wpedantic parse_catalog.cpp $(pkg-config --cflags --libs pugixml) -o parse_catalog
./parse_catalog catalog.xml

Captured output:

101 | C++ & XML
102 | Parsing <feeds>

document_element() selects the root element; children("book") visits its direct book children. The program checks the attribute and title before accepting each record. std::from_chars plus the end-pointer check rejects values such as 101oops and integers outside the range of int. It stores validated records before printing, so a bad later record does not leave partial success output.

The 1 MiB input cap is an example application policy, not a library limit or a recommended size for every workload. load_buffer() copies the bounded input into the document. Keep nodes and pointers into the document within its lifetime. See the pugixml loading and error reference.

Check malformed XML and missing fields separately​

Save these two additional files:

malformed.xml:

<catalog><book id="101"><title>Broken</book></catalog>

missing-title.xml:

<catalog><book id="101"/></catalog>

Run each command separately:

./parse_catalog malformed.xml

Captured stderr, exit status 1:

XML parse error: Start-end tags mismatch (offset 39)
./parse_catalog missing-title.xml

Captured stderr, exit status 1:

Input error: book requires exactly one <title>

The missing-title file parses successfully but fails the application's field checks. Treat those as separate errors. Do not continue extracting from a partially parsed document after a parse failure. pugixml's reported offset is not a line number and can refer to the converted buffer when encodings differ; do not assume it is an original-file byte position in every case.

The downloadable test runner also checks a missing file, empty input, blank titles, missing/invalid/overflowing IDs, wrong or multiple roots, duplicate titles, nested markup in titles, an empty catalog, a DOCTYPE, oversized input, and an invalid record after a valid one. Every failing catalog case produced no stdout.

Parse an XML string with TinyXML-2​

For a simple string, TinyXML-2 exposes Parse(), RootElement(), and FirstChildElement(). Check every pointer before calling GetText(). Save this as tiny_title.cpp:

#include <tinyxml2.h>
#include <iostream>
#include <string>

int main(int argc, char* argv[]) {
if (argc != 2) {
std::cerr << "Usage: tiny_title 'XML string'\n";
return 1;
}
tinyxml2::XMLDocument doc;
if (doc.Parse(argv[1]) != tinyxml2::XML_SUCCESS) {
std::cerr << "XML parse error at line " << doc.ErrorLineNum()
<< ": " << doc.ErrorStr() << '\n';
return 1;
}
const auto* root = doc.RootElement();
if (!root || std::string(root->Name()) != "book") {
std::cerr << "Expected <book> root\n";
return 1;
}
const auto* title = root->FirstChildElement("title");
if (!title || !title->GetText() ||
std::string(title->GetText()).find_first_not_of(" \t\r\n") ==
std::string::npos) {
std::cerr << "Missing or blank book/title\n";
return 1;
}
std::cout << title->GetText() << '\n';
}

Compile and run:

c++ -std=c++17 -Wall -Wextra -Wpedantic tiny_title.cpp $(pkg-config --cflags --libs tinyxml2) -o tiny_title
./tiny_title '<book><title>C++ &amp; XML</title></book>'

Captured output:

C++ & XML

For <book/>, the program prints Missing or blank book/title and exits with status 1. A mismatched closing tag produces a parse error with a line number. Both cases were executed in the test runner.

This smaller example returns text only when the first title's first child is a text node. It does not implement the catalog program's input cap, duplicate-title checks, or text-only policy. Use it for the shown simple string shape; extend the checks before accepting arbitrary records. For files, TinyXML-2 provides LoadFile(), whose result must also be checked. The TinyXML-2 documentation describes its UTF-8 input and error APIs.

Build both programs with CMake​

Put both .cpp files and this CMakeLists.txt in one directory:

cmake_minimum_required(VERSION 3.16)
project(xml_examples LANGUAGES CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(CMAKE_CXX_EXTENSIONS OFF)
find_package(pugixml CONFIG REQUIRED)
find_package(tinyxml2 CONFIG REQUIRED)
add_executable(parse_catalog parse_catalog.cpp)
target_link_libraries(parse_catalog PRIVATE pugixml::pugixml)
add_executable(tiny_title tiny_title.cpp)
target_link_libraries(tiny_title PRIVATE tinyxml2::tinyxml2)

With the libraries installed in a standard prefix:

cmake -S . -B build
cmake --build build
./build/parse_catalog catalog.xml
./build/tiny_title '<book><title>C++ &amp; XML</title></book>'

For Homebrew, supply its prefix when configuring:

cmake -S . -B build -DCMAKE_PREFIX_PATH="$(brew --prefix)"

Both targets were built with CMake as well as the direct compiler commands. If CMake cannot find a package, point CMAKE_PREFIX_PATH at its installation prefix and confirm it contains the package's CMake configuration files. A header include alone does not link the library.

Handle untrusted XML deliberately​

An XML parser can become an external resource loader when DTD or entity processing is enabled. An attacker-controlled reference can cause local file access or outbound requests. Entity expansion, deep nesting, huge text nodes, and compressed inputs can also exhaust resources.

Use an explicit policy for XML from feeds, uploads, or remote responses:

  • Limit bytes before parsing, including the decompressed size. Bound processing time and memory separately; the sample's file-size cap does not impose a read deadline or a nesting-depth limit.
  • Reject DTDs if your format does not require them. pugixml does not expand DTD-declared entities; the catalog program additionally retains and rejects DOCTYPE nodes. That does not make it a fully conforming XML validator. See its conformance notes.
  • With libxml2, do not enable XML_PARSE_NOENT, external DTD loading, or XInclude casually. XML_PARSE_NOENT enables entity substitution despite its name. XML_PARSE_NO_XXE, available since libxml2 2.13.0, disables external DTD/entity loading. XML_PARSE_NONET alone is not a complete external-resource policy, and XML_PARSE_HUGE relaxes resource limits. Check the options and resource-loader behavior of your installed version in the libxml2 parser reference.
  • With Xerces-C++, review validation mode, setLoadExternalDTD(false), and setDisableDefaultEntityResolution(true). External DTD loading can still occur in validating modes, so use an explicit resolver policy for any schemas or resources you intentionally permit. See the Xerces DOM configuration guide.

For security-sensitive ingestion, choose a parser that enforces the XML rules you require and run it within resource limits. Missing-field checks and a byte cap address specific failures; they are not a complete hostile-input sandbox.

DOM, streaming, namespaces, and validation​

DOM versus streaming: pugixml and TinyXML-2 build a tree in memory, which makes repeated traversal and editing convenient. For a large export that you read once, consider SAX callbacks or a pull reader from Xerces-C++, libxml2, or libxml++. Retain only the current record where possible. Streaming still needs limits on text size, depth, and any records your application accumulates; it does not guarantee constant memory on its own.

Namespaces: an XML prefix is an alias for a namespace URI, and producers can change it without changing the element's identity. The correct namespace-aware match uses the URI and local name. This fixture has no namespaces. Literal calls such as child("book") do not establish namespace identity; do not strip prefixes and assume correctness. Use namespace-aware APIs for feeds or protocols that require them. See Namespaces in XML.

Validation: parsing checks syntax to the extent implemented by the library. DTD/XSD validation checks a declared grammar; application validation checks required fields and business rules. These are separate steps. The catalog program checks selected fields, ignores unselected elements/attributes, and does not verify unique IDs or an entire schema. Its parsing options discard comments and processing instructions; the title check rejects child elements and collects the retained text and CDATA. Add further rules if your input contract requires them. Restrict schema loading to trusted resources rather than following arbitrary locations supplied by the input.

RapidXML in existing code: its default parse<0>() does not validate closing tag names. Enable parse_validate_closing_tags when you need that check, catch rapidxml::parse_error, and retain the mutable, zero-terminated input buffer for as long as nodes reference it. This advice is from the RapidXML manual; no RapidXML, Xerces-C++, or libxml++ program was executed for this refresh.

If your source is HTML, use an HTML parser: recovery rules for web pages differ from XML rules. See the separate C++ HTML parsing guide. If you are working in Python instead, see XML parsing in Python.

Examples tested on 2026-10-07 with pugixml 1.16, TinyXML-2 10.0.0, Apple Clang 21.0.0, C++17, and CMake 3.28.1 on macOS 26.6.2. All 24 test cases passed. Source, fixtures, execution limits, and captured outputs. Package installation and Linux/Windows execution were not tested.

This refresh was prepared with AI assistance; the C++ examples and failure cases were executed in the environment described above.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster