Build an AI Scraper with MCP and Validate Its Extracted Records

Replaced the conceptual agent pipeline with five captured MCP retrievals, 15 real Claude extractions, and an offline validator. The evidence packet preserves the raw results, including a formatting failure and quarantined records. The September 29 correction removed unsupported reliability and vendor-comparison claims; this refresh supplies an executed implementation.
An MCP tool can return page content without proving that a model extracted the correct price. A JSON Schema can accept a record that contains the wrong value. Keep retrieval, extraction and acceptance as separate steps, and retain the evidence at each boundary.
This walkthrough retrieves an owned catalog through ScrapingAnt MCP, sends the saved Markdown to a pinned Claude model, and checks its records against a schema, source excerpts and an independent fixture oracle. You can replay the complete validation offline before using either API.
What was actually run
The packet uses Python 3.12.11, MCP SDK 2.2.0, httpx2 2.13.1 for MCP transport, httpx 0.28.1 for the model API, and jsonschema 4.25.1. The dependency lock is included. ScrapingAnt discovery negotiated protocol 2025-11-25 and advertised server version 1.30.0 in the saved capture.
Five small fixtures exercise different conditions:
| Fixture | What changes |
|---|---|
| Complete | Two products with explicit IDs, titles, prices and currencies |
| Changed layout | Different wrappers and classes around the same product statements |
| Missing price | One product has no available price |
| Conflicting value | One product has two contradictory price/currency statements |
| Untrusted instruction | Page text asks the extractor to fabricate a product |
Each fixture was fetched once with get_web_page_markdown, browser=false and proxy_type=datacenter. The URLs point to immutable HTML files in the owned examples repository. These are static pages; this experiment does not test JavaScript rendering. The complete and changed-layout HTML fixtures produced identical Markdown, so this case checks the retrieval representation and does not demonstrate model-side adaptation to a changed representation.
The model stage used claude-haiku-4-5-20251001, temperature 0 and a 2,000-output-token limit, with three calls per saved page. The model received the prompt, application schema, source URL and retrieved text. It received neither the oracle nor tools for further browsing. There was no extraction retry or model-based repair loop.
Start with the saved evidence
Open the packet at its tested commit, or check it out locally:
git clone https://github.com/ScrapingAnt/scrapingant-examples.git
cd scrapingant-examples
git checkout 178e5e1847939c8419bf8b348c8fadc5586af338
cd examples/building-ai-driven-scrapers-in-2025-agents-mcp-and
python3.12 -m venv .venv
. .venv/bin/activate
python -m pip install -r requirements.lock.txt
./run.sh
The script runs 27 offline checks and replays the saved Anthropic responses. It makes no provider calls, even if API keys are configured. Inspect expected_output/anthropic-2026-09-30/validation.json, accepted.jsonl and quarantine.jsonl for the per-run results.
The offline tests include malformed JSON, wrong types, missing and extra keys, duplicate IDs, invented values, incorrect source URLs, unsupported quotes, omitted products and truncated model output. Those tests inject failures; they are not additional live provider observations.
The first failure: the response was not raw JSON
Every Claude response enclosed its JSON in a Markdown code fence despite the prompt requesting only JSON. Passing the untouched message to the strict JSON parser therefore failed in 15 of 15 runs.
We added a framing adapter after observing that failure. It accepts raw JSON or exactly one complete json code fence. It removes only the enclosing fence before parsing; it does not extract an object from surrounding prose, combine multiple blocks, fill missing fields or rewrite values. Incomplete fences and truncated API responses still fail. The raw provider responses remain unchanged, and each validation result records whether decoding occurred.
This is a post-hoc parser change, not evidence that the original prompt produced directly usable JSON. The results below are explicitly after fence decoding. No additional model calls were made to obtain them.
Validate structure, then source support
The application schema requires six fields: id, title, price_minor, currency, source_url and evidence. It rejects additional properties. The validator requires actual integers for available prices rather than coercing strings or floating-point values. Price and currency may be null so an incomplete record can be represented, but that does not authorize storing it as a complete product.
A captured accepted record looks like this:
{
"id": "AA101",
"title": "Desk Lamp",
"price_minor": 3499,
"currency": "USD",
"source_url": "https://raw.githubusercontent.com/ScrapingAnt/scrapingant-examples/dbf3261c7d3960558ba52fa200f8136e9da28bb2/fixtures/mcp-catalog/complete.html",
"evidence": "AA101 | Desk Lamp | Price: 34.99 USD (3499 minor units)."
}
The validator applies separate checks:
- Parse the decoded JSON and enforce the exact object/record structure.
- Reject duplicate or unknown product IDs and compare typed values with the independently authored oracle.
- Require the acquisition URL and a nonempty exact excerpt from the saved Markdown. The excerpt must also support the oracle's product statement; whitespace normalization is used only for that support comparison.
- Quarantine unavailable or ambiguous records, even when their null fields satisfy the schema and correctly match the oracle.
For example, this is a real quarantine entry from the missing-price case:
{
"run": "missing-price-1",
"record": {
"id": "BB202",
"title": "Ceramic Mug",
"price_minor": null,
"currency": "USD",
"source_url": "https://raw.githubusercontent.com/ScrapingAnt/scrapingant-examples/dbf3261c7d3960558ba52fa200f8136e9da28bb2/fixtures/mcp-catalog/missing-price.html",
"evidence": "BB202 | Ceramic Mug | Price: unavailable; currency: USD."
},
"reasons": [
"oracle_unavailable"
]
}
The oracle makes this fixture test measurable. It is not a truth service that can be applied unchanged to arbitrary websites. A production application needs its own acceptance rules and a review path for uncertain records; matching an excerpt alone does not prove that every field is correct.
Results and their denominators
All five MCP retrievals and all 15 Anthropic requests completed. Each model response contained two records after framing was decoded.
| Measurement | Captured result |
|---|---|
| Raw message accepted by strict JSON parsing | 0 / 15 runs |
| Messages requiring one JSON-fence decode | 15 / 15 runs |
| Schema-conforming records after decoding | 30 / 30 emitted records |
| Oracle value precision after decoding | 30 matching records / 30 emitted records |
| Oracle value recall after decoding | 30 matching records / 30 expected records |
| Accepted records | 24 / 30 |
| Quarantined records | 6 / 30 |
| Runs accepting every expected record | 9 / 15 attempted Anthropic runs |
| Omitted, duplicated or invented product IDs | 0 in this fixture capture |
Value precision and recall count correct null reporting as an oracle match. Acceptance is stricter: the three missing-price runs and three conflicting-value runs each quarantine one record. That explains why perfect fixture value matching does not mean every record or run was accepted.
The complete, changed-layout and instruction-bearing cases each accepted both records in all three repetitions. This is evidence about these fixtures and this pinned prompt/model only. Three repetitions of one small page are not three independent web samples, and one ignored page instruction is not a prompt-injection security guarantee.
Run live extraction explicitly
To reproduce the model stage, configure ANTHROPIC_API_KEY through your local secret mechanism. You can reuse the saved MCP results without paying for another retrieval:
MODEL_BUDGET_USD=10 python live.py extract --provider anthropic \
--retrieval-capture expected_output/live-2026-09-30 \
--output /tmp/mcp-model-resume
python replay.py --capture /tmp/mcp-model-resume --write /tmp/mcp-model-resume
Use a new output directory for each experiment. The runner refuses to overwrite a model ledger, records attempts before sending them, and stops on a provider failure without retrying. To repeat retrieval too, configure SCRAPINGANT_API_KEY and use the separate retrieve command in the packet README. Keep credentials out of transcripts and commits.
The recorded Anthropic usage was 9,228 input tokens and 4,632 output tokens, corresponding to $0.032388 at the standard prices recorded in the packet. This is a usage-based estimate, not a billing receipt. The runner conservatively reserves $0.21 per attempt, or $3.15 for the planned 15 calls, below the authorized $10 limit. The MCP envelopes did not expose billed ScrapingAnt credits, so that cost remains unavailable.
An earlier OpenAI request returned 429 insufficient_quota and produced no model output. Its failed attempt remains separately recorded; the table above describes the subsequent Anthropic experiment, not a comparison of model quality.
Where MCP and ScrapingAnt fit
The ScrapingAnt MCP server exposes get_web_page_markdown, which fetches a URL and returns its content as Markdown; the tools accept url, browser, proxy_type and proxy_country. Here it supplies the retrieval step. The application supplies Claude extraction, the JSON Schema, oracle scoring and quarantine behavior.
If you already have saved HTML or Markdown, the replay and validation steps need no ScrapingAnt request. For static pages available through ordinary HTTP, a direct fetch may be sufficient. Use MCP when your client needs that tool interface, and start with the ScrapingAnt MCP documentation for connection and authentication details.
This test does not compare HTML with Markdown, measure autonomous navigation or establish production reliability. Keep transport retries and storage policy separate from model-record validation.
Examples tested on 2026-09-30 with Python 3.12.11, mcp 2.2.0, httpx2 2.13.1, httpx 0.28.1 and jsonschema 4.25.1. Code: executed evidence packet.
This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is the named owner responsible for reviewing the code, measurements and corrections before publication.