NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

Test a Production Scraper: Retries, Timeouts and Invalid Data

· 9 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Test a Production Scraper: Retries, Timeouts and Invalid Data

Correction (2026-09-29)

The earlier article included unsupported avoidance, uptime, anti-bot prevalence and vendor-comparison claims. Those claims have been removed. This refresh replaces the broad recommendations with an executed local failure harness, including incorrect records that passed validation and a deadline that overran under host load. No production-site reliability benchmark is claimed.

A scraper returning HTTP 200 can still produce an empty page, malformed data or the wrong price. Before connecting it to a downstream dataset, test what it does when retrieval, extraction, validation and storage fail separately.

The Python example below runs against an owned, local catalog. It checks retry limits, timeouts, record validation and duplicate delivery. Its most useful result is a failure of the data checks: three deliberately wrong prices passed the schema and reached storage.

Start with the failure HTTP status cannot reveal​

The fixture serves a product record inside an HTML script element. Other routes return an empty body, broken JSON, a missing price, a string price or a plausible but wrong numeric price. The naive control treats a successful HTTP status as success. Its captured output is:

Naive control: valid=200, empty=200, malformed=200, missing_price=200, wrong_type=200, wrong_price=200

All six requests returned 200. That check cannot distinguish the valid catalog from the invalid variants. The complete harness therefore separates fetching the body, extracting its record, validating it and writing it to SQLite.

Run the local harness​

Download or clone the committed example, then run these commands from its directory:

python3.12 -m venv .venv
. .venv/bin/activate
pip install -r requirements.txt
./run.sh

The packet pins HTTPX, jsonschema and their dependencies. It starts its own server on 127.0.0.1; no API key, paid service or external target is needed. Your environment must allow a loopback TCP listener.

The files have separate responsibilities:

  • fixture_server.py controls response sequences and slow bodies.
  • scraper.py implements acquisition, extraction, validation and delivery.
  • fixtures/expected_cases.json specifies expected outcomes independently of the scraper.
  • run_matrix.py checks results, target-side request counts and stored records.

./run.sh writes captures under expected_output/ and exits unsuccessfully if the observed behavior differs from the expectations. Read those failures before adapting the example to your own target.

Bound retries and total elapsed work separately​

The tested settings in scraper.py are:

MAX_ATTEMPTS = 3
JOB_SECONDS = 6.0
IO_SECONDS = 1.0
MAX_BODY_BYTES = 65536

These are laboratory settings, not recommended timeouts for a remote scraping service. Three attempts means the initial request plus at most two retries.

The client retries the declared transient GET statuses—429, 500, 502, 503 and 504—and network or timeout exceptions. Its direct-target 401, 403 and 404 cases stop immediately. Content, schema, business-rule and storage failures do not enter the retry loop. The fixture matrix exercises 429 and 503; the other transient statuses are configured behavior, not separately measured cases.

HTTPX's timeouts distinguish connect, read, write and connection-pool waits. A read timeout limits inactivity while waiting for data. It does not give this example a total budget across retries and backoff.

The slow-body cases make the distinction visible. One server response stalls; another sends a byte every 0.05 seconds. In the final run, the stalled response exhausted three read-timeout attempts. The trickling response kept progressing but was interrupted by the separate total deadline after one attempt.

That deadline wraps the complete asynchronous retry operation using asyncio.timeout_at. It also covers retry sleeps. Cancellation is cooperative: this is not a way to forcibly interrupt arbitrary synchronous parsing or database work.

Respect Retry-After without extending the job forever​

The actual retry-wait code takes the larger of its jittered backoff and the server's requested wait:

backoff = rng.uniform(0, min(0.2, 0.05 * 2 ** (attempts - 1)))
delay = max(backoff, retry_after if retry_after is not None else 0)
if loop.time() + delay >= deadline:
return finish('retry_deferred', reason=reason, requested_wait_s=delay)
event('retry_wait', reason=reason, seconds=delay)
await asyncio.sleep(delay)

This is an excerpt from the runnable function. The one-second Retry-After fixture was retried after the required wait; the 60-second fixture returned retry_deferred. The code never shortened that wait to squeeze in another attempt. Here “deferred” is a terminal result for a caller to handle, not a built-in queue or scheduled retry.

The parser handles the HTTP-date and delay-seconds forms. Date parsing, past dates and malformed values have unit tests; the network scenarios use seconds.

Validate shape, business requirements and actual values​

The versioned JSON Schema requires an ID, title, integer price in minor units, currency and source URL; it rejects extra properties. The client does not coerce a string price into an integer. A separate rule requires USD for this particular synthetic feed.

The final run produced 30 extracted records. Of those, 24 passed the schema. Missing and string prices failed validation; the EUR record passed the schema but failed the USD rule. Pages that never produced an extracted record are outside that schema denominator.

The wrong-price fixture reveals the remaining gap. Both 1299 and 9999 are valid integers. The schema and USD rule accepted the latter, even though the independently authored fixture oracle says the expected price is 1299. All three wrong-price repetitions were stored.

Across the isolated scenario databases, only 12 of 15 stored records matched that oracle. The scraper itself never reads the oracle: it belongs to the evaluator. Real production records do not arrive with known correct answers. Keep reviewed examples and source captures when assessing factual accuracy; a passing schema alone cannot provide it.

Make delivery and failures observable​

Each scenario has a fresh SQLite database with a primary key on record ID. Replaying the same successful job leaves one row and reports duplicate. A query-only SQLite connection produces a real write failure, reported as delivery_failed, without fetching the page again.

This demonstrates suppression of an identical replay, not durable recovery or distributed exactly-once delivery. The example also does not resolve two different records sharing the same ID.

The event log records the job ID, attempt, elapsed time, response status, retry wait, body hash, schema errors, delivery result and terminal outcome. Bounded previews contain only synthetic fixture data. There were exactly 60 terminal events for 60 started jobs, so the run can reconcile every job with its final result.

What the final run showed​

The following table summarizes the captured report. Each scenario ran three times.

Controlled caseObserved behavior per repetition
503 followed by valid contentTwo attempts, one stored record
Persistent 503Three attempts, then stop
Short / long Retry-AfterWait and retry / return retry_deferred
Stalled / trickling bodyExhaust read-timeout attempts / hit total deadline
Direct-target 401, 403, 404One attempt, no stored record
Empty, synthetic challenge or malformed JSONContent failure, no stored record
Missing / string priceSchema failure, no stored record
EUR / wrong numeric priceBusiness-rule rejection / incorrect record stored
Identical replay / failed SQLite writeOne row retained / delivery failure

Selected lines copied from the final command output:

Cases: 19 x 3 repetitions = 57 checks
Passed: 57/57
Jobs: 60; HTTP attempts: 78; terminal events: 60
Schema-valid: 24/30 extracted records
Oracle matches: 12/15 stored records
Wrong-price deliveries despite valid schema: 3
Deadline violations beyond 6.0 s + 1.0 s tolerance: 0

The duplicate scenario starts two jobs, which is why 57 scenario checks cover 60 jobs. The 78 attempts exclude the six separate naive-control requests. Passing every expected-behavior check includes correctly observing the wrong-price failure; it does not mean 57 successful scrapes or a production success rate.

The development failures are part of the evidence​

An earlier, tighter timeout configuration passed 56 of 57 checks because a trickle response unexpectedly hit its inactivity timeout and retried. A later loaded-host run passed 55 of 57: stalled responses reached the total deadline instead of exhausting their read timeouts. One terminal event was logged at 2.820 seconds despite a 2.5-second budget.

Both diagnostics remain in the packet. The final settings separate the intended timing cases more widely and allow one second of scheduling tolerance. The zero-violation count above therefore means zero jobs exceeded six seconds plus that allowance. It is not a hard wall-clock guarantee.

Where ScrapingAnt fits​

ScrapingAnt is not needed for this harness, already available HTML or a static page your HTTP client can retrieve. Validation and delivery failures remain your application's responsibility.

For a target needing managed acquisition, the ScrapingAnt request interface can replace the retrieval stage. With browser=true, ScrapingAnt renders the target page with JavaScript and returns its HTML; do not set return_page_source=true for that mode. A JavaScript-rendered request through a datacenter proxy costs 10 API credits. A request without a browser through a datacenter proxy costs 1 API credit for a non-Google target.

That adapter was not exercised here. Keep client transport failures, ScrapingAnt API status, target status when available, and record validation separate. The fixture's direct-target status policy must not be copied blindly to a provider API.

Limits before you reuse the example​

This is a small local acceptance harness. It did not test live anti-bot systems, browser readiness, high concurrency, remote storage or persistent restart recovery. The synthetic challenge page tests missing expected content, not reliable CAPTCHA detection. The body-size cap is implemented but has no oversized-response network case in this matrix.

Replace the fixture contract with your own required fields and reviewed expected values, then keep transport, extraction and delivery outcomes separate. Treat an untested failure mode as work remaining, even when the current matrix passes.

📚Related Reading

Requests vs HTTPX in 2026: Measured Differences and Migration Gotchas

requests 2.34 vs httpx 0.28 measured on one server: connection reuse, threads vs asyncio, HTTP/2, streaming, retries, and the errors a migration hits.

📚Related Reading

Handling Scrapy Failure URLs - A Comprehensive Guide

Learn how to handle failure URLs in Scrapy, a popular web scraping framework. This guide covers strategies for managing failed requests, retrying requests, and logging errors to ensure a smooth and efficient web scraping process.

Examples tested on 2026-09-29 with Python 3.12.11, HTTPX 0.28.1 and jsonschema 4.25.1. Code and captured outputs: scrapingant-examples.

This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is responsible for the published article and corrections.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster