Test a Production Scraper: Retries, Timeouts and Invalid Data

The earlier article included unsupported avoidance, uptime, anti-bot prevalence and vendor-comparison claims. Those claims have been removed. This refresh replaces the broad recommendations with an executed local failure harness, including incorrect records that passed validation and a deadline that overran under host load. No production-site reliability benchmark is claimed.
A scraper returning HTTP 200 can still produce an empty page, malformed data or the wrong price. Before connecting it to a downstream dataset, test what it does when retrieval, extraction, validation and storage fail separately.
The Python example below runs against an owned, local catalog. It checks retry limits, timeouts, record validation and duplicate delivery. Its most useful result is a failure of the data checks: three deliberately wrong prices passed the schema and reached storage.
Start with the failure HTTP status cannot reveal
The fixture serves a product record inside an HTML script element. Other routes return an empty body, broken JSON, a missing price, a string price or a plausible but wrong numeric price. The naive control treats a successful HTTP status as success. Its captured output is:
Naive control: valid=200, empty=200, malformed=200, missing_price=200, wrong_type=200, wrong_price=200
All six requests returned 200. That check cannot distinguish the valid catalog from the invalid variants. The complete harness therefore separates fetching the body, extracting its record, validating it and writing it to SQLite.
Run the local harness
Download or clone the committed example, then run these commands from its directory:
python3.12 -m venv .venv
. .venv/bin/activate
pip install -r requirements.txt
./run.sh
The packet pins HTTPX, jsonschema and their dependencies. It starts its own server on 127.0.0.1; no API key, paid service or external target is needed. Your environment must allow a loopback TCP listener.
The files have separate responsibilities:
fixture_server.pycontrols response sequences and slow bodies.scraper.pyimplements acquisition, extraction, validation and delivery.fixtures/expected_cases.jsonspecifies expected outcomes independently of the scraper.run_matrix.pychecks results, target-side request counts and stored records.
./run.sh writes captures under expected_output/ and exits unsuccessfully if the observed behavior differs from the expectations. Read those failures before adapting the example to your own target.
Bound retries and total elapsed work separately
The tested settings in scraper.py are:
MAX_ATTEMPTS = 3
JOB_SECONDS = 6.0
IO_SECONDS = 1.0
MAX_BODY_BYTES = 65536
These are laboratory settings, not recommended timeouts for a remote scraping service. Three attempts means the initial request plus at most two retries.
The client retries the declared transient GET statuses—429, 500, 502, 503 and 504—and network or timeout exceptions. Its direct-target 401, 403 and 404 cases stop immediately. Content, schema, business-rule and storage failures do not enter the retry loop. The fixture matrix exercises 429 and 503; the other transient statuses are configured behavior, not separately measured cases.
HTTPX's timeouts distinguish connect, read, write and connection-pool waits. A read timeout limits inactivity while waiting for data. It does not give this example a total budget across retries and backoff.
The slow-body cases make the distinction visible. One server response stalls; another sends a byte every 0.05 seconds. In the final run, the stalled response exhausted three read-timeout attempts. The trickling response kept progressing but was interrupted by the separate total deadline after one attempt.
That deadline wraps the complete asynchronous retry operation using asyncio.timeout_at. It also covers retry sleeps. Cancellation is cooperative: this is not a way to forcibly interrupt arbitrary synchronous parsing or database work.
Respect Retry-After without extending the job forever
The actual retry-wait code takes the larger of its jittered backoff and the server's requested wait:
backoff = rng.uniform(0, min(0.2, 0.05 * 2 ** (attempts - 1)))
delay = max(backoff, retry_after if retry_after is not None else 0)
if loop.time() + delay >= deadline:
return finish('retry_deferred', reason=reason, requested_wait_s=delay)
event('retry_wait', reason=reason, seconds=delay)
await asyncio.sleep(delay)
This is an excerpt from the runnable function. The one-second Retry-After fixture was retried after the required wait; the 60-second fixture returned retry_deferred. The code never shortened that wait to squeeze in another attempt. Here “deferred” is a terminal result for a caller to handle, not a built-in queue or scheduled retry.
The parser handles the HTTP-date and delay-seconds forms. Date parsing, past dates and malformed values have unit tests; the network scenarios use seconds.
Validate shape, business requirements and actual values
The versioned JSON Schema requires an ID, title, integer price in minor units, currency and source URL; it rejects extra properties. The client does not coerce a string price into an integer. A separate rule requires USD for this particular synthetic feed.
The final run produced 30 extracted records. Of those, 24 passed the schema. Missing and string prices failed validation; the EUR record passed the schema but failed the USD rule. Pages that never produced an extracted record are outside that schema denominator.
The wrong-price fixture reveals the remaining gap. Both 1299 and 9999 are valid integers. The schema and USD rule accepted the latter, even though the independently authored fixture oracle says the expected price is 1299. All three wrong-price repetitions were stored.
Across the isolated scenario databases, only 12 of 15 stored records matched that oracle. The scraper itself never reads the oracle: it belongs to the evaluator. Real production records do not arrive with known correct answers. Keep reviewed examples and source captures when assessing factual accuracy; a passing schema alone cannot provide it.
Make delivery and failures observable
Each scenario has a fresh SQLite database with a primary key on record ID. Replaying the same successful job leaves one row and reports duplicate. A query-only SQLite connection produces a real write failure, reported as delivery_failed, without fetching the page again.
This demonstrates suppression of an identical replay, not durable recovery or distributed exactly-once delivery. The example also does not resolve two different records sharing the same ID.
The event log records the job ID, attempt, elapsed time, response status, retry wait, body hash, schema errors, delivery result and terminal outcome. Bounded previews contain only synthetic fixture data. There were exactly 60 terminal events for 60 started jobs, so the run can reconcile every job with its final result.
What the final run showed
The following table summarizes the captured report. Each scenario ran three times.
| Controlled case | Observed behavior per repetition |
|---|---|
| 503 followed by valid content | Two attempts, one stored record |
| Persistent 503 | Three attempts, then stop |
| Short / long Retry-After | Wait and retry / return retry_deferred |
| Stalled / trickling body | Exhaust read-timeout attempts / hit total deadline |
| Direct-target 401, 403, 404 | One attempt, no stored record |
| Empty, synthetic challenge or malformed JSON | Content failure, no stored record |
| Missing / string price | Schema failure, no stored record |
| EUR / wrong numeric price | Business-rule rejection / incorrect record stored |
| Identical replay / failed SQLite write | One row retained / delivery failure |
Selected lines copied from the final command output:
Cases: 19 x 3 repetitions = 57 checks
Passed: 57/57
Jobs: 60; HTTP attempts: 78; terminal events: 60
Schema-valid: 24/30 extracted records
Oracle matches: 12/15 stored records
Wrong-price deliveries despite valid schema: 3
Deadline violations beyond 6.0 s + 1.0 s tolerance: 0
The duplicate scenario starts two jobs, which is why 57 scenario checks cover 60 jobs. The 78 attempts exclude the six separate naive-control requests. Passing every expected-behavior check includes correctly observing the wrong-price failure; it does not mean 57 successful scrapes or a production success rate.
The development failures are part of the evidence
An earlier, tighter timeout configuration passed 56 of 57 checks because a trickle response unexpectedly hit its inactivity timeout and retried. A later loaded-host run passed 55 of 57: stalled responses reached the total deadline instead of exhausting their read timeouts. One terminal event was logged at 2.820 seconds despite a 2.5-second budget.
Both diagnostics remain in the packet. The final settings separate the intended timing cases more widely and allow one second of scheduling tolerance. The zero-violation count above therefore means zero jobs exceeded six seconds plus that allowance. It is not a hard wall-clock guarantee.
Where ScrapingAnt fits
ScrapingAnt is not needed for this harness, already available HTML or a static page your HTTP client can retrieve. Validation and delivery failures remain your application's responsibility.
For a target needing managed acquisition, the ScrapingAnt request interface can replace the retrieval stage. With browser=true, ScrapingAnt renders the target page with JavaScript and returns its HTML; do not set return_page_source=true for that mode. A JavaScript-rendered request through a datacenter proxy costs 10 API credits. A request without a browser through a datacenter proxy costs 1 API credit for a non-Google target.
That adapter was not exercised here. Keep client transport failures, ScrapingAnt API status, target status when available, and record validation separate. The fixture's direct-target status policy must not be copied blindly to a provider API.
Limits before you reuse the example
This is a small local acceptance harness. It did not test live anti-bot systems, browser readiness, high concurrency, remote storage or persistent restart recovery. The synthetic challenge page tests missing expected content, not reliable CAPTCHA detection. The body-size cap is implemented but has no oversized-response network case in this matrix.
Replace the fixture contract with your own required fields and reviewed expected values, then keep transport, extraction and delivery outcomes separate. Treat an untested failure mode as work remaining, even when the current matrix passes.
Examples tested on 2026-09-29 with Python 3.12.11, HTTPX 0.28.1 and jsonschema 4.25.1. Code and captured outputs: scrapingant-examples.
This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is responsible for the published article and corrections.