Requests vs Playwright vs ScrapingAnt: Cost per Valid Record

Choose a scraping method by the records it delivers. A fast response containing a loading placeholder cannot satisfy a price-monitoring job. A browser may add unnecessary work when the required table is already in the HTML.
The recorded benchmark compared Python Requests, direct Playwright and both ScrapingAnt modes against two controlled, already published fixtures. All four extracted the static catalog correctly. Only the browser arms passed the delayed JavaScript diagnostic. For the catalog, ScrapingAnt raw delivered the same validated records for 90% fewer API credits than ScrapingAnt rendered: 0.3333 credits per accepted record observation, versus 3.3333 credits with rendering. Direct Requests and Playwright had lower observed latency; their total operating cost was not measured.
The measured results: output before speed
These are public-origin results from October 9, 2026: ten randomized, interleaved attempts per target and arm, no retries, 80 acquisitions in total. The 40 ScrapingAnt calls consumed exactly 220 credits according to their response receipts. This is a small fixture study, not evidence of production reliability or a universal cost winner.
The static catalog has three product records with four required fields. Each arm accepted all ten attempts: thirty record observations, repeating the same three unique products.
Static catalog
Observed October 9, 2026 · Existing public-origin fixtures · 10 attempts per arm · No retries
Linear scale from 0 to 1331.14 ms. All-attempt times and observed ranges are below.
Every attempt remains in the accounting. On small screens, scroll the table horizontally.
| Arm | Valid / attempts | Accepted observations | All-attempt latency | Actual API credits |
|---|---|---|---|---|
| Requests | 10 / 10 | 30 catalog records | Median 55.77 msRange 52.95–61.32 ms | No API chargeTotal USD unmeasured |
| Playwright | 10 / 10 | 30 catalog records | Median 237.26 msRange 219.49–277.16 ms | No API chargeTotal USD unmeasured |
| ScrapingAnt raw | 10 / 10 | 30 catalog records | Median 692.28 msRange 648.56–785.27 ms | 10 total0.3333 per accepted observation |
| ScrapingAnt rendered | 10 / 10 | 30 catalog records | Median 1331.14 msRange 1193.03–1414.95 ms | 100 total3.3333 per accepted observation |
30 accepted observations per arm repeat the same 3 unique records. Repeated fixture attempts do not estimate production reliability.
Requests had the lowest observed median time for this complete HTML table. Within ScrapingAnt, raw retrieval satisfied the same contract at one tenth of the rendered arm's API credit cost. Rendering bought no additional accepted catalog data in this sample.
The second fixture changes a text element after a configured 1500 ms delay and creates a readiness marker. It is a DOM readiness diagnostic, not a product catalog. Ten accepted sentinel observations are not ten extracted commercial records.
Delayed DOM diagnostic
Observed October 9, 2026 · Existing public-origin fixtures · 10 attempts per arm · No retries
Linear scale from 0 to 3106.86 ms. All-attempt times and observed ranges are below.
Every attempt remains in the accounting. On small screens, scroll the table horizontally.
| Arm | Valid / attempts | Accepted observations | All-attempt latency | Actual API credits |
|---|---|---|---|---|
| Requests | 0 / 10 | 0 diagnostic sentinels | Median 51.90 msRange 44.31–148.62 msAll outputs invalid | No API chargeTotal USD unmeasured |
| Playwright | 10 / 10 | 10 diagnostic sentinels | Median 2028.79 msRange 2021.13–2173.07 ms | No API chargeTotal USD unmeasured |
| ScrapingAnt raw | 0 / 10 | 0 diagnostic sentinels | Median 695.07 msRange 643.28–1566.77 msAll outputs invalid | 10 totalUndefined per accepted observation |
| ScrapingAnt rendered | 10 / 10 | 10 diagnostic sentinels | Median 3106.86 msRange 3013.67–6754.54 ms | 100 total10 per accepted observation |
Each successful browser arm repeats one text sentinel 10 times. This is not a commercial product-record sample. Repeated fixture attempts do not estimate production reliability.
Both raw arms returned HTTP 200 but kept the initial Web Scraping is hard text, so all their diagnostic outputs failed the contract. ScrapingAnt raw still incurred ten credits. Its cost per accepted sentinel is undefined, because it accepted none. The fast raw responses do not replace valid output.
Direct Playwright and ScrapingAnt with browser=true and wait_for_selector='#loaded' passed all ten diagnostic attempts. Their observed medians were 2028.79 ms and 3106.86 ms respectively, including the fixture's intentional delay. The rendered API range was 3013.67–6754.54 ms; the slowest attempt remains in the dataset. No attempt timed out and no retry was made.
Which approach should you start with?
| Your extraction contract | Starting approach | What to verify |
|---|---|---|
| All required fields are already in response HTML | Requests, or ScrapingAnt raw when managed retrieval fits your deployment | Exact values, complete records, target status and total cost |
| Required data is created by JavaScript | Direct Playwright, or ScrapingAnt rendered with a readiness selector | The page-created signal, then full field and completeness validation |
| You need scripted interactions or browser-context control | Direct Playwright | The state and steps required by your application |
| You want raw and rendered retrieval through one managed API | Evaluate both ScrapingAnt modes on your own permitted target | Accepted output, actual credit receipts and your operational cost model |
Start with the simplest acquisition that passes the validator. Switch to rendering when raw HTML lacks required data, and keep the same validator after the switch. This benchmark shows that rule on controlled pages; it does not settle the cost of running your application.
For HTTP client choices, see Requests vs HTTPX. For a worked missing-JavaScript diagnosis, use the dynamic Python scraping tutorial.
Define a valid record before making requests
For the catalog, success means exactly these ordered records. All four displayed cells must match. A missing row, duplicate table, extra row or wrong value fails the entire page contract.
| Product | Displayed price | Units sold | Share |
|---|---|---|---|
| Ant Farm Deluxe | $1,234.50 | 12,345 | 45.5% |
| Ant Farm Mini | $99.00 | 1,020 | 3.8% |
| Magnifier | $12.25 | 987,654 | 50.7% |
The diagnostic requires exactly one #test element with I ❤️ ScrapingAnt (after 1500 ms) and exactly one page-created #loaded marker. Original text and marker absence fail. The literal oracle in contract.py is written independently of the fixture code.
HTTP 200 alone is insufficient. Every arm must also supply target status 200 and pass the full output contract. An API result with missing target-status evidence would be unaccepted under this protocol. All captured API responses supplied the required status header. A selector tells us when to inspect the DOM; its appearance does not establish completeness.
Four methods, one parser and validator
| Arm | Retrieval | Readiness |
|---|---|---|
| Python Requests | One GET, no configured proxy or browser | Parse response HTML |
| Direct Playwright | Bundled Chromium, reused browser, fresh context and page per attempt | DOMContentLoaded, then attached contract selector |
| ScrapingAnt raw | /v2/general, browser=false, datacenter proxy, timeout=10 | Parse returned HTML; no selector parameter |
| ScrapingAnt rendered | Same endpoint, proxy type and timeout; browser=true, return_page_source=false | wait_for_selector uses the same contract selector as Playwright |
With browser=true, ScrapingAnt renders the target page with JavaScript and returns its HTML. With browser=true, wait_for_selector tells ScrapingAnt to wait for the specified DOM element to appear. See the rendering documentation and selector parameter.
The rendered arm explicitly sets return_page_source=false. The documented browser-enabled page-source mode costs two credits through a datacenter proxy and does not render JavaScript. It is a separate mode outside this four-arm comparison. See the request parameter reference and credit schedule.
The Playwright acquisition uses the remaining readiness budget for navigation and its selector wait:
response = page.goto(url, wait_until='domcontentloaded', timeout=remaining())
row['status'] = row['target_status'] = response.status if response else None
if row['status'] == 200:
page.wait_for_selector(TARGETS[target]['selector'], state='attached', timeout=remaining())
body = page.content().encode('utf-8')
The managed parameters are built by api_parameters():
params = dict(url=url, browser='true' if arm == 'api_rendered' else 'false',
proxy_type='datacenter', timeout=10)
if arm == 'api_rendered':
params.update(return_page_source='false', wait_for_selector=TARGETS[target]['selector'])
These are excerpts from the runnable harness, not standalone scripts. Every acquisition feeds the same parser and literal assertions. No hidden JSON endpoint, resource blocking, stealth configuration or application retry was added. Requests and the API client disable redirects; Playwright retains normal browser navigation. These fixed fixture URLs returned 200 without redirects in the source checks.
API credits per valid result
For one target and one record contract:
API credits per accepted record observation
= all attempted requests' actual credit receipts
/ accepted record observations
Count charged invalid outputs in the numerator. A missing receipt makes actual billed cost unknown. Zero accepted records makes the ratio undefined. Neither case becomes “free.” Repeated observations count extraction work; unique records count distinct data. Report both, and keep catalog records separate from diagnostic sentinels.
Receipt-based API cost
Actual Ant-credits-cost receipts · 10 attempts per arm and target · Units kept separate
Linear scale 0–3.3333. Raw: 10 credits / 30 observations. Rendered: 100 / 30.
Linear scale 0–10. Raw spent 10 credits with no accepted result. Rendered: 100 credits / 10 sentinels.
Undefined means zero accepted outputs, not zero cost. These are API credits, not dollars or measured total operating costs.
For the catalog, the raw calculation is 10 / 30 = 0.3333 credits per record observation; rendered is 100 / 30 = 3.3333. On the diagnostic, raw is 10 / 0, which is undefined; rendered is 100 / 10 = 10 credits per sentinel. That last number measures this readiness diagnostic, not the cost of extracting a commercial product.
For these ordinary non-Google targets, the documented rates are one credit for raw datacenter retrieval and ten for JavaScript rendering through a datacenter proxy. Every captured Ant-credits-cost receipt matched the arm's expected rate. List pricing planned the budget; actual receipts supply the measured numerator. Credit reference.
Credits are not total cash cost
Requests and direct Playwright incur no ScrapingAnt API charge. Their compute, proxy, setup and maintenance costs were not measured. A subscription's allocated dollar cost also depends on your plan and utilization. This evidence therefore cannot declare a total-cost winner across all four methods.
Use a separate model with your own assumptions:
total cost per accepted observation
= (allocated subscription + compute + additional proxies
+ amortized setup and maintenance labor)
/ accepted observations
Supply plan utilization, host/proxy rates, labor hours and an amortization period. Avoid charging for managed browser/proxy work again when it is already included in the API allocation. No labor-saving figure or break-even point is measured here.
Reproduce and inspect the evidence
Use the immutable source and evidence packet, or download the complete packet (ZIP). It includes pinned dependencies, fixture copies, the harness, regression tests, raw HTML captures and offline reports. The original local acquisition revision was 05cbdf3080507b8f3f2d41ad65bb3ddd42b5d587; the published runtime files are byte-identical, and their source hashes are embedded in the public-origin report.
From its examples/http-vs-browser-cost/ directory, on a POSIX system with Python 3.12:
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
python -m playwright install chromium
./run.sh
Browser and package downloads require network access. On a fresh Linux host, follow Playwright's system dependency installation instructions. The captured environment here is macOS.
The default command runs regression tests and forty loopback acquisitions. It never reads an API key or calls ScrapingAnt. It refuses to overwrite an output directory. Use a fresh --out path for another run:
python runner.py --out run_output/another-local-run
python verify_capture.py expected_output/local/report.json
To verify the captured four-arm evidence without network requests or credits:
python verify_capture.py expected_output/live/report.json
The strict verifier requires complete live provenance maps, expected completed receipts and a total within the approved ceiling. It then checks the randomized schedule, all initiated attempts, raw-body and runtime-source hashes, contract decisions, credit totals and recomputed summary. File hashes protect artifact integrity; they do not authenticate provider billing. Legacy local captures without runtime hashes are explicitly labeled. Interrupted work remains as unfinished or aborted with uncertain cost, rather than disappearing from the denominator.
Inspect the complete public-origin report (JSON), all eighty attempts (CSV) and source/public fixture checks (JSON). Empty credit fields on direct arms mean no API receipt, not zero-dollar operating cost.
A new live run requires a bounded budget
The existing owned targets are the static catalog and delayed DOM fixture. The runner checks that both return HTTP 200 and match tracked source bytes before any API work. The October 9 checks passed. Source drift stops the run for review.
The recorded run used prior owner approval for at most 40 API calls and 220 credits, with no pilot, paid warmup or retry. Every new live run needs its own explicit budget approval. Configure SCRAPINGANT_API_KEY securely in the process environment; the harness sends it only in the fixed service endpoint's authentication header. Do not place a key in command arguments or target URLs.
python runner.py --live --repeats 10 --approved-credits 220 --out run_output/live
This command reproduces the design; the captured execution used the distinct output directory run_output/live-approved-20261009. The flag guards spending but does not grant permission. The runner reserves documented cost before each API call and stops on a missing or unexpected receipt. It permits the expected arm charge, or zero with an API error status. Interrupted calls require billing reconciliation before further paid work. No subscription purchase, residential proxy, model request or new endpoint is required.
When ScrapingAnt fits the task
ScrapingAnt provides raw and JavaScript-rendered HTML through one API, with proxy type and readiness in the request configuration. Use browser=false when the required data is already in the response, or browser=true with wait_for_selector when JavaScript creates it. The service runs the rendered browser and proxy retrieval; your application requests and validates the returned HTML.
The measured run makes that choice concrete. Raw mode delivered every catalog record at one tenth of rendered mode's API credit cost. Rendered mode with a readiness selector delivered the JavaScript-created output that raw retrieval missed. One managed API lets you choose the acquisition mode for each extraction contract.
These are service capabilities and observed fixture results. Their operational value depends on your deployment and maintenance needs; setup time, engineering savings and production reliability were not measured. Requests is sufficient for this static catalog. Direct Playwright also handles the delayed diagnostic and remains an option when you need browser-context or interaction control.
Start with raw retrieval on your own permitted target, validate every required field, and enable rendering when the contract needs JavaScript-created data. Use the ScrapingAnt request reference to configure that first request. The Playwright alternative page covers the broader managed-service choice.
Environment, timing and limits
The recorded runs used Python 3.12.10, Requests 2.34.2, BeautifulSoup 4.15.0, Playwright 1.62.0 and bundled Chromium 151.0.7922.34 on macOS 26.6.2 arm64. Client geography was not supplied. API worker browser version and reuse are unknown.
Ten rounds shuffle targets and arms with seed 607; concurrency is one. At least one second separates public-origin acquisitions. One Chromium browser is reused, with a fresh context and page per attempt. Requests and API Sessions are reused. Timings include acquisition, context/page creation where applicable, extraction and validation. They exclude browser launch and context cleanup. The public-origin browser launch was 228.74 ms, reported separately.
Direct acquisition and browser readiness share a ten-second budget; total acquisition plus validation has a fifteen-second caller deadline. The API worker timeout is ten seconds, with transit inside the caller deadline. Worker execution is not observable as direct client timing. Requests connect/read timeouts alone are not a wall-clock deadline, so this POSIX runner enforces one. Requests timeout documentation, Playwright navigation and waits.
The direct and managed arms use different network/proxy paths. These observations compare their acquisition stacks from one client environment, not intrinsic library speed. No concurrent blog build or rendering audit ran during the public-origin acquisition, but unrelated host activity was not controlled.
Ten attempts are a small sample of deterministic pages. The per-fixture Wilson 95% interval is approximately 72.2–100% for ten accepted attempts and 0–27.8% for zero. These are descriptive binomial intervals; shared fixtures and one environment limit their interpretation. They do not estimate production reliability or establish p95, p99 or statistical superiority.
Separate local loopback report (JSON) and forty local attempts (CSV) remain available. Their final capture overlapped a local blog build; local timing excludes Internet transit and is illustrative. Earlier captures are retained in the packet outside the final denominators. None is substituted for an inconvenient public-origin result.
These fixtures do not test production load, browser memory, proxy quality, geography, anti-bot behavior, pagination, authenticated sessions, long-run maintenance or service uptime. Compare methods after they meet your record contract, using every attempt's measured cost and explicitly labeled operational assumptions.
Examples tested on 2026-10-09 with Python 3.12.10, Requests 2.34.2 and Playwright 1.62.0. Code and captures: evidence packet.
ScrapingAnt publishes this comparison. AI agents drafted this article and executed the recorded benchmark and checks. Oleg Kulyk approved publication and is responsible for the code, measurements and corrections.