NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

Web Scraping for Ad Exchanges: Check ads.txt and sellers.json

· 7 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

Web Scraping for Ad Exchanges: Check ads.txt and sellers.json

Correction (2026-09-29)

The earlier article attributed audience insights and bid optimization to scraping without supporting data. This refresh replaces those promises with a tested public-declaration checker. Its inputs are owned synthetic fixtures, not live exchange or auction data.

You have a publisher's ads.txt and an advertising system's sellers.json. Does each declared account ID resolve to a seller record? Which results need investigation, and which cannot be checked because an input is missing or malformed?

This is a useful, bounded web-data task for ad operations. It collects declared relationships rather than ad creatives, audience profiles or performance metrics. We'll run a local Python checker, retain the source of each finding, and distinguish missing evidence from a negative lookup.

Know what the files describe​

In ads.txt v1.1, a publisher lists an advertising-system domain, an account ID and a relationship. DIRECT means the publisher controls that account. RESELLER means another entity controls it with permission to resell the publisher's inventory. A fourth certification identifier is optional.

The sellers.json 1.0 specification maps seller IDs to identity records. Its PUBLISHER, INTERMEDIARY and BOTH types describe the seller's inventory/payment relationship. They are separate from ads.txt's account-control labels; the checker preserves both instead of mechanically equating them.

Neither file is an auction log. A particular bid request's SupplyChain object supplies transaction-specific context that this example does not have. A lookup cannot tell you which ads served, what anyone bid, who saw an ad, or how a campaign performed.

Run the owned fixtures​

Use the pinned example directory in a local checkout of ScrapingAnt/scrapingant-examples. The captured run used Python 3.10.2 and only its standard library. No packages, API key or network retrieval are required.

cd examples/web-scraping-ad-exchanges
./run.sh

The runner checks the hand-authored expectations, prints both reports and writes captures under expected_output/. It stops with a nonzero exit code if a check fails. Its verification output was:

Fixture classifications: 16/16 matched the independent oracle.
Malformed-input regressions: 5/5 rejected before lookup.
Leading-zero IDs, independent relationship/type labels, provenance, variables and duplicate handling: passed.

fixtures/sources.json explicitly maps each synthetic advertising-system domain to a local file. There is no recursive crawl. The collection is declared complete for the exercise; that assumption must not be transferred to an arbitrary downloaded response.

The ads fixture includes this exact line:

EXCHANGE.EXAMPLE , 00123 , direct , cert-demo # whitespace and a leading-zero ID

The parser removes the comment and surrounding field whitespace, lowercases the domain and normalizes the relationship. It leaves 00123 as a string. In the same fixture, 123 is a different account ID and is absent. Coercing both to integers would lose that distinction.

Comments and blank lines are ignored. OWNERDOMAIN and MANAGERDOMAIN declarations are stored separately, including their values, but are not interpreted. This keeps variable lines out of the seller-row totals without pretending to implement their business rules.

Keep failure states separate​

A tempting shortcut is to turn every read error into an empty seller list. That would make an unavailable file look like evidence that an account is absent. Here, load_sellers() returns an error before any lookup if the input cannot be read, parsed or accepted by the checked schema.

It also retains a list of records for each ID rather than letting a dictionary overwrite duplicate IDs. The following decision function is copied from the complete checker:

def classify(index, error, seller_id):
if error:
return 'uncheckable', error, []
matches = index.get(seller_id, [])
if len(matches) > 1:
return 'ambiguous', 'multiple_seller_records', matches
if not matches:
return 'absent', 'seller_id_not_in_complete_fixture', []
if matches[0]['is_confidential'] == 1:
return 'confidential', 'identity_withheld', matches
return 'matching', 'seller_id_found', matches

These status names are this example's diagnostic vocabulary, not an IAB-defined result format. matching means one public record has that exact ID. It does not certify the publisher's declaration or prove an actual sale.

Confidential records still carry a seller ID and type. Under the pinned specification, is_confidential is an integer: 0 or 1; confidential identities can omit name and domain. This strict parser rejects JSON booleans and strings in that field instead of relying on truthiness.

The first five lookup results, copied from the captured text report, show each principal outcome:

Owned synthetic fixture: declaration lookups, not transaction validation.
line system seller_id relationship status reason
3 exchange.example 00123 DIRECT matching seller_id_found
4 exchange.example reseller-a RESELLER matching seller_id_found
5 exchange.example missing DIRECT absent seller_id_not_in_complete_fixture
6 exchange.example private DIRECT confidential identity_withheld
7 exchange.example duplicate RESELLER ambiguous multiple_seller_records

Use the JSON version when you need to investigate a finding:

python3 check_declarations.py --format json

Each row identifies the publisher, ads source file and line. Successful source reads have SHA-256 body hashes in the sources section; matches include their zero-based sellers-array positions. A missing file has no body hash. This lets you trace an observation back to its exact local input without inventing a retrieval timestamp.

For account 00123, the inventory domain is publisher.example, while the business domain is publisher-business.example. That difference is deliberately accepted. IAB's implementation guidance, Case A illustrates the same distinction between a publisher's site and its owner's business domain. Do not turn it into an automatic fraud flag.

Read the totals with their denominators​

Deduplication uses the normalized advertising-system domain, exact seller ID, relationship and optional certification ID together. The repeated row stays visible with a reference to its first line, but is not joined twice.

Captured measureResult
Data rows, excluding comments, blanks and variables16
Invalid data rows2
Repeated data rows1
Unique valid join candidates13
Matching / absent / confidential / ambiguous / uncheckable candidates3 / 2 / 1 / 1 / 6
Variable declarations stored separately2
Unavailable mapped sellers inputs1 of 6

All 16 row classifications matched the independent fixture oracle. Five additional malformed-input checks passed, including duplicate JSON members and a malformed unrelated seller record that invalidates the whole sellers input. These are fixture coverage results, not measurements of exchange quality or market prevalence.

Use the checker within its limits​

This is an offline diagnostic subset. It does not fetch files, establish freshness, interpret variable directives, traverse subdomains, validate Public Suffix List domains or reconstruct a SupplyChain. Unsupported extension syntax can be rejected by this parser even when a broader implementation could understand it.

The ads fixture intentionally mixes usable and malformed rows so you can inspect individual findings. The specification says obviously corrupted or malformed file contents should be ignored. Continuing the diagnostic scan does not establish that the overall file is usable for authorization. Likewise, the duplicate seller-ID case is deliberately inconsistent; the checker reports ambiguity rather than selecting an entity.

ScrapingAnt is not needed for this local text/JSON task. No browser-dependent content is involved, and the captured run used zero API calls and zero credits. For a separate retrieval stage, the headless-browser documentation explains static versus rendered requests; no ScrapingAnt integration was tested here.

Examples tested on 2026-09-29 with Python 3.10.2 (standard library only). Code: versioned evidence packet.

This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is responsible for the published article and corrections.

Forget about getting blocked while scraping the Web

Try out ScrapingAnt Web Scraping API with thousands of proxy servers and an entire headless Chrome cluster