HTML vs Markdown vs Plain Text for LLM Extraction

Choose the smallest representation that preserves the information your extraction task needs. Plain text can be sufficient for visible values. Markdown is a useful candidate when headings, tables and link destinations matter. Keep HTML when the answer depends on DOM attributes or script data that conversion removes.
We tested the same six-field extraction contract on three saved inputs, then on 15 hosted MCP returns. HTML produced correct records in each tested case. Markdown worked for the purchase-link case but returned empty arrays for the table. Local text worked for the table; hosted raw-mode text did not. These small, fixed-model results make the practical rule sharper: preserve the fields, then validate complete records before comparing cost.
This study includes reference transformations, field preservation, tokenization, 15 authenticated static-page MCP retrievals and 58 fixed-model extraction responses. Two additional requests failed before inference and remain in the evidence. The three cases contain seven unique products; repeated responses are not independent websites. Hosted acquisition and same-capture comparisons are reported separately. The named author reviewed the methodology, conclusions and byline before publication.
Which format should you send to an LLM?
| Your required information | Candidate to test first | Check before accepting it |
|---|---|---|
| Visible prose or isolated displayed values | Plain text | Labels, boundaries and every required value still survive |
| Headings, ordinary tables, named links and their destinations | Markdown | Row/column associations and actual URLs survive conversion |
data-* attributes, element identity, DOM relationships or script-carried data | HTML, or a purpose-built parser before the model | The field exists in the acquired state and remains in the input |
| A complex table, code example or ambiguous price | Compare the actual outputs | The structure your task depends on survives; validate the model's records |
These are candidates to test, not universal accuracy rankings. Our table experiment shows why: Markdown preserved every required literal but the fixed model returned no records. Use an exact-record validator after extraction, whichever format you choose.
ScrapingAnt's MCP server exposes HTML, Markdown and text tools with shared retrieval controls. That gives you a practical way to evaluate format choice in one managed workflow. The existing URL-to-Markdown product guide covers conversion and its earlier token study; this article focuses on what your task loses.
Same contract, same saved input
The fixed task selects the table with Product, Price, Units sold, Share headers, or the offers under Desk lamps / Travel bottles on the other pages. Other demonstration tables and documentation links are outside that product region. Every format still receives the complete source conversion.
The contract requests those products in document order with six fields: name, displayed_price, units_sold, share, sku, and buy_url. Preserve displayed strings, resolve relative purchase URLs using the supplied source URL, and return null when that product region lacks a field. Do not guess a SKU from its URL.
We used one existing, authorized catalog capture and two authored local fixtures:
| Case | Products | Required non-null slots | What it isolates |
|---|---|---|---|
| Visible table | 3 | 12 | Names, displayed prices, units and percentage shares |
| Purchase links | 2 | 8 | Visible names/prices/SKUs, with identical “Buy” labels and different URLs |
| Attribute-only SKUs | 2 | 8 | Names/prices/links, with SKUs stored only in data-sku |
Absent fields still belong in the JSON contract as nulls, but do not inflate preservation scores. The oracle is a separate literal JSON file, not generated from the parser's output. The same AI execution session wrote the local fixtures and oracle; independent review must check their agreement. These deliberately chosen examples explain specific failures, not their prevalence across websites.
Every arm starts from identical UTF-8 source bytes. HTML keeps those bytes. Markdown follows the inspected converter recipe: BeautifulSoup removes script and noscript, then html2text converts the remaining document with links enabled. Text uses a local Chromium selection reference with page requests blocked and document scripts disabled. The browser worker's inspected selection recipe informed that reference; it does not verify the hosted tool's output, especially browser=false text behavior.
The table capture comes unchanged from the Requests/Playwright acquisition study. We do not repeat its timing or credit comparison here. The richer fixtures remain local files; no new public fixture endpoint was created.
What survived, before asking a model
Visible table
| Format | Tokens | Fields |
|---|---|---|
| HTML | 1,522 | 12/12 |
| Markdown | 573 | 12/12 |
| Text | 415 | 12/12 |
All required values survive in all three reference outputs.
Purchase links
| Format | Tokens | Fields |
|---|---|---|
| HTML | 173 | 8/8 |
| Markdown | 72 | 8/8 |
| Text | 55 | 6/8 |
Text loses two destinations. Markdown retains the links.
Attribute-only SKUs
| Format | Tokens | Fields |
|---|---|---|
| HTML | 143 | 8/8 |
| Markdown | 45 | 6/8 |
| Text | 26 | 4/8 |
Markdown and text lose both data-sku values.
“Present” means the expected non-null literal exists in the representation, allowing a relative link path with its separately supplied base URL. All tested purchase URLs lack query or fragment components; the current relative-URL checker tests pathname only and must be extended before evaluating those components. It is a field-slot survival check. Repeated equal values count as separate expected slots; this check does not establish product associations or accepted model records.
The purchase-link Markdown includes these two captured excerpts, in order:
[Buy](/buy/lamp-a)
[Buy](/buy/lamp-b) Newsletter | Contact
The text output contains Buy twice, but neither destination. A downstream process cannot recover which URLs those labels pointed to from that text alone. The footer also illustrates that conversion is not automatic main-content extraction.
For the SKU case, the source contains this complete product element:
<article data-sku="BOTTLE-RED"><h2>Red Bottle</h2><p>Price: $12.00</p><a href="/buy/red-bottle">Buy</a></article>
The Markdown retains the name, price and link; it contains no BOTTLE-RED. Supplying that SKU in an answer would require some other source or a guess, which this contract forbids.
A stronger loss test: change only the missing fact
We changed only the two purchase destinations in the link fixture, then only the two data-sku values in the attribute fixture. The corresponding source HTML changed, while these derived outputs stayed byte-identical:
For these transforms, two different required answers become indistinguishable to a consumer limited to the lossy representation. That demonstrates information loss without attributing a failure to a model. It does not prove that a model will correctly extract every field from the richer input.
Absolute token counts, with the field denominator beside them
These counts cover the exact representation content. They exclude the task, schema, source URL, tool envelope and model protocol overhead. UTF-8 bytes, Unicode code points and model tokens are different units. All three selected inputs happen to be ASCII, so byte and code-point counts coincide here; that does not generalize to Unicode pages.
| Case / format | Bytes / code points | o200k tokens | cl100k tokens | Present slots |
|---|---|---|---|---|
| Table / HTML | 3,688 / 3,688 | 1,522 | 1,520 | 12/12 |
| Table / Markdown | 1,504 / 1,504 | 573 | 576 | 12/12 |
| Table / text | 1,051 / 1,051 | 415 | 416 | 12/12 |
| Links / HTML | 491 / 491 | 173 | 170 | 8/8 |
| Links / Markdown | 218 / 218 | 72 | 72 | 8/8 |
| Links / text | 174 / 174 | 55 | 57 | 6/8 |
| Attributes / HTML | 399 / 399 | 143 | 139 | 8/8 |
| Attributes / Markdown | 128 / 128 | 45 | 44 | 6/8 |
| Attributes / text | 75 / 75 | 26 | 26 | 4/8 |
Using o200k content counts, the table's HTML/Markdown ratio is about 2.66 and HTML/text about 3.67. The attribute fixture has the larger HTML/text ratio, 5.50, while text loses four of eight required slots. These are payload ratios, not model-bill discounts or measured extraction success rates.
Five local repeats produced identical bytes for each case/arm. They check transform determinism on these inputs, not repeated model outcomes or independent sites. The downloadable report includes every output hash, slot check, tokenizer vocabulary digest and runtime version.
What the fixed model experiment returned
We froze gpt-4.1-mini-2025-04-14 in Chat Completions with temperature 0, top_p 1, one response, max_completion_tokens=1000, default service tier and store=false. Every arm received the same system instruction, task, six-field schema and source URL, with the full representation as untrusted page data. There were no tools, external retrieval, conversation history, automatic retries or prompt revisions. The oracle never entered the prompt.
The seeded, interleaved plan contained five rounds of the nine same-capture inputs, followed by the 15 saved hosted outputs. Two pre-inference failures consumed their planned jobs; explicitly authorized continuations preserved both failures and skipped them. The remaining 58 requests completed. The missing responses are offline links/Markdown repeat 0 and offline table/text repeat 0. All available responses, including those outside matched comparisons, remain downloadable.
For the fairest comparison within each case, this table uses only repeat labels with inference responses in all three arms. “Accepted” means every ordered record and all six fields exactly match the withheld oracle, including required nulls.
| Input / matched repeats | HTML accepted | Markdown accepted | Text accepted |
|---|---|---|---|
| Same-capture table / repeats 1–4 | 4/4 | 0/4 | 4/4 |
| Same-capture purchase links / repeats 1–4 | 4/4 | 4/4 | 0/4 |
| Same-capture attribute SKUs / repeats 0–4 | 5/5 | 0/5 | 0/5 |
| Saved hosted table outputs / repeats 0–4 | 5/5 | 0/5 | 0/5 |
These denominators are repeated model responses on the selected inputs. Five successes out of five do not establish 100% reliability. The report includes nominal 95% Wilson repeat intervals: 5/5 corresponds to approximately 56.6–100%, and 0/5 to 0–43.4%, under an independent-Bernoulli approximation. Repeated identical inputs may be correlated, so those intervals do not estimate performance across websites. We do not pool these different tasks into a best-format score.
Read the failures, not just the percentages
Across all 58 inference responses, 28 pages passed and 30 failed the strict contract. Twenty-five failures returned the valid JSON array []: every table/Markdown response, every links/text and attributes/text response, and every hosted table/text response. The remaining five failures were attribute-case Markdown: both products and their links were returned, but sku was null instead of the source's data-sku value. Those nulls avoid inventing a SKU, yet fail a task that requires it.
The table's Markdown contains every required value; its empty arrays therefore cannot be explained by absent literals alone. We did not vary prompts or converters to identify the cause. Likewise, local browser-selection text passed the table task while hosted raw-mode text returned empty arrays. Their payloads differ in title, hidden row and cell layout, but this run does not isolate which difference caused the result.
The useful decisions are narrower than a format ranking. Test text for a displayed-value task, Markdown for link-sensitive output, and HTML or a deterministic parser for attributes. If complete records fail, investigate the actual payload and prompt before scaling. This run gives no basis for assuming Markdown or text succeeds merely because it uses fewer tokens.
Validate the records, then compare cost per correct output
The runnable contract.py checks exact ordered records and field sets. A missing or extra product, reordered or duplicated row, invented destination, or numeric coercion of a displayed price fails the page. Null fields remain part of this strict record check.
This executable example exercises that scorer against the literal oracle, then introduces a wrong destination. It does not call a model:
from copy import deepcopy
from contract import TRUTH, score
candidate = deepcopy(TRUTH["links"])
print(score("links", candidate)["accepted_page"])
candidate[0]["buy_url"] = "https://catalog.example/buy/wrong"
print(score("links", candidate)["accepted_page"])
Its captured output is:
True
False
When adapting this experiment, keep the source, prompt, schema, model snapshot, settings and retries fixed. Interleave arms, retain every response and failure, and score against a withheld oracle. Report exact accepted pages as the primary denominator and correct field slots separately.
Model cost per accepted page = all model charges / accepted pages
Retrieval credits per accepted page = all charged credits / accepted pages
Incorrect, truncated and failed charged attempts stay in the numerator. If no page is accepted, the ratio is undefined. Keep model cash and API credits separate unless you have an explicit subscription allocation. Token counts alone cannot supply the denominator or observed charges.
Here are all available model observations. Costs are calculated from reported usage at the published standard rates, not billing receipts. “Known cost” excludes the two failures with unknown charges; the last column includes every attempt in that case/arm only when all costs are known. USD amounts are rounded to eight decimal places.
| Case / format | Inference / attempts | Accepted pages | Known token-rate cost (USD) | All-attempt cost / accepted page (USD) |
|---|---|---|---|---|
| Table / HTML | 5/5 | 5 | 0.00318920 | 0.00063784 |
| Table / Markdown | 5/5 | 0 | 0.00190400 | Undefined: zero accepted |
| Table / text | 4/5 | 4 | 0.00237920 | Unknown: one failure's charge |
| Links / HTML | 5/5 | 5 | 0.00208000 | 0.00041600 |
| Links / Markdown | 4/5 | 4 | 0.00150240 | Unknown: one failure's charge |
| Links / text | 5/5 | 0 | 0.00084600 | Undefined: zero accepted |
| Attributes / HTML | 5/5 | 5 | 0.00206800 | 0.00041360 |
| Attributes / Markdown | 5/5 | 0 | 0.00178400 | Undefined: zero accepted |
| Attributes / text | 5/5 | 0 | 0.00078800 | Undefined: zero accepted |
| Hosted table / HTML | 5/5 | 5 | 0.00269000 | 0.00053800 |
| Hosted table / Markdown | 5/5 | 0 | 0.00190400 | Undefined: zero accepted |
| Hosted table / text | 5/5 | 0 | 0.00163200 | Undefined: zero accepted |
The 58 responses reported 48,709 input tokens, including 14,976 cached tokens, and 4,860 output tokens. Using $0.40/M uncached input, $0.10/M cached input and $1.60/M output gives $0.0227668 of known token-rate cost. The two other attempts have no reported usage, so total actual cost remains unknown. Their combined $0.0072 allowance reservation is not observed billing. Published rates are from the official model page, checked October 11, 2026.
Repeated inputs received caching discounts, and a successful full record costs more output tokens than []. Consequently, these run-specific costs are not isolated effects of format size or forecasts for fresh pages. The protocol download records the complete plan, usage, retained failures and accounting. The execution report separates all-available observations from matched repeats and keeps unknown-charge ratios null.
For a captured model-to-validator workflow, see the AI scraper guide.
What the hosted MCP tools actually returned
On October 11, 2026, we made five interleaved calls to each tool for the existing static table URL, with browser=false, proxy_type=datacenter and proxy_country=US fixed. There were no retries or browser/proxy fallbacks. Server 1.30.0 negotiated MCP protocol 2025-11-25.
| Hosted format | Attempts | Outputs containing all 12 required literals | UTF-8 bytes | o200k tokens |
|---|---|---|---|---|
| HTML | 5 | 5 | 3,688 | 1,522 |
| Markdown | 5 | 5 | 1,504 | 573 |
| Text | 5 | 5 | 1,083 | 436 |
Each arm returned identical content in its five observations. All five HTML responses matched the frozen catalog bytes. The Markdown matched its local reference, but hosted text did not match browser selection: it retained the document title and the CSS-hidden stock row, and placed table cells on separate lines. All required priced-product values remained present. These token counts still cover content only.
That difference matters when choosing a tool: raw-mode text is not evidence of what a browser would visibly select. Inspect the actual payload rather than assuming that “text” removes hidden elements or preserves browser table boundaries. We did not test hosted browser-mode text or the two richer local-only fixtures.
A scoped read-only account-counter snapshot before the run and after every call increased by one each time: 15 credits in total, consistent with the documented schedule. The content-only MCP response supplied no per-request credit receipt, and concurrent account usage cannot be ruled out. We report these contemporaneous deltas separately from documented pricing and model-provider costs.
Each of the 15 saved hosted returns subsequently entered the same fixed-model task, without another acquisition. HTML produced five accepted pages; Markdown and text each produced five empty arrays. For this arm's five accepted HTML pages, the contemporaneous retrieval-counter delta was five credits; that is an account observation, not an exclusive per-request receipt. No accepted page exists for the other hosted arms, so a retrieval-cost-per-accepted-page ratio for them is undefined. The model usage costs above remain separate.
This is a captured MCP-tool retrieval followed by a fixed extraction request. It does not test an interactive agent choosing tools, following links, retrying errors or operating a production pipeline.
Live calls reacquire the URL. Matching HTML and Markdown output hashes do not expose the underlying source bytes of separately acquired Markdown/text calls. Keep this hosted acquisition check separate from the offline comparison in which every arm starts with exactly the same saved source.
Choosing the corresponding ScrapingAnt MCP tool
Public, unauthenticated discovery on October 11, 2026 advertised server 1.30.0, protocol 2025-11-25, and these tools:
get_web_page_html: url, browser, proxy_type, proxy_country
get_web_page_markdown: url, browser, proxy_type, proxy_country
get_web_page_text: url, browser, proxy_type, proxy_country
The runnable discovery script only initialized, sent the initialized notification and listed tools. It read no credentials and called no retrieval tool. Full schemas and exchanges are in the discovery JSON.
After configuring a host using the MCP documentation, select a tool by the required fields. The following is an illustrative tool-call request, not an executed retrieval capture. An initialized MCP client can send it with its configured API key:
{
"jsonrpc": "2.0",
"id": 3,
"method": "tools/call",
"params": {
"name": "get_web_page_markdown",
"arguments": {
"url": "https://scrapingant.github.io/scrapingant-examples/fixtures/html-tables.html",
"browser": false,
"proxy_type": "datacenter",
"proxy_country": "US"
}
}
}
Use get_web_page_html when you need attributes, or test get_web_page_text when its actual text payload meets your contract. A real MCP call reacquires the URL, so check source drift and retrieval failures separately from same-capture format comparison. Transport success is insufficient evidence of usable content.
The observed MCP schema has no js_snippet, wait_for_selector or extraction-schema parameter. Those direct-API controls should not be pasted into this tool request. The separate structured JSON extractor is not one of the three observed MCP tools. Rendering and proxy choices also change acquisition behavior; changing them while comparing formats confounds the result.
Try the format choice on a workload worth running
Saved HTML and adequate target APIs or feeds may already satisfy your task. ScrapingAnt is relevant when you want managed URL retrieval, browser rendering or proxy selection alongside the three output choices. This fixture study does not measure production reliability or the operating effort those capabilities might save.
Try these formats on one of your pages, then use the MCP setup documentation. ScrapingAnt's verified signup offer includes 10,000 free API credits monthly with no credit card required. Start with one permitted URL, specify the fields, inspect actual returned content, and validate complete records before expanding the job.
MCP requests use the normal retrieval credit schedule. For a non-Google URL with browser=false and a datacenter proxy, the documented cost is one API credit per request; our account-counter observations were consistent with that amount. These are separate forms of evidence; account deltas do not establish exclusive per-request charges. Model-provider usage is additional. Evaluate your successful workload, failure/retry costs and the current plans before choosing a paid subscription; no paid price or universal savings guarantee is inferred here.
Reproduce and adapt the evidence
Download the complete reproduction ZIP, CSV, full report, or literal oracle.
With Python 3.12 installed, unpack the ZIP and run from its example directory:
python3.12 -m venv .venv
. .venv/bin/activate
python -m pip install -r requirements.txt
python -m playwright install chromium
./run.sh
Installation and the first tokenizer-cache load need internet access. After those prerequisites, the replay uses local files and blocks page requests. It makes no paid retrieval/model calls even if credentials are configured, and compares fresh outputs to the saved evidence without overwriting it.
Change your own oracle when changing your task, then record a new observation separately. Do not update expected results merely to make a failing validation pass. Preserve failed output and distinguish absent information from a model mistake.
Limits and common questions
Is Markdown better than HTML for LLM extraction? It retains the purchase links we tested and produced four accepted link-case responses. It loses attribute-only SKUs and returned empty arrays for our table task. Choose using your required fields and measured records, not a universal Markdown preference.
Does Markdown remove page noise? ScrapingAnt's documented conversion removes script/noscript elements; navigation, footers, cookie banners and advertising can remain. The link fixture's captured Markdown includes its footer. See the conversion reference.
What about tables, code and images? Inspect the conversion output. The documented converter emits ordinary tables as pipe-separated rows, preformatted code as indented text, and images with alt text/source. That does not prove preservation of merged-cell relationships, code language metadata or image meaning. This study measured a simple table and catalog fields, not those richer tasks.
Can text or Markdown make an unsafe source safe? Treat every representation as untrusted page data. Format conversion is not an instruction-safety boundary. The fixed prompt tells the model to treat page content as data, but this study does not measure resistance to hostile instructions.
Does this prove a production winner? No. There are seven unique fixture records, one original catalog capture and two selected local cases. The hosted check makes 15 requests to one static URL. The same execution session created the local cases/oracle. Different converters, browser state, target content, prompts, model versions or tokenizers can change results. We measured repeated fixed-model outcomes, but unknown charges prevent a total all-attempt model cost. Production reliability, prompt robustness and paying-customer impact remain unmeasured.
Study evidence recorded on 2026-10-11 with model gpt-4.1-mini-2025-04-14, Python 3.12.11, BeautifulSoup 4.12.2, html2text 2020.1.16, tiktoken 0.11.0, Playwright 1.62.0 and Chromium 151.0.7922.34. Code and source provenance are included in the reproduction ZIP.
AI agents produced the local fixtures, transformations, authenticated MCP captures, fixed-model measurements and draft. Oleg Kulyk is the publication owner and reviewed the methodology, conclusions and byline before publication.