NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content

One post tagged with "layout robustness"

View All Tags

From HTML to Embeddings - ML-Based Parsers That Survive Layout Changes

· 16 min read
Oleg Kulyk
Co-Founder @ ScrapingAnt

From HTML to Embeddings: ML-Based Parsers That Survive Layout Changes

Correction (2026-09-14)

An earlier version of this article claimed that ML parsers achieve "F1 improvements of 10–25 percentage points over rule-based baselines under layout changes" and "F1 above 0.9 in field deployments" without identifying a source. We could not trace either number to a published experiment. Sections 4.1 and 5.3 now cite the SWDE few-shot results from Li et al. (ACL 2022) and state what those results do and do not show.

Traditional web scraping pipelines rely heavily on brittle, hand-crafted rules – CSS selectors, XPath queries, and regular expressions – that tend to break as soon as a website’s layout or DOM structure changes. With the rapid evolution of front-end frameworks, A/B testing, and personalized content, these brittle approaches impose high maintenance costs and limit scalability.