Web Scraping with Playwright Java: A Runnable Maven Example

The earlier version contained disconnected Maven fragments, a subclass accessing a private field, an undefined robots helper, and an incorrect retry helper. Those were source-inspection findings, not compiler results from this experiment. This refresh replaces them with a tested project and removes unsupported hardware and performance claims.
A Java scraper needs to know when a page's data is complete, what a valid record looks like, and whether an empty result means success or failure. This walkthrough runs Playwright against an owned catalog whose products appear after JavaScript executes. It exports validated JSON and checks failure outcomes without contacting a customer site.
The complete Maven project includes the entry point, fixture server, scraper, build configuration and tests. Its captured run passed all six controlled scenarios. That is evidence about this fixture, not a success rate for arbitrary websites.
Set up the complete Java project
Use Docker for the reproduced setup. The tested Linux ARM64 container supplied OpenJDK 25.0.4, Maven 3.9.12 and Chromium 153.0.8010.12. The POM pins Playwright 1.63.0, Jackson 2.18.3 and JUnit 5.12.1, plus build-plugin versions. It targets Java 21 bytecode; Java 21 itself was not the runtime tested here.
Get the exact source snapshot:
git clone https://github.com/ScrapingAnt/scrapingant-examples.git
cd scrapingant-examples
git checkout 7f1780fe62699d70598a7998e352af51965219e5
cd examples/web-scraping-playwright-java
The files have separate responsibilities:
| File | Responsibility |
|---|---|
pom.xml | Dependencies, compiler target and build plugins |
src/main/java/example/Main.java | CLI outcomes and JSON publication |
src/main/java/example/CatalogScraper.java | Browser lifetime, readiness and record validation |
src/main/java/example/FixtureServer.java | Owned loopback HTTP fixture |
fixtures/catalog.html | Delayed catalog and failure scenarios |
src/test/java/example/CatalogScraperTest.java | Exact expected values and failure checks |
run.sh | Build, tests, dependency capture and real CLI invocations |
The container supplies the browser binaries. Initial image and Maven dependency downloads need internet access; the scraping exercise itself only uses loopback. The README also provides an untested native-install alternative and links to the official browser installation instructions. Do not assume the Java or Maven installation on your host matches the captured environment.
Run the scraper and inspect its JSON
From the example directory, run this tested command:
docker run --rm --ipc=host \
-v "$PWD:/work" -w /work \
mcr.microsoft.com/playwright/java:v1.63.0-noble \
bash -c './run.sh && java -cp "target/classes:$(cat target/classpath.txt)" example.Main success output/catalog.json > expected_output/readme-single.stdout.txt 2> expected_output/readme-single.stderr.txt'
It builds and tests the project, then runs the successful scenario in the same container. The JSON appears in output/catalog.json in your mounted directory. The final invocation's captured stdout, saved as expected_output/readme-single.stdout.txt, is:
case=success outcome=success records=2
browser=153.0.8010.12
[ {
"id" : "mug-001",
"title" : "Café mug ☕",
"price" : 12.50,
"currency" : "EUR"
}, {
"id" : "book-002",
"title" : "Field notes — Київ",
"price" : 8.00,
"currency" : "EUR"
} ]
The final invocation's stderr was empty. Keep compilation and this invocation in the same disposable container: a generated classpath file points to Maven's downloaded dependencies, which a new container does not inherit. The packet retains the failed fresh-container attempt as a diagnostic rather than hiding it.
For a fixed container reference, the README records the resolved multi-platform image digest. The command above uses the versioned tag that was actually run.
Wait for a catalog contract, then read records
The owned fixture begins with a loading catalog. After 250 milliseconds, its script appends the records, sets data-count, then marks data-state="ready". A missing readiness marker must not be interpreted as an empty catalog.
Here is the complete CatalogScraper.java from the executed project. Use it with the project's POM and entry point, rather than as a standalone file:
package example;
import com.microsoft.playwright.*;
import com.microsoft.playwright.options.WaitForSelectorState;
import java.math.BigDecimal;
import java.util.*;
public final class CatalogScraper {
public record Product(String id, String title, BigDecimal price, String currency) {}
public record Result(List<Product> products, String browserVersion) {}
public static final class InvalidCatalog extends RuntimeException {
public InvalidCatalog(String message) { super(message); }
}
public static Result scrape(String url, double timeoutMs) {
// All Playwright calls stay on this thread; each scope closes on success or failure.
try (Playwright playwright = Playwright.create();
Browser browser = playwright.chromium().launch();
BrowserContext context = browser.newContext()) {
Page page = context.newPage();
page.navigate(url, new Page.NavigateOptions().setTimeout(timeoutMs));
Locator catalog = page.locator("#catalog[data-state='ready']");
catalog.waitFor(new Locator.WaitForOptions().setState(WaitForSelectorState.ATTACHED).setTimeout(timeoutMs));
String expectedText = catalog.getAttribute("data-count");
if (expectedText == null || !expectedText.matches("[0-9]+"))
throw new InvalidCatalog("Invalid expected count");
int expected;
try { expected = Integer.parseInt(expectedText); }
catch (NumberFormatException e) { throw new InvalidCatalog("Invalid expected count"); }
Locator cards = catalog.locator(".product");
int actual = cards.count(); // Readiness, not count(), establishes list completeness.
if (actual != expected) throw new InvalidCatalog("Count mismatch: expected " + expected + ", found " + actual);
Set<String> ids = new HashSet<>();
List<Product> products = new ArrayList<>();
for (int i = 0; i < actual; i++) {
Locator card = cards.nth(i);
String id = required(card.getAttribute("data-id"), "id");
String title = field(card, ".title", "title");
String priceText = field(card, ".price", "price");
String currency = required(card.getAttribute("data-currency"), "currency");
if (!ids.add(id)) throw new InvalidCatalog("Duplicate id: " + id);
if (!priceText.matches("[0-9]+\\.[0-9]{2}")) throw new InvalidCatalog("Invalid price: " + priceText);
if (!currency.matches("[A-Z]{3}")) throw new InvalidCatalog("Invalid currency: " + currency);
products.add(new Product(id, title, new BigDecimal(priceText), currency));
}
return new Result(List.copyOf(products), browser.version());
}
}
private static String field(Locator card, String selector, String name) {
Locator value = card.locator(selector);
if (value.count() != 1) throw new InvalidCatalog("Missing or repeated field: " + name);
return required(value.textContent(), name);
}
private static String required(String value, String name) {
if (value == null || value.isBlank()) throw new InvalidCatalog("Missing field: " + name);
return value.strip();
}
}
There are two distinct checks. The locator wait establishes the fixture's ready state. The subsequent count check verifies that the number of cards matches its declared total. Calling count() alone is not a completeness wait; see Playwright's locator documentation.
Each title and price locator is scoped to one product card. The scraper rejects missing or repeated fields, blank required values, duplicate IDs and invalid price/currency shapes before returning records. BigDecimal represents the parsed price. The price grammar deliberately accepts only nonnegative values with two decimal places; the currency check only requires three uppercase characters, not membership in an ISO currency list.
The fixture promises that its contents remain stable after readiness. For your own target, replace that contract, the card selectors and the schema together. This example does not implement pagination, streaming updates or arbitrary page interactions.
Distinguish empty results from failed extraction
Main accepts CASE OUTPUT.json. The test runner invokes the real Java process for each scenario and records its exit status:
| Owned scenario | Observed exit | Observed result |
|---|---|---|
success | 0 | Two exact validated records |
empty | 0 | Ready catalog, zero records, JSON empty array |
missing-field | 2 | Missing price rejected; no new output file |
timeout | 3 | Readiness never appears; no new output file |
duplicate | 2 | Duplicate product ID rejected; no new output file |
mismatch | 2 | Three declared cards, two found; no new output file |
For example, the missing-field scenario writes this captured stderr:
case=missing-field outcome=invalid reason=Missing or repeated field: price
The timeout scenario writes:
case=timeout outcome=timeout
The entry point uses exit 1 for unexpected errors and exit 64 for the wrong argument count. Those mappings come from the source; they are not additional scenarios in the six-case measurement. Navigation and readiness each receive a 1,500-millisecond timeout. These are separate operation limits, not a total execution deadline.
Retries and request pacing are outside this example. A delayed Playwright action is not a tested request-rate policy. For bounded retry and deadline design, continue with the dedicated guide:
Publish complete output and close resources
After extraction passes, Main serializes the complete list with Jackson, writes UTF-8 to a temporary file beside the destination, then requests an atomic replacement. A filesystem without atomic-move support returns an error; the example does not fall back to a potentially partial write. Inspect the complete entry point for its error handling and temporary-file cleanup.
A test confirms that an invalid extraction preserves an existing output file. Consumers should still check the process exit status: the presence of an older JSON file does not prove that the latest run succeeded. The packet does not simulate a disk-full or atomic-move failure.
Try-with-resources closes the context, browser and Playwright instance on both successful and failed extraction; the fixture server has its own resource scope. All Playwright operations here stay on one thread, consistent with the official Java threading guidance.
The three JUnit tests cover the six-case exact-value matrix, previous-output preservation and fixture-port reuse. Each matrix case also checked that the JVM had no new live descendant processes after the scraper returned. These observations are useful cleanup checks in the captured container, not proof that every possible operating-system resource leak is absent.
When to use local Playwright or hosted rendering
You do not need ScrapingAnt for this local exercise. If ordinary HTTP already supplies the fields you need, start there; use this browser example when you need to work with the page's rendered DOM and define its readiness condition explicitly.
With browser=true, ScrapingAnt renders the target page with JavaScript and returns its HTML. A request with JavaScript rendering through a datacenter proxy costs 10 API credits. If that acquisition model fits your application, see the headless-browser documentation. That is a documented option, not an API comparison measured by this packet; no paid requests were made. You still need to validate the returned data.
The evidence is limited to an owned synthetic catalog in Chromium. It does not establish third-party reliability, cross-browser behavior, a throughput improvement or a complete production scraper. Adapt the readiness contract and schema to your target, and keep invalid results visible before sending them downstream.
Examples tested on 2026-09-30 with Playwright 1.63.0, Chromium 153.0.8010.12, OpenJDK 25.0.4 and Maven 3.9.12. Code and captured evidence.
This article was drafted with AI assistance from a tested evidence packet. Oleg Kulyk is the named owner responsible for reviewing the code, measurements and corrections before publication.