fix: opt-in CDP re-validation of headless-browser navigations (SSRF)

The reqwest DNS resolver can't see Chromium navigations, so a scraped page
that redirects the browser to an internal target (302 -> 127.0.0.1:5432, or
a rebinding hostname) was still loaded. Add crawler::intercept: an opt-in CDP
Fetch guard that re-validates every Document navigation/redirect through the
same ensure_public_target + resolved-IP check, failing internal targets.

Gated behind CRAWLER_SSRF_INTERCEPT (default OFF): enabling Fetch is a fragile
critical-path hook and the CDP wiring is not CI-verifiable (no Chromium in
CI). When off, open_page is byte-identical to browser.new_page; when on, a
guard-install failure falls back to an unguarded navigation rather than
wedging the crawl. The pure decision logic (verdict/is_blocked) is unit-tested;
an #[ignore] smoke test covers the CDP path. Validate with a manual crawl
before enabling in production.

Bump to 0.124.11.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-07-08 06:29:17 +02:00
parent ff4ca964f5
commit bf425cf8e6
15 changed files with 329 additions and 10 deletions

View File

@@ -62,6 +62,43 @@ async fn headless_browser_can_navigate_and_read_title() {
handle.close().await.expect("close cleanly");
}
/// Smoke-test the opt-in SSRF navigation guard (`CRAWLER_SSRF_INTERCEPT`).
/// With interception ON, a normal (allowed) navigation must still complete —
/// i.e. enabling CDP `Fetch` and installing the request handler must NOT wedge
/// page loads. This is the regression the interceptor's fragility risks; it
/// can only be exercised with a real Chromium, hence `#[ignore]`.
///
/// (The private-target *blocking* path is covered by the pure-logic unit tests
/// in `crawler::intercept` — `verdict` / `is_blocked` — which don't need a
/// browser.)
#[tokio::test]
#[ignore = "downloads Chromium; run with --ignored"]
async fn ssrf_interception_does_not_wedge_allowed_navigation() {
use mangalord::crawler::intercept;
const PAGE: &str =
"data:text/html,<html><head><title>Guarded%20OK</title></head><body></body></html>";
intercept::set_enabled(true);
let handle = browser::launch(LaunchOptions::headless())
.await
.expect("launch headless chromium");
// Route through the guarded opener (blank page -> Fetch.enable -> handler
// -> goto). If the handler failed to continue the navigation, this would
// hang until the test harness times out.
let page = intercept::open_page(handle.browser(), PAGE)
.await
.expect("guarded open_page");
page.wait_for_navigation().await.expect("wait for navigation");
let title = page.get_title().await.expect("get title");
assert_eq!(title.as_deref(), Some("Guarded OK"));
handle.close().await.expect("close cleanly");
intercept::set_enabled(false);
}
/// Live end-to-end: navigate to a real page, get the rendered HTML, and
/// parse it with `scraper`. ipify.org renders the visitor's public IP
/// into the page DOM, so a successful run proves browser → render →