feat(crawler): CRAWLER_ALLOW_ANY_HOST bypasses the host allowlist (0.44.0)
Operators whose sources shard images across numbered CDN subdomains (cdn1/cdn2/…) can't realistically pre-enumerate every host in CRAWLER_DOWNLOAD_ALLOWLIST. The new flag short-circuits the host check in DownloadAllowlist::contains while leaving the scheme, localhost, and private-IP defenses in is_safe_url untouched, so scraped URLs pointing at 10.x / 169.254.169.254 / file:// stay refused. Wired in both the server (config::build_download_allowlist) and the one-shot bin/crawler.rs entry point. Default is false — fail-closed posture is preserved unless the operator opts in. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
[package]
|
||||
name = "mangalord"
|
||||
version = "0.43.1"
|
||||
version = "0.44.0"
|
||||
edition = "2021"
|
||||
default-run = "mangalord"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user