* refactor(web-crawler): remove Playwright rendering, keep static SPA shells
Drop the Playwright-based fallback rendering path so SPA pages are stored
as their static shell HTML instead of being rendered headlessly. SPA
shells now surface the <noscript> notice (e.g. "You need to enable
JavaScript to run this app.") rather than producing an empty document,
and robots.txt-blocked entry pages get a human-readable error message.
- Delete playwright_renderer.py, render_heuristics.py and their tests
- Strip fallback_playwright/playwright_timeout config, fallback_rendered
counter, and render-hint plumbing from the crawler and web importer
- HTMLParser: drop the SPA-empty-pattern stripping and fall back to
<noscript> text when trafilatura extracts nothing
- Humanize the robots.txt entry-failure message
- Migrate the orphaned _convert_to_raw_url tests to HTTPAccessor (the
method moved there in an earlier reorg) and remove the stale file
* refactor(web-importer): simplify robots.txt-blocked import message
Replace the verbose robots.txt explanation with a short compliance hint
that omits the URL and points users to local-file import instead.
* Refactor recursive web import into HTTP accessor
Move ordinary web page import routing into HTTPAccessor and materialize crawled pages as a temporary directory via WebImporter.
Relocate Scrapy/Playwright crawling under parse.accessors.web_crawler, keep trafilatura extraction inside HTMLParser, and avoid repeated ResourceService.add_resource calls.
Add recursive crawl controls, safe request validation, page/download classification, and focused unit coverage.
* Document recursive web crawler options
* fix(web-crawler): stop SSRF sub-resource block from failing whole render
The playwright fallback validated every sub-resource request against the
SSRF guard and raised on the first disallowed host, failing the entire
page render. volcengine docs load a probe resource on an internal host,
so rendering always failed and the crawler stored the static anti-bot
"Please wait..." challenge page as content.
Now a blocked sub-resource is only aborted; the main document and final
URL still gate the result. Also wait past JS interstitials, retry reads
through in-flight navigation, and reject shell/challenge pages instead of
storing them.
* fix(web-crawler): surface renderer error hint on entry-page failure
When Playwright is unavailable, the renderer returns an actionable install
hint via RenderResult.error, but the spider silently kept the static shell
and WebImporter raised only the generic "Failed to fetch entry page". The
hint never reached the user.
Now the spider records rendered.error on the failed page, and WebImporter
appends the entry page's failure reason to the raised message so the CLI
shows the Playwright install instructions.
* fix(web-crawler): surface render hints and enforce crawl limits
* fix(web-crawler): avoid rendering SSR app pages
* perf(web-crawler): bound render concurrency and cap networkidle wait
Playwright renders were dispatched from parse callbacks without any
concurrency limit, so a page with many child links could spawn dozens of
Chromium pages at once (observed peak 28 for a 20-page crawl), risking OOM
on large sites and starting ~2.3x more renders than needed before
max_pages stopped the crawl. Gate renders with a semaphore sized to
config.concurrency and re-check the success limit after acquiring a slot
so queued callbacks skip rendering once the crawl is already done.
Also cap the networkidle wait at 8s: pages with continuous background
activity (e.g. GraphiQL) never go idle and previously blocked until the
full render timeout, turning a ~3s page into ~38s. Content is ready after
domcontentloaded and _wait_past_challenge covers late-arriving text.
Bump default concurrency 5 -> 10.
* fix(web-crawler): route .html/.htm URLs through recursive WebImporter
An explicit .html/.htm URL is detected as DOWNLOAD_HTML via the extension
map, so access() previously only routed URLType.WEBPAGE to WebImporter and
these URLs fell through to single-file download, silently ignoring
depth/max_pages. Route DOWNLOAD_HTML through WebImporter too, treating a
single-page import as the depth=0 case.
* fix(web-crawler): improve HTML extraction and rendering heuristics
- Drop trafilatura favor_precision=True: it stripped the full body of
link-dense pages, keeping only headers.
- Only render __NEXT_DATA__ pages with Playwright when their static body
is too thin; SSR/SSG Next.js pages already ship full text.
- Disable Scrapy telnet console to avoid opening port 6023.
* fix(web-crawler): keep code-hosting single-file URLs off recursive crawler
GitHub/GitLab blob and GitHub raw URLs resolve to a single file, not a
site. Route them through the single-file download path instead of the
recursive WebImporter, which otherwise crawls the hosting UI shell.
* docs(resources): add recursive web crawler usage examples
Add depth/max_pages crawl examples to the HTTP, Python SDK, and CLI
blocks in both the zh and en resource API docs, plus path-prefix
filtering and skip_download_links variants.