fix(web): fall back to a non-browser fetch on Cloudflare bot challenges - #89
Open
Sarath1018 wants to merge 2 commits into
Open
Sarath1018 wants to merge 2 commits into
Sarath1018 wants to merge 2 commits into
Conversation
Some Cloudflare-fronted sites (e.g. w3.org) serve headless Chromium a 403 "Just a moment..." challenge (`cf-mitigated: challenge`) instead of the page, so the crawl indexed nothing and failed with "Skipped indexing ...?__cf_chl_rt_tk=... due to HTTP 403 response" (Darwin cc_pair 935). These sites challenge clients that claim to be Chrome but fail fingerprinting, while serving the real page to plainly declared non-browser clients. When the browser gets a challenge, re-fetch the original URL once with requests' default user agent and index that HTML. This also stops the challenge's JS-rewritten `?__cf_chl_rt_tk=` URL from being treated as a redirect and used as the document id. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…orker load_from_state raised "No valid pages" / "Skipped indexing ..." without stopping Playwright, and also leaked an instance restarted after a batch boundary when the remaining pages produced no documents. A sync_playwright instance that isn't stopped leaves its asyncio loop marked running on the thread, so every later start_playwright() on that thread fails with "It looks like you are using Playwright Sync API inside the asyncio loop". Dask workers run --nthreads=1, so one leaked instance broke every later web connector on that pod until it restarted: yesterday's w3.org Cloudflare 403 failures (attempts 583631/583632) poisoned both workers, and today's attempts 584050/584051 failed immediately. Wrap the crawl in try/finally so the current instance is always stopped (stop() is idempotent), and stop the instance in start_playwright if browser launch / context / OAuth setup fails after start(). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR fixes two web-connector bugs that together broke every web connector on both Darwin dask workers.
It looks like you are using Playwright Sync API inside the asyncio loop.1. Cloudflare bot challenge
Problem
The web connectors for
https://www.w3.org/TR/WCAG22/(cc_pair 935) andhttps://www.w3.org/WAI/WCAG22/Understanding/fail with:w3.org sits behind Cloudflare, which returns a 403 "Just a moment…" bot challenge (
cf-mitigated: challenge) to headless Chromium instead of the page. The challenge's JavaScript also rewritespage.urlto a one-time?__cf_chl_rt_tk=URL, which the connector logged as a redirect. The challenge page has no links, so nothing else was queued and the run ended with 0 documents.The challenge is triggered by clients that claim to be a browser but fail Cloudflare's fingerprinting. The same URL returns the real page to plainly declared non-browser clients:
python-requests,curl, a custom bot nameDEFAULT_USER_AGENT), HeadlessChromecf-mitigated: challengeFix
In
WebConnector.load_from_state, when Playwright's response is a Cloudflare challenge (403/503 pluscf-mitigated: challenge):requests, using its default user agent, and parse that HTML in place of the browser content (link extraction and recursion work as usual)?__cf_chl_rt_tk=URL no longer becomes the document idPages that aren't challenged take exactly the same path as before. The global Chrome user agent is unchanged because other sites need it.
2. Leaked Playwright instance poisons the worker thread
Problem
Today every web connector on both workers failed immediately (attempts 584050 and 584051) with:
A
sync_playwrightinstance keeps its asyncio loop marked as running on the thread untilstop()is called, and the nextsync_playwright().start()on that thread refuses to start.load_from_stateleaked an instance in two cases:RuntimeError("Skipped indexing …")/"No valid pages found."without stopping Playwright. This is the w3.org case above.The dask workers run
dask worker --nworkers=1 --nthreads=1, and every task reuses that one thread. So one leak breaks every later web connector on the pod until it restarts. Yesterday's w3.org failures (583631 onbstcr, 583632 onzssm4) were the last web-connector runs on each pod before today's failures.Fix
try/finallyso the current instance is always stopped, whether the crawl produces no documents, hits an unexpected exception or is abandoned by the caller.stop()is idempotent in Playwright 1.41.2 (_exit_was_calledguard), so instances already stopped at batch boundaries are unaffected.start_playwrightnow stops the instance if browser launch, context creation or the OAuth token fetch fails aftersync_playwright().start().Most of this diff is re-indentation. Review with "Hide whitespace"; the real change is about 25 lines.
Testing
test_cloudflare_fallback.py(14 cases): challenge detection and the fallback's accept/reject rules.test_playwright_cleanup.py(3 cases): a fake browser checks that every started instance is stopped for the 0-document, restart-after-batch and failed-launch cases. All 3 fail on the pre-fix code.feature/darwin): reproduces the prod 403 error exactly.https://www.w3.org/TR/WCAG22/(the full WCAG 2.2 spec, about 144k chars) and follows a link torelative-luminance.html, both with clean document ids.docs.uipath.comandexample.comstill go through the browser path unchanged.Deploy note
Both dask worker pods are poisoned right now and stay that way until they restart. Deploying this PR restarts them. Then re-run the two w3.org connectors.
If the workers are restarted before this is deployed, pause the two w3.org connectors first. Otherwise their next scheduled run fails the old way and poisons the workers again.
🤖 Generated with Claude Code