Skip to content

fix(web): fall back to a non-browser fetch on Cloudflare bot challenges - #89

Open
Sarath1018 wants to merge 2 commits into
feature/darwinfrom
fix/web-connector-cloudflare-challenge
Open

Sarath1018 wants to merge 2 commits into
feature/darwinfrom
fix/web-connector-cloudflare-challenge

Conversation

@Sarath1018

@Sarath1018 Sarath1018 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

This PR fixes two web-connector bugs that together broke every web connector on both Darwin dask workers.

  1. Cloudflare bot challenge: w3.org pages return a 403 challenge to headless Chromium, so those crawls index nothing.
  2. Leaked Playwright instance: that failure left a Playwright instance running on the worker thread, so every later web connector on the same pod failed with It looks like you are using Playwright Sync API inside the asyncio loop.

1. Cloudflare bot challenge

Problem

The web connectors for https://www.w3.org/TR/WCAG22/ (cc_pair 935) and https://www.w3.org/WAI/WCAG22/Understanding/ fail with:

RuntimeError: Skipped indexing https://www.w3.org/TR/WCAG22/?__cf_chl_rt_tk=… due to HTTP 403 response

w3.org sits behind Cloudflare, which returns a 403 "Just a moment…" bot challenge (cf-mitigated: challenge) to headless Chromium instead of the page. The challenge's JavaScript also rewrites page.url to a one-time ?__cf_chl_rt_tk= URL, which the connector logged as a redirect. The challenge page has no links, so nothing else was queued and the run ended with 0 documents.

The challenge is triggered by clients that claim to be a browser but fail Cloudflare's fingerprinting. The same URL returns the real page to plainly declared non-browser clients:

Client identifies itself as Response
python-requests, curl, a custom bot name 200, the full spec
Chrome 120 (our DEFAULT_USER_AGENT), HeadlessChrome 403, cf-mitigated: challenge

Fix

In WebConnector.load_from_state, when Playwright's response is a Cloudflare challenge (403/503 plus cf-mitigated: challenge):

  • re-fetch the original URL once with requests, using its default user agent, and parse that HTML in place of the browser content (link extraction and recursion work as usual)
  • skip redirect handling for the challenge, so the ?__cf_chl_rt_tk= URL no longer becomes the document id
  • if the fallback also fails (non-2xx, non-HTML, or a network error), skip the page with an explicit "blocked by a Cloudflare bot challenge" error and don't retry the browser, which would only get the same challenge

Pages that aren't challenged take exactly the same path as before. The global Chrome user agent is unchanged because other sites need it.

2. Leaked Playwright instance poisons the worker thread

Problem

Today every web connector on both workers failed immediately (attempts 584050 and 584051) with:

playwright._impl._errors.Error: It looks like you are using Playwright Sync API inside the asyncio loop.

A sync_playwright instance keeps its asyncio loop marked as running on the thread until stop() is called, and the next sync_playwright().start() on that thread refuses to start. load_from_state leaked an instance in two cases:

  • 0 documents: it raised RuntimeError("Skipped indexing …") / "No valid pages found." without stopping Playwright. This is the w3.org case above.
  • Restart after a batch: after the batch-boundary restart, if the remaining pages produced no documents, the loop ended with an empty batch and the new instance was never stopped. This case leaked silently, with no error.

The dask workers run dask worker --nworkers=1 --nthreads=1, and every task reuses that one thread. So one leak breaks every later web connector on the pod until it restarts. Yesterday's w3.org failures (583631 on bstcr, 583632 on zssm4) were the last web-connector runs on each pod before today's failures.

Fix

  • Wrap the crawl in try/finally so the current instance is always stopped, whether the crawl produces no documents, hits an unexpected exception or is abandoned by the caller. stop() is idempotent in Playwright 1.41.2 (_exit_was_called guard), so instances already stopped at batch boundaries are unaffected.
  • start_playwright now stops the instance if browser launch, context creation or the OAuth token fetch fails after sync_playwright().start().

Most of this diff is re-indentation. Review with "Hide whitespace"; the real change is about 25 lines.

Testing

  • Unit tests (17 in total, all passing):
    • test_cloudflare_fallback.py (14 cases): challenge detection and the fallback's accept/reject rules.
    • test_playwright_cleanup.py (3 cases): a fake browser checks that every started instance is stopped for the 0-document, restart-after-batch and failed-launch cases. All 3 fail on the pre-fix code.
  • End-to-end runs of the real connector (Playwright 1.41.2, Python 3.11):
    • w3.org, before the fix (on feature/darwin): reproduces the prod 403 error exactly.
    • w3.org, after the fix: indexes https://www.w3.org/TR/WCAG22/ (the full WCAG 2.2 spec, about 144k chars) and follows a link to relative-luminance.html, both with clean document ids.
    • Leak, before the fix: two crawls run back to back on one thread, the first forced to fail. The second fails with the exact prod asyncio error.
    • Leak, after the fix: the second crawl succeeds.
  • Regression check: docs.uipath.com and example.com still go through the browser path unchanged.
  • black, ruff, reorder-python-imports and mypy are clean on the changed files.

Deploy note

Both dask worker pods are poisoned right now and stay that way until they restart. Deploying this PR restarts them. Then re-run the two w3.org connectors.

If the workers are restarted before this is deployed, pause the two w3.org connectors first. Otherwise their next scheduled run fails the old way and poisons the workers again.

🤖 Generated with Claude Code

Sarath1018 and others added 2 commits September 23, 2026 15:55
Some Cloudflare-fronted sites (e.g. w3.org) serve headless Chromium a 403
"Just a moment..." challenge (`cf-mitigated: challenge`) instead of the page,
so the crawl indexed nothing and failed with "Skipped indexing
...?__cf_chl_rt_tk=... due to HTTP 403 response" (Darwin cc_pair 935).

These sites challenge clients that claim to be Chrome but fail
fingerprinting, while serving the real page to plainly declared
non-browser clients. When the browser gets a challenge, re-fetch the
original URL once with requests' default user agent and index that HTML.
This also stops the challenge's JS-rewritten `?__cf_chl_rt_tk=` URL from
being treated as a redirect and used as the document id.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…orker

load_from_state raised "No valid pages" / "Skipped indexing ..." without
stopping Playwright, and also leaked an instance restarted after a batch
boundary when the remaining pages produced no documents. A sync_playwright
instance that isn't stopped leaves its asyncio loop marked running on the
thread, so every later start_playwright() on that thread fails with "It
looks like you are using Playwright Sync API inside the asyncio loop".

Dask workers run --nthreads=1, so one leaked instance broke every later web
connector on that pod until it restarted: yesterday's w3.org Cloudflare 403
failures (attempts 583631/583632) poisoned both workers, and today's
attempts 584050/584051 failed immediately.

Wrap the crawl in try/finally so the current instance is always stopped
(stop() is idempotent), and stop the instance in start_playwright if
browser launch / context / OAuth setup fails after start().

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant