fix(ci): expose runner cargo and reclaim stale coverage images - #2514
Conversation
Every Rust coverage-evidence run on cwlab-s1-04 failed with "could not run trusted cargo vendor ...: FileNotFoundError" (e.g. the fast-mlsirm#2157 dispatch, run 36529717504): rustup's cargo is in ~/.cargo/bin, which the runner's .path omits. A failed coverage gate keeps opencode-agent from approving, so no fast-mlsirm PR could get approval. - New "Expose runner Rust toolchain" step adds ~/.cargo/bin to GITHUB_PATH when cargo is not already on PATH, and warns when it is absent. - The materializer reports a missing cargo as a runner toolchain gap. - New "Reclaim stale coverage images" step removes only opencode-coverage-tools images older than 2 h, dangling images and build cache older than 24 h; failure warns and never blocks measurement. The guest disk refilled to 96% in hours from per-run images left by runs dispatched before per-run cleanup (#2491). Behavioral tests run each step with fake docker/HOME; the materializer test reproduces the production message before the fix. Dispatch blob pin updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtgtJk1WjqieqvLDb1w8uW
|
Bypass merge record (coordinator-authorized, 🤖 Generated with Claude Code |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (6)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Why
cwlab-s1-04fails.Could not materialize base Rust dependencies: could not run trusted cargo vendor for base manifest Cargo.lock, crates/fast-mlsirm-py/Cargo.lock, fuzz/Cargo.lock: FileNotFoundError, seen on the fast-mlsirm#2157 dispatch, run 36529717504, job 109392516857.cargois in~/.cargo/bin, butactions-runner/.pathdoesn't include it.opencode-coverage-tools:<run>-<attempt>images (2.4–4.6 GB each) accumulate. fix(ci): reclaim per-run OpenCode coverage resources #2491 removes each run's own image at job end, but dispatches created before fix(ci): reclaim per-run OpenCode coverage resources #2491 run the older workflow and leave their images behind.Change
Expose runner Rust toolchain(before measurement). Ifcargois already on PATH (hosted), it does nothing. If~/.cargo/bin/cargoexists, it appends that directory toGITHUB_PATH. Otherwise it emits::warning::, since repositories without Rust don't need cargo.materialize_base_rust_dependencies.pynow reportsFileNotFoundErroras "cargo is not installed or not on PATH on this runner (self-hosted runners keep it in ~/.cargo/bin)".Reclaim stale coverage images(before measurement). It removes onlyreference=opencode-coverage-toolsimages withuntil=2h, then dangling images (image prune -f), then build cache older than 24 h. It never runs-aorsystem prune, so pinned base images stay. A failure emits::warning::and doesn't block measurement. One OpenCode job runs per runner, so the current run's image, built after this step, is never touched.REVIEW_DISPATCH_BLOB_SHAis recomputed for the changed workflow.Tests
tests/test_opencode_coverage_runner_hygiene.py(new, 6 tests) runs each step's actualrun:script with a fakedockerand fakeHOME:-a;~/.cargo/bingets appended;6 failed before the change and 6 pass after.
test_missing_cargo_is_reported_as_a_runner_toolchain_gapreproduced the production message before the fix.All 36 test files referencing the dispatch workflow or the materializer: 900 passed, 1 skipped, both plain and with
GITHUB_ACTIONS=true. The single failure,test_maturin_offline_build_contract, is environmental and fails identically on main in this local environment (the uv venv has no pip).actionlinton the workflow.Follow-up (not in this PR)
Content-addressed coverage image tags (a hash of the Dockerfile, lock inputs and pinned base digest) would reuse identical images across runs. That saves build time, but it changes the
--pull --no-cachefreshness semantics, so it needs its own review. Disk usage is already bounded by #2491 plus this reclaim step.🤖 Generated with Claude Code
https://claude.ai/code/session_01PtgtJk1WjqieqvLDb1w8uW