Repository navigation
Conversation
…l figure .github/workflows/published-wheel.yml installs the published wheel from PyPI on an OS matrix, so each wheel-install job is one real-interpreter download of our own. It moved to a daily cron and began firing 2026-10-03, which added about four a day to pypi.off_publish_days.*.real_interpreter with no outside user involved. That is the figure the 10-16 activation decision turns on, so the snapshot now names the exclusion and reports the figure raw and net. The refresh reads published-wheel.yml runs from the GitHub Actions API, dates them by the UTC day the run started, and counts the jobs that completed the step "Install the published wheel and run the offline example". The workflow's release job installs nothing, so a bare job count would over-subtract. Three rules keep the subtraction honest. A run on a publish day is listed and not subtracted, because that whole day is already outside the off-publish sum. Each day's subtraction stops at that day's real-interpreter rows, and any surplus is reported as jobs_not_subtracted rather than taken from another day. If the Actions API is unavailable, or its 45-day run inventory starts after the window does, the net field reads unavailable with the reason; the raw figure is never presented as the net one. No row is deleted. downloads and mean_per_day keep their existing meaning as the raw figures, so older snapshots and the classifier do not change under readers' feet. The net sits beside them in net_of_own_ci, with the method and the upper-bound caveat stated once on the parent block. Fresh snapshot, retrieved 2026-10-05T10:21:28Z, series ends 2026-10-04. Seven days: 19 raw real-interpreter events off the publish days, 2.7 a day; net of our own CI 11, 1.6 a day. Thirty days: 36 raw, 1.4 a day; net 24, 0.9 a day. Four of those 26 days have no pypistats rows, so both 30-day means are floors. The exclusion does not explain the move the 2026-10-04 focus review reported. The review's clean week, 2026-09-27 to 2026-10-03, reproduces from frozen feeds at raw 15, 2.1 a day, net 11, 1.6 a day. Like for like, the week before is raw 12, 2.0 a day, net 8, 1.3 a day. Each window holds exactly one four-job run, so the net figure rose. What this change does show is that the raw figure overstated both weeks by four events, and that from 2026-10-03 the daily cron adds about four a day to every window it touches: 8 of the 19 raw events in the committed 7-day window are our own jobs. The net figure is still an upper bound: other CI, reinstalls and mirrors-excluded tooling also use real interpreters, an interpreter does not identify a person, and one job is not provably one download. published-wheel.yml itself is unchanged. It is a real release check and it is working. Refs proof/published-wheel-ci-exclusion-20261005 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: f379b278-63aa-4e9a-86c9-8be52c766129) |
🤖 Claude reviewReviewProcess concern —
|
Summary
.github/workflows/published-wheel.ymlinstalls the published wheel from PyPI on an OS matrix, so a wheel-install job that fetches the wheel adds one real-interpreter download of our own. It moved to a daily cron on 2026-10-03, which added about four a day topypi.off_publish_days.*.real_interpreterwith no outside user involved. That is the figure AG-08 (AG-08: Run the week-four activation and native-alternative decision gate #737) turns on.published-wheel.ymlruns from the GitHub Actions API, counts the jobs that completed the stepInstall the published wheel and run the offline example(the workflow'sreleasejob resolves the version and installs nothing), and reports the off-publish-day real-interpreter figure twice:downloads/mean_per_daystay raw, and a newnet_of_own_ciblock holds the net. A newexclusions.published_wheel_cinames the workflow andexclusions.published_wheel_ci_runslists each run id, date and job count.jobs_not_subtractedrather than taken from another day. If the Actions API is unavailable, or its run inventory starts after the window does, the net field readsunavailablewith the reason; the raw figure is never presented as the net one.published-wheel.ymlitself is unchanged. It is a real release check and it is working.Test plan
python -m pytest sdk/tests/test_activation_evidence.py -q -m "not integration"-> 38 passed. That is the only suite that imports either changed script; the full suite runs in CI.httpsredirect and is dropped on a host change and on anhttps -> httpdowngrade.--published-wheel-runs-from 2026-09-25the 7-day window nets out and the 30-day window reportsunavailable, names itself insources.github_published_wheel_runs.windows_not_subtracted, and adds its own unknown.python -m ruff check scripts/refresh_activation_snapshot.py scripts/activation_weekly_report.py sdk/tests/test_activation_evidence.py-> clean.python scripts/refresh_activation_snapshot.py --fetch-public --retrieved-at 2026-10-05T10:21:28Z --out docs/guides/activation-snapshot-2026-10-05.json.net_of_own_ci.status == "unavailable"with a reason, nodownloadskey, and the classifier reportsunavailable; ...rather than the raw number.jobs_not_subtractedis 1.python scripts/check_docs.py --repository bmdhodl/agent47-> passed (the design doc is on its list).Numbers
Fresh snapshot, retrieved 2026-10-05T10:21:28Z, series ends 2026-10-04.
Four of the 26 days in the 30-day window have no pypistats rows (2026-09-05,
09-06, 09-09, 09-14), so both 30-day means are floors. The 7-day window has no
missing day.
Both figures in that table come from live reads and are not reproducible
from the frozen feeds in this PR. Those feeds reproduce the clean week below and
nothing else; re-running the documented
--fetch-publiccommand later reads alater window. The proof README says the same.
The review's clean week, 2026-09-27 to 2026-10-03, reproduced from the frozen
feeds in
proof/published-wheel-ci-exclusion-20261005/: raw 15 over 7 days(2.1/day), net of the four jobs on run
3711136341111 over 7 days, 1.6/day.The exclusion does not explain the move the review reported. Like for like,
the week before (2026-09-21..2026-09-27, from
docs/guides/activation-snapshot-2026-09-28.json) is raw 12 over 6 days(2.0/day) and net 8 (1.3/day), because it holds run
36190628659with four jobson a day with 8 real-interpreter rows. Each window holds exactly one four-job
run, so net rose 1.3 -> 1.6. What the exclusion does show is that the raw figure
overstated both weeks by four events, and that from 2026-10-03 the daily cron
adds about four a day to every window it touches: the committed 7-day window
carries 8 of its 19 raw events from our own jobs. The first QA round let the
wrong conclusion stand in this PR body; the second caught it.
The net figure is still an upper bound: other CI, reinstalls and
mirrors-excluded tooling also use real interpreters, an interpreter does not
identify a person, and one job is not provably one download.
Artifacts
Plan, verification commands and QA verdict
Approach. Add a fourth named exclusion beside
mirrors,simulated_reportsand
publish_dates. Count wheel-install jobs by step name rather than by jobcount, because the workflow's
releasejob installs nothing and a bare countwould over-subtract. Keep the existing raw keys where they are so the classifier
and the existing snapshots do not change meaning, and add the net beside them.
Tradeoff. chosen: reliability / rejected: simplicity / because: a step-name
job count and an explicit unavailable state cost more code than a bare job count,
but a silent over- or under-subtraction is the dishonest number this change
exists to remove.
Risky areas named before the assertions were written. (a) the boundary to the
Actions API, where an outage must not present raw as net; (b) shared state with
the existing publish-day exclusion, where a run on a publish day must not be
excluded twice; (c) input we do not control, where the workflow's non-install
releasejob must not be counted. Each has a test.Review round. QA lenses read this change twice and found real problems, all
fixed here. Round 1: the GitHub token could ride a cross-host redirect (now
dropped by a custom opener), the subtraction had no per-day cap and
over-subtracted silently (now capped, with the surplus reported), a 35-day run
lookback could under-cover the 30-day window (now 45 days), and
excluded_jobsdid not match the repo's
<noun>_excludednaming. Round 2: the token alsosurvived an
https -> httpdowngrade, a per-window coverage failure wasinvisible outside its own block, the lookback was asserted rather than measured
against the single run page, the offline path could not declare partial
coverage, and
jobs_not_subtractedwas not surfaced to the classifier.Verification commands.
python -m pytest sdk/tests/test_activation_evidence.py -q -m "not integration" python -m ruff check scripts/refresh_activation_snapshot.py scripts/activation_weekly_report.py sdk/tests/test_activation_evidence.py python scripts/refresh_activation_snapshot.py --fetch-public --retrieved-at 2026-10-05T10:21:28Z --out docs/guides/activation-snapshot-2026-10-05.json python scripts/activation_weekly_report.py docs/guides/activation-snapshot-2026-10-05.jsonCitations. Workflow trigger and matrix:
.github/workflows/published-wheel.ymllines 7-13 and 52-73. Install step name: line 73. Prior card in this lineage:
#806, which built the off-publish-day baseline. Live readback 2026-10-05:
actions/workflows/published-wheel.yml/runsandactions/runs/<id>/jobs.QA verdict. Two lenses, two rounds, both WARN in round 2.
Correctness and security lens: confirmed the round-1 fixes by execution,
including re-deriving both windows' arithmetic by hand, and raised 5 more
actionable findings. All 5 are fixed here, each with a test. Its 6 remaining nits
are listed below, un-actioned on purpose.
House style and claim-honesty lens: confirmed the key renames, the single copy of
own_ci_method/own_ci_upper_bound, thatreport.jsonis byte-equal to theclassifier's stdout, and every figure in the proof README against the snapshot.
It then caught the thing that matters most in this change: the clean-week
conclusion was not like-for-like, and the honest comparison says our own CI does
not explain the move. That is corrected above and in the proof README, together
with a false feed date range, an inverted statement of the GitHub rate limit, a
missing-days caveat on the 30-day means, and the "one job is one download"
overclaim.
showwork receipt.
qw-20261004-activation-exclude-published-wheel-ci, sixclaims, shipped in this diff under
.showwork/.verifyis GREEN, 6 of 6checks pass, but the session closed
blocked, notok: I recorded the claimsbefore declaring acceptance requirements, and the gate refuses a clean close
without them (
Outcome: UNVERIFIED. No acceptance requirements declared.).Declaring them after the fact was refused by design. Retracting six true claims
to reorder the ledger would be gaming the receipt, so the receipt says
blockedand names the gap. That is a receipt-discipline gap on my side, not agap in the verification above: every command in the test plan was run and its
result is stated.
Sign-off: Anthropic | claude-opus-5 | high
Known findings not actioned
The correctness lens raised these as nits, not blockers. They are recorded here
so they are not lost. None of them changes a figure in this snapshot.
filter=latest, so a re-run's superseded attempt could be read.created_atrather thanrun_started_at, which can differ for a queued run.--published-wheel-runs-jsonis not validated; a malformed file raises aKeyErrorrather than a clear error.exclusions.published_wheel_ciprose is emitted even when the Actions read failed and nothing was subtracted.runs_excludedlists the runs considered, so itswheel_install_jobsdo not have to sum tojobs_excludedonce the per-day cap binds.tmp_pathfixture it does not use.The house-style lens also left these, all cosmetic:
OWN_CI_EXCLUSIONis stilllonger than its sibling strings, and the proof README's run table no longer
carries the run event because the frozen feed does not hold one.
Risk
Low. The change is measurement maintenance: it adds keys and leaves every
existing key's meaning intact, touches no product code path, and adds no
dependency. The only new external read is this repository's own public Actions
API, and a failed read is reported as
unavailablerather than guessed.Held for review, not merged
This PR changes 1365 lines across 14 files (1335 insertions, 30 deletions): 619
in
scripts/andsdk/tests/, 260 in the snapshot the card requires, 253 inthe classifier output beside it, 140 in the proof README, 41 in the design doc,
49 in the
.showwork/receipt and 3 in the frozen feeds. The queue worker'smerge checklist holds any PR over 400 changed lines, so this hold is automatic
and does not reflect a doubt about the content. Every other merge gate passes;
the risk tier re-derives GREEN at tier 2 across all 14 paths.
The diff was not shrunk to fit. The card requires a dated snapshot and a test
beside the script, and splitting a measurement change from the snapshot it
produces would ship a figure no artifact supports.
Rollback:
git revert <merge-sha>. Nothing outside these files depends on thenew keys, and the raw figures keep their existing meaning, so a revert restores
the previous snapshot behaviour exactly.
Label:
needs:patrick-review.