Skip to content

feat(activation): subtract our own published-wheel CI from the install figure - #843

Open
bmdhodl wants to merge 1 commit into
mainfrom
fix/activation-snapshot-exclude-published-wheel-ci
Open

bmdhodl wants to merge 1 commit into
mainfrom
fix/activation-snapshot-exclude-published-wheel-ci

Conversation

@bmdhodl

@bmdhodl bmdhodl commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • .github/workflows/published-wheel.yml installs the published wheel from PyPI on an OS matrix, so a wheel-install job that fetches the wheel adds one real-interpreter download of our own. It moved to a daily cron on 2026-10-03, which added about four a day to pypi.off_publish_days.*.real_interpreter with no outside user involved. That is the figure AG-08 (AG-08: Run the week-four activation and native-alternative decision gate #737) turns on.
  • The refresh script now reads published-wheel.yml runs from the GitHub Actions API, counts the jobs that completed the step Install the published wheel and run the offline example (the workflow's release job resolves the version and installs nothing), and reports the off-publish-day real-interpreter figure twice: downloads / mean_per_day stay raw, and a new net_of_own_ci block holds the net. A new exclusions.published_wheel_ci names the workflow and exclusions.published_wheel_ci_runs lists each run id, date and job count.
  • No row is deleted, and a run that fell on a publish day is listed but not subtracted, because that whole day is already outside the off-publish sum. Each day's subtraction stops at that day's real-interpreter rows, and any surplus is reported as jobs_not_subtracted rather than taken from another day. If the Actions API is unavailable, or its run inventory starts after the window does, the net field reads unavailable with the reason; the raw figure is never presented as the net one.
  • published-wheel.yml itself is unchanged. It is a real release check and it is working.

Test plan

  • python -m pytest sdk/tests/test_activation_evidence.py -q -m "not integration" -> 38 passed. That is the only suite that imports either changed script; the full suite runs in CI.
  • The new tests fail against the pre-change scripts and pass after (verified by checking out HEAD copies of only the two scripts, then restoring), so they are falsifiable.
  • Token containment asserted directly on the redirect handler: the header survives a same-host https redirect and is dropped on a host change and on an https -> http downgrade.
  • Partial coverage asserted end to end: with --published-wheel-runs-from 2026-09-25 the 7-day window nets out and the 30-day window reports unavailable, names itself in sources.github_published_wheel_runs.windows_not_subtracted, and adds its own unknown.
  • python -m ruff check scripts/refresh_activation_snapshot.py scripts/activation_weekly_report.py sdk/tests/test_activation_evidence.py -> clean.
  • Documented refresh command run live: python scripts/refresh_activation_snapshot.py --fetch-public --retrieved-at 2026-10-05T10:21:28Z --out docs/guides/activation-snapshot-2026-10-05.json.
  • Actions-API-unavailable path asserted: net_of_own_ci.status == "unavailable" with a reason, no downloads key, and the classifier reports unavailable; ... rather than the raw number.
  • Per-day cap asserted: on a fixture day with four of our jobs but three real-interpreter rows, three are subtracted and jobs_not_subtracted is 1.
  • python scripts/check_docs.py --repository bmdhodl/agent47 -> passed (the design doc is on its list).
  • Every committed JSON artifact parses as UTF-8 with no BOM, and every figure in the proof README and the design doc was read back out of the committed snapshot.

Numbers

Fresh snapshot, retrieved 2026-10-05T10:21:28Z, series ends 2026-10-04.

Window Raw real-interpreter Own-CI jobs Net Net per day
7d, 2026-09-28..2026-10-04 19 (2.7/day) 8 11 1.6
30d, 2026-09-05..2026-10-04 36 (1.4/day) 12 24 0.9

Four of the 26 days in the 30-day window have no pypistats rows (2026-09-05,
09-06, 09-09, 09-14), so both 30-day means are floors. The 7-day window has no
missing day.

Both figures in that table come from live reads and are not reproducible
from the frozen feeds in this PR. Those feeds reproduce the clean week below and
nothing else; re-running the documented --fetch-public command later reads a
later window. The proof README says the same.

The review's clean week, 2026-09-27 to 2026-10-03, reproduced from the frozen
feeds in proof/published-wheel-ci-exclusion-20261005/: raw 15 over 7 days
(2.1/day), net of the four jobs on run 37111363411 11 over 7 days, 1.6/day.

The exclusion does not explain the move the review reported. Like for like,
the week before (2026-09-21..2026-09-27, from
docs/guides/activation-snapshot-2026-09-28.json) is raw 12 over 6 days
(2.0/day) and net 8 (1.3/day), because it holds run 36190628659 with four jobs
on a day with 8 real-interpreter rows. Each window holds exactly one four-job
run, so net rose 1.3 -> 1.6. What the exclusion does show is that the raw figure
overstated both weeks by four events, and that from 2026-10-03 the daily cron
adds about four a day to every window it touches: the committed 7-day window
carries 8 of its 19 raw events from our own jobs. The first QA round let the
wrong conclusion stand in this PR body; the second caught it.

The net figure is still an upper bound: other CI, reinstalls and
mirrors-excluded tooling also use real interpreters, an interpreter does not
identify a person, and one job is not provably one download.

Artifacts

Plan, verification commands and QA verdict

Approach. Add a fourth named exclusion beside mirrors, simulated_reports
and publish_dates. Count wheel-install jobs by step name rather than by job
count, because the workflow's release job installs nothing and a bare count
would over-subtract. Keep the existing raw keys where they are so the classifier
and the existing snapshots do not change meaning, and add the net beside them.

Tradeoff. chosen: reliability / rejected: simplicity / because: a step-name
job count and an explicit unavailable state cost more code than a bare job count,
but a silent over- or under-subtraction is the dishonest number this change
exists to remove.

Risky areas named before the assertions were written. (a) the boundary to the
Actions API, where an outage must not present raw as net; (b) shared state with
the existing publish-day exclusion, where a run on a publish day must not be
excluded twice; (c) input we do not control, where the workflow's non-install
release job must not be counted. Each has a test.

Review round. QA lenses read this change twice and found real problems, all
fixed here. Round 1: the GitHub token could ride a cross-host redirect (now
dropped by a custom opener), the subtraction had no per-day cap and
over-subtracted silently (now capped, with the surplus reported), a 35-day run
lookback could under-cover the 30-day window (now 45 days), and excluded_jobs
did not match the repo's <noun>_excluded naming. Round 2: the token also
survived an https -> http downgrade, a per-window coverage failure was
invisible outside its own block, the lookback was asserted rather than measured
against the single run page, the offline path could not declare partial
coverage, and jobs_not_subtracted was not surfaced to the classifier.

Verification commands.

python -m pytest sdk/tests/test_activation_evidence.py -q -m "not integration"
python -m ruff check scripts/refresh_activation_snapshot.py scripts/activation_weekly_report.py sdk/tests/test_activation_evidence.py
python scripts/refresh_activation_snapshot.py --fetch-public --retrieved-at 2026-10-05T10:21:28Z --out docs/guides/activation-snapshot-2026-10-05.json
python scripts/activation_weekly_report.py docs/guides/activation-snapshot-2026-10-05.json

Citations. Workflow trigger and matrix: .github/workflows/published-wheel.yml
lines 7-13 and 52-73. Install step name: line 73. Prior card in this lineage:
#806, which built the off-publish-day baseline. Live readback 2026-10-05:
actions/workflows/published-wheel.yml/runs and actions/runs/<id>/jobs.

QA verdict. Two lenses, two rounds, both WARN in round 2.

Correctness and security lens: confirmed the round-1 fixes by execution,
including re-deriving both windows' arithmetic by hand, and raised 5 more
actionable findings. All 5 are fixed here, each with a test. Its 6 remaining nits
are listed below, un-actioned on purpose.

House style and claim-honesty lens: confirmed the key renames, the single copy of
own_ci_method / own_ci_upper_bound, that report.json is byte-equal to the
classifier's stdout, and every figure in the proof README against the snapshot.
It then caught the thing that matters most in this change: the clean-week
conclusion was not like-for-like, and the honest comparison says our own CI does
not explain the move. That is corrected above and in the proof README, together
with a false feed date range, an inverted statement of the GitHub rate limit, a
missing-days caveat on the 30-day means, and the "one job is one download"
overclaim.

showwork receipt. qw-20261004-activation-exclude-published-wheel-ci, six
claims, shipped in this diff under .showwork/. verify is GREEN, 6 of 6
checks pass, but the session closed blocked, not ok: I recorded the claims
before declaring acceptance requirements, and the gate refuses a clean close
without them (Outcome: UNVERIFIED. No acceptance requirements declared.).
Declaring them after the fact was refused by design. Retracting six true claims
to reorder the ledger would be gaming the receipt, so the receipt says
blocked and names the gap. That is a receipt-discipline gap on my side, not a
gap in the verification above: every command in the test plan was run and its
result is stated.

Sign-off: Anthropic | claude-opus-5 | high

Known findings not actioned

The correctness lens raised these as nits, not blockers. They are recorded here
so they are not lost. None of them changes a figure in this snapshot.

  1. Skipped jobs count as zero wheel installs, and the runs list does not pass filter=latest, so a re-run's superseded attempt could be read.
  2. Runs are dated by created_at rather than run_started_at, which can differ for a queued run.
  3. The offline run JSON passed to --published-wheel-runs-json is not validated; a malformed file raises a KeyError rather than a clear error.
  4. The exclusions.published_wheel_ci prose is emitted even when the Actions read failed and nothing was subtracted.
  5. runs_excluded lists the runs considered, so its wheel_install_jobs do not have to sum to jobs_excluded once the per-day cap binds.
  6. The shipped snapshot's own figures come from live reads and are not reproducible from the frozen feeds; those feeds reproduce the clean week only. Stated in the proof README too.
  7. One test takes a tmp_path fixture it does not use.

The house-style lens also left these, all cosmetic: OWN_CI_EXCLUSION is still
longer than its sibling strings, and the proof README's run table no longer
carries the run event because the frozen feed does not hold one.

Risk

Low. The change is measurement maintenance: it adds keys and leaves every
existing key's meaning intact, touches no product code path, and adds no
dependency. The only new external read is this repository's own public Actions
API, and a failed read is reported as unavailable rather than guessed.

Held for review, not merged

This PR changes 1365 lines across 14 files (1335 insertions, 30 deletions): 619
in scripts/ and sdk/tests/, 260 in the snapshot the card requires, 253 in
the classifier output beside it, 140 in the proof README, 41 in the design doc,
49 in the .showwork/ receipt and 3 in the frozen feeds. The queue worker's
merge checklist holds any PR over 400 changed lines, so this hold is automatic
and does not reflect a doubt about the content. Every other merge gate passes;
the risk tier re-derives GREEN at tier 2 across all 14 paths.

The diff was not shrunk to fit. The card requires a dated snapshot and a test
beside the script, and splitting a measurement change from the snapshot it
produces would ship a figure no artifact supports.

Rollback: git revert <merge-sha>. Nothing outside these files depends on the
new keys, and the raw figures keep their existing meaning, so a revert restores
the previous snapshot behaviour exactly.

Label: needs:patrick-review.

…l figure

.github/workflows/published-wheel.yml installs the published wheel from PyPI on
an OS matrix, so each wheel-install job is one real-interpreter download of our
own. It moved to a daily cron and began firing 2026-10-03, which added about
four a day to pypi.off_publish_days.*.real_interpreter with no outside user
involved. That is the figure the 10-16 activation decision turns on, so the
snapshot now names the exclusion and reports the figure raw and net.

The refresh reads published-wheel.yml runs from the GitHub Actions API, dates
them by the UTC day the run started, and counts the jobs that completed the step
"Install the published wheel and run the offline example". The workflow's
release job installs nothing, so a bare job count would over-subtract.

Three rules keep the subtraction honest. A run on a publish day is listed and
not subtracted, because that whole day is already outside the off-publish sum.
Each day's subtraction stops at that day's real-interpreter rows, and any
surplus is reported as jobs_not_subtracted rather than taken from another day.
If the Actions API is unavailable, or its 45-day run inventory starts after the
window does, the net field reads unavailable with the reason; the raw figure is
never presented as the net one. No row is deleted.

downloads and mean_per_day keep their existing meaning as the raw figures, so
older snapshots and the classifier do not change under readers' feet. The net
sits beside them in net_of_own_ci, with the method and the upper-bound caveat
stated once on the parent block.

Fresh snapshot, retrieved 2026-10-05T10:21:28Z, series ends 2026-10-04. Seven
days: 19 raw real-interpreter events off the publish days, 2.7 a day; net of our
own CI 11, 1.6 a day. Thirty days: 36 raw, 1.4 a day; net 24, 0.9 a day. Four of
those 26 days have no pypistats rows, so both 30-day means are floors.

The exclusion does not explain the move the 2026-10-04 focus review reported.
The review's clean week, 2026-09-27 to 2026-10-03, reproduces from frozen feeds
at raw 15, 2.1 a day, net 11, 1.6 a day. Like for like, the week before is raw
12, 2.0 a day, net 8, 1.3 a day. Each window holds exactly one four-job run, so
the net figure rose. What this change does show is that the raw figure overstated
both weeks by four events, and that from 2026-10-03 the daily cron adds about
four a day to every window it touches: 8 of the 19 raw events in the committed
7-day window are our own jobs.

The net figure is still an upper bound: other CI, reinstalls and
mirrors-excluded tooling also use real interpreters, an interpreter does not
identify a person, and one job is not provably one download.

published-wheel.yml itself is unchanged. It is a real release check and it is
working.

Refs proof/published-wheel-ci-exclusion-20261005

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 5, 2026 10:30
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: f379b278-63aa-4e9a-86c9-8be52c766129)

@bmdhodl bmdhodl added the needs:patrick-review Requires Patrick personal review label Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

🤖 Claude review

Review

Process concern — .showwork session closed as blocked

.showwork/sessions/qw-20261004-activation-exclude-published-wheel-ci.jsonl shows two session.finish.refused events followed by "status": "blocked", "outcome": {"verdict": "UNVERIFIED"}. All 6 individual checks are GREEN, but the gate could not grant a clean close because acceptance requirements were not declared before the claims. The audit file calls this a "receipt discipline gap, not a work gap." Per the repo's CLAUDE.md showwork rules, REFUSED means a gap must be fixed or claims retracted — passing --status blocked is the fallback for genuinely stuck sessions, which this appears to be. Worth confirming with the owner before merging whether this close is acceptable or needs a follow-up session.


scripts/refresh_activation_snapshot.py

_net_of_own_ci — downloads can never go negative (OK), but the relationship between raw["downloads"] and rows_by_date is implicit

raw["downloads"] is computed by _off_publish() over the same interpreter_rows list, so the sum of non-publish-day values in rows_by_date should equal it. This invariant holds as long as callers pass the same interpreter_rows to both functions — which they do at build_snapshot:354. No bug, but a silent coupling. If that call site ever diverges, downloads could go silently wrong.

_DropAuthOnHostChange — strips only Authorization (scripts/refresh_activation_snapshot.py, new _DropAuthOnHostChange class)

Only Authorization is added to requests (no other secret headers), so stripping only that header on host/scheme changes is correct. The test at test_the_github_token_never_leaves_its_host_or_its_scheme covers same-host, other-host, and https→http cases. LGTM.

own_ci_reason is set before the sources dict is written, then the sources dict is mutated with windows_not_subtracted later (build_snapshot around line 381). This is fine — the dict is mutable and the mutation happens unconditionally when not_netted is non-empty.

published_wheel_runs — page-full coverage check

if len(runs) >= RUNS_PAGE_SIZE and oldest_seen is not None and oldest_seen > earliest:
    covers_from = oldest_seen

If the page is full but all runs predate earliest, oldest_seen < earliest, the condition is false, and covers_from = earliest (correct — no runs in the lookback means nothing to subtract). Edge case is handled correctly.


sdk/tests/test_activation_evidence.py

_mock_public_feeds now raises OSError for any GitHub API URL (line ~692). This is the right defensive default — it forces any test that wants the unavailable path to assert it explicitly, preventing silent pass-throughs. Good.

test_own_ci_net_figure_is_lower_than_the_raw_figure asserts "capped" in net["note"]. The actual note text in _net_of_own_ci ends with "capped at that day's rows" — the assertion holds. Fine.


No blocking issues in the core subtraction logic

The per-day cap (min(jobs, rows_by_date.get(day, 0))), publish-date exclusion, unavailable-on-partial-coverage, and graceful degradation when the Actions API is unreachable all look correct and are well-covered by the new tests.


Summary: One non-blocking process flag (showwork blocked close); no code bugs or security problems found. The token-stripping redirect handler is correct and tested. Recommend confirming the showwork close intent with the owner before merge.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs:patrick-review Requires Patrick personal review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants