Skip to content

release: prepare 2.0.0 (hold merge until 2026-10-06) - #849

Merged
bmdhodl merged 6 commits into
mainfrom
claude/release-v2-prep
Oct 6, 2026
Merged

bmdhodl merged 6 commits into
mainfrom
claude/release-v2-prep

Conversation

@bmdhodl

@bmdhodl bmdhodl commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Summary

Release prep for AgentGuard 2.0.0. Patrick authorized the release on 2026-10-05 to go live on 2026-10-06. sdk/pyproject.toml, agentguard --version and the release markers were already 2.0.0; this PR changes the public text from "candidate" to release wording.

  • CHANGELOG.md: the 2.0.0 section no longer says "Unreleased candidate".
  • README.md and the generated sdk/PYPI_README.md (this is the PyPI page): the Python 3.11 note, and the receipt, hook and run sections say "new in 2.0.0". The "not published yet" paragraph is gone.
  • Docs: getting started, Python 3.11 migration, try-release (pins agentguard47==2.0.0, Python 3.11+), free local clients, OpenAI Responses (installs agentguard47==2.0.0 from PyPI instead of a local wheel), enforcement boundary, compatibility, Vercel comparison, three discussion drafts, docs index.
  • Site: index labels, receipt note, hosted-tools rule, enforcement table row, two blog lines.
  • examples/openai_agents_sdk_budget.py docstring.
  • memory/state.md, memory/decisions.md and the roadmap row record the approval and the timing.

Not changed: the dated 1.4.1 candidate receipts in docs/compatibility.md, and the install-intent-candidate metric name on the site (activation tests and reports use it).

Goal: the PyPI page, GitHub docs and site describe 2.0.0 as released at tag time. Scope: text only, plus proof. Non-goals: code changes, the tag, PyPI upload, social posts. Done: no candidate wording outside dated receipts, PyPI README in sync, release gates pass, the built wheel installs and runs in a clean Python 3.11 venv, pages pass at 375/768/1440, proof saved (ops/04-DEFINITION_OF_DONE.md).

Merge timing

Hold this PR until 2026-10-06, then merge just before the v2.0.0 tag. If it merges earlier, GitHub and the site say "new in 2.0.0" while PyPI still serves 1.4.0. I did not turn on auto-merge for this reason.

After the merge, the runbook in docs/RELEASING.md tags the merge commit. publish.yml publishes to PyPI with Trusted Publishing, creates the GitHub Release, and dispatches release-content.yml (subscriber email; social drafts only with queue_social=true, which the dispatch does not set).

Related Issues

Release gate left open by #831. Python 3.11 policy in memory/decisions.md.

Proof

Saved under proof/v2.0.0/ (Windows 11, Python 3.13.2 for the suite, Python 3.11 for the clean venv, Git Bash 5.2.37; make is not installed, so each command ran directly).

Command Result
browser_check.py on the origin/main site 0 of 12 pass (every page shows "2.0.0 candidate")
python scripts/sdk_preflight.py exit 0
python scripts/review_readiness_guard.py exit 0
python scripts/ci_tools_requirements_guard.py exit 0
python scripts/sdk_release_guard.py --check-price-table-age passed
ruff check (the publish.yml set plus the changed example) exit 0
bandit -r sdk/agentguard/ -s B101,B110,B112,B311 -q exit 0
pytest sdk/tests/test_architecture.py 9 passed
PyPI README sync, release guard, release example, activation, first-run tests 94 passed
pytest sdk/tests/ --cov=agentguard --cov-fail-under=80 1592 passed, 3 skipped; coverage 93%
clean_install.sh: build like publish.yml, install the wheel in a new Python 3.11 venv --version prints agentguard 2.0.0; doctor, demo, report, receipt, quickstart, raw starter, starter report, hook claude-code --help, run --help all exit 0
same wheel in a Python 3.10 venv refused: requires a different Python: 3.10.11 not in '>=3.11'
python proof/v2.0.0/browser_check.py (Playwright, Chromium) 12 of 12 pass: new labels shown, no candidate text, no horizontal overflow at 375/768/1440
python proof/v2.0.0/verify.py Verified v2.0.0 release prep

Price-table note: the OpenAI and Google rows were checked on 2026-07-15. The publish.yml age gate passes until 2026-10-13 and fails from 2026-10-14.

Not run here: actionlint and shellcheck are not installed; no workflow changed.

showwork session v200-release-prep: VERIFIED, with verify.py as the declared acceptance check.

Review follow-up (3d2e718, 1d72cfa, 4be010d, fff367e)

  • Removed the X/LinkedIn approval line from memory/decisions.md (outreach detail stays out of this repo).
  • verify.py compares 10-clean-install.txt with the step count in clean_install.sh, not a fixed 17. Negative test: one extra step fails with both counts.
  • proof/v2.0.0/README.md: publish.yml runs the price-table age check before the build; run it again on tag day before pushing the tag.
  • showwork v200-review-fixes-outcome: VERIFIED (verify.py is the declared acceptance check). v200-review-fixes closed checks-only because I declared its acceptance check late.
  • Second review (1d72cfa, audit record 563c68a): clean_install.sh exits 1 outside Windows Git Bash. verify.py matches any Python 3.10 patch release in the refusal check. The proof README says the sdist hash changes per build (same contents); the wheel hash is stable. showwork v200-review-fixes-2: VERIFIED.
  • Third review (4be010d): clean_install.sh exits 1 if mkdir or cd fails (no set -e, so step still logs the expected 3.10 refusal). verify.py flags "not published yet" only on 2.0.0 lines. showwork rewrote the v200-review-fixes-2 audit report to its final GREEN state. showwork v200-review-fixes-3: VERIFIED.
  • Fourth review (it covered 563c68a): no code change. The audit report was already GREEN in 4be010d. verify.py reads the saved 09-test.txt, not a live suite. The try-release.md 2.0.0 pin stays: the tag follows the merge (docs/RELEASING.md steps 7-8). Reply.
  • Fifth review (fff367e): clean_install.sh deletes the work path only if it is missing, empty or marked by an earlier run, so . cannot remove the repo. verify.py notes that 09-test.txt is the saved suite run. showwork v200-review-fixes-4: VERIFIED.

Review Readiness

  • Public positioning claims have a source/fact ledger. The wording changes only the release status; the facts come from CHANGELOG.md and the existing proof.
  • State, lock, file, or process-concurrency changes include cross-platform failure proof. N/A.
  • External API collectors include response-shape, pagination, null, and partial-failure tests. N/A.
  • Proof artifacts include command, exit code, platform, and regenerated-after-review status. Yes; not yet regenerated after review.
  • Workflow changes explain trigger scope, timeouts, concurrency, artifacts, and spend impact. N/A, no workflow change.

Risk And Rollback

Low for code: no SDK code changes. The risk is timing (see Merge timing). If the tag publish fails after this merges, revert this commit to restore the candidate wording until the publish succeeds.

Scope

  • Related ops/doc(s): docs/RELEASING.md, ops/03-ROADMAP_NOW_NEXT_LATER.md
  • Does this change the public API? No.
  • Does this shift roadmap priority? No. It records the approved release.

Checklist

  • make check passes: 1592 passed, 3 skipped, coverage 93%; ruff exit 0
  • make structural passes (9 passed)
  • make security passes (bandit, the publish.yml gate)
  • New functionality has tests: N/A (text only); Playwright page checks and verify.py
  • No new hard dependencies in core SDK
  • No hardcoded absolute paths
  • If __init__.py exports changed, ops/02-ARCHITECTURE.md updated. N/A.

🤖 Generated with Claude Code


Note

Low Risk
Text, proof, and verification metadata only—no SDK runtime changes; main risk is publishing/merge timing and misleading “released” wording before PyPI updates.

Overview
Prepares AgentGuard 2.0.0 for publication by treating it as a shipped release in public copy and recording machine-checkable proof that the tree is ready to tag.

CHANGELOG.md drops “Unreleased candidate” language and frames 2.0.0 as the major release that supersedes the planned 1.4.1 patch (Python 3.11+ breaking notes unchanged).

.showwork/ adds claims, session logs, audit reports, and a tree snapshot for v200-release-prep and follow-up v200-review-fixes-* rounds. Declared acceptance is python proof/v2.0.0/verify.py exiting with Verified v2.0.0 release prep; sessions finish VERIFIED after review fixes (Git Bash-only clean_install.sh, safer workdir handling, dynamic step-count checks, regex-based Python 3.10 wheel refusal, scoped “not published yet” scans, proof README notes).

proof/v2.0.0/ gains saved transcripts for review-readiness, CI-tools, and release-guard (--check-price-table-age) runs alongside the existing release-prep bundle.

Merge timing: hold until tag day so GitHub/site “2.0.0” messaging does not lead PyPI still on 1.4.0.

Reviewed by Cursor Bugbot for commit fff367e. Bugbot is set up for automated code reviews on this repo. Configure here.

Patrick authorized the 2.0.0 release on 2026-10-05 to go live on
2026-10-06. The source was already 2.0.0. This moves the changelog,
README, generated PyPI README, docs, site and Agents SDK example from
"candidate" to release wording, and records the approval and timing in
memory/ and the roadmap. Dated 1.4.1 receipts keep their wording.

Merge just before the v2.0.0 tag so GitHub and the site do not say
"new in 2.0.0" while PyPI still serves 1.4.0.

Proof in proof/v2.0.0/: release gates, 1592 passed / 3 skipped at 93%
coverage, clean Python 3.11 install of the built wheel, Python 3.10
refusal, and Playwright checks of four pages at 375/768/1440.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 00:23
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@cursor

cursor Bot commented Oct 6, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: a445d8af-fe9d-496b-aacf-e42a271607c2)

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🤖 Claude review

Review

1. Business-sensitive outreach detail in a public repo — memory/decisions.md

+- Each X or LinkedIn post still needs his approval in chat.

CLAUDE.md explicitly states: "do not put business-sensitive plans or outreach data in this repo." Social-media approval workflow is outreach data. This line should be removed from the committed file (it belongs in a private channel or local note).


2. Price-table gate expires in 8 days — operational risk

proof/v2.0.0/README.md documents this:

OpenAI and Google rows were checked on 2026-07-15, so the publish.yml gate fails from 2026-10-14 (more than 90 days). Tag before then, or re-check those prices first.

Today is 2026-10-06. The window is narrow. If the publish slips past 2026-10-14, release-guard will block the publish.yml run. This is already noted, but it's worth a hard reminder: re-run sdk_release_guard.py --check-price-table-age on tag day before pushing.


3. verify.py hardcodes the exit-code count — fragile but not wrong

proof/v2.0.0/verify.py:44:

check(len(exits) == 17 and set(exits[:-1]) == {"0"} and exits[-1] == "1", ...)

This passes today (17 step calls in clean_install.sh, last one intentionally exits 1). If someone adds a step to the script without regenerating the proof, this silently fails with a misleading message. Not a blocking bug, but len(exits) >= 17 with an explicit comment explaining the expected structure would be more maintainable.


Everything else looks correct

  • All "candidate / unpublished" wording swept consistently across docs, site, examples, CHANGELOG, PyPI README, and memory files.
  • showwork session correctly went through four refused finish calls (unverified claims and undeclared changes) before reaching GREEN — the ledger is honest.
  • verify.py's STALE regex won't false-positive on the intentionally-retained "1.4.1 candidate receipts" in docs/compatibility.md.
  • Clean-install proof confirms Python 3.10 rejects the wheel (requires a different Python: 3.10.11 not in '>=3.11').
  • Browser checks pass 12/12 after the site changes.

Blocking: item 1 (outreach data in public repo, violates CLAUDE.md). Items 2–3 are non-blocking but worth acting on before publish day.

…e tag-day price check

- memory/decisions.md: remove the X/LinkedIn approval line (outreach detail
  stays out of this public repo).
- proof/v2.0.0/verify.py: compare 10-clean-install.txt with the step count
  in clean_install.sh instead of a fixed 17. A new step without a rerun now
  fails with both counts in the message.
- proof/v2.0.0/README.md: publish.yml runs the price-table age check before
  the build; run it again on tag day before pushing the tag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor

cursor Bot commented Oct 6, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: a1b55c32-8663-4131-b58c-5ba392887ef3)

@bmdhodl

bmdhodl commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Addressed the Claude review in 3d2e718:

  1. Outreach detail (blocking): removed the X/LinkedIn approval line from memory/decisions.md.
  2. Price-table gate: publish.yml already runs --check-price-table-age before the build, so a late tag stops before any PyPI upload. proof/v2.0.0/README.md now says to run the check again on tag day before pushing the tag. The gate passes today, and the tag is planned for 2026-10-06.
  3. Fixed step count: verify.py now counts the step calls in clean_install.sh and compares that count with 10-clean-install.txt. I kept the check exact instead of >= 17, because >= would pass a stale proof. Negative test: one extra step gives 10-clean-install.txt has 17 results but clean_install.sh has 18 steps; rerun clean_install.sh (exit 1). The real tree gives Verified v2.0.0 release prep.

showwork v200-review-fixes-outcome: VERIFIED, with verify.py as the declared acceptance check. Lint, preflight, review-readiness and release guards pass.

The Copilot, Codex and Bugbot notices are usage-limit messages with no review content.

🤖 Addressed by Claude Code

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🤖 Claude review

Review

This is a documentation release-prep PR: converting "candidate" wording to release language across docs, site, CHANGELOG, and memory files. No SDK code changes. Evidence is extensive and the showwork ledger closed GREEN.

No blocking issues. Three minor quality findings:


1. proof/v2.0.0/clean_install.sh is Windows-only without declaration

File: proof/v2.0.0/clean_install.sh:17,18,23

py311="$work/venv311/Scripts/python.exe"   # Windows path
step py -3.11 -m venv "$work/venv311"       # Windows Python Launcher

The script is a .sh file but unconditionally uses Scripts/python.exe and the py launcher. On Linux/Mac it will silently pass the step commands (which log exit=?) but produce wrong results. The proof README documents "Platform: Windows 11" but the script header only says #!/usr/bin/env bash. Someone trying to reproduce on Linux would get a misleading partial run.

Suggestion: Add a guard at the top: [[ "$OSTYPE" == "msys" || "$OSTYPE" == "cygwin" ]] || { echo "This script targets Windows/Git Bash"; exit 1; }.


2. proof/v2.0.0/verify.py pins exact Python patch version in an assertion

File: proof/v2.0.0/verify.py:56

check("requires a different Python: 3.10.11 not in '>=3.11'" in log,
      "Python 3.10 did not refuse the wheel")

This string includes 3.10.11. If the proof is ever re-run against Python 3.10.9 or 3.10.12, this check silently fails even though the rejection behavior is correct. The meaningful part is not in '>=3.11'; the patch version is incidental.

Suggestion: Match re.search(r"requires a different Python: 3\.10\.\d+ not in '>=3\.11'", log) instead.


3. Price table expires 2026-10-14 — 8 days from today

File: proof/v2.0.0/README.md (documented)

OpenAI and Google rows were checked on 2026-07-15, so the publish.yml gate fails from 2026-10-14 (more than 90 days). Tag before then, or re-check those prices first.

The proof README calls this out clearly and publish.yml gates on it before any upload, so publication cannot silently proceed past the deadline. No action needed here beyond what's already documented, but flagging it as a merge-day dependency.


Everything else is consistent: all candidate/unpublished wording is removed from public-facing files, the STALE regex in verify.py catches the right patterns, the 10-clean-install.txt correctly shows Python 3.10 refusing the wheel, and browser checks pass 12/12.

…erify.py

- clean_install.sh: exit 1 with a clear message outside Windows Git Bash
  (it uses the py launcher and Scripts/python.exe), before touching the
  work dir.
- verify.py: match the Python 3.10 refusal with a regex, not 3.10.11.
- proof README: a rebuild gives the same wheel hash; the sdist hash changes
  per build (setuptools mtimes, gzip header) with identical contents.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor

cursor Bot commented Oct 6, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: c67ae1e9-09df-4024-bd7d-f9b367f61620)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cursor

cursor Bot commented Oct 6, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: a70b281a-aaf1-4611-bf19-058095eed555)

@bmdhodl

bmdhodl commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Addressed the second Claude review in 1d72cfa (audit record in 563c68a):

  1. Windows-only script: clean_install.sh now exits 1 with a clear message outside Windows Git Bash, before it touches the work dir. Tested with OSTYPE=linux-gnu: exit 1, and no work dir was created. On Git Bash the full rerun gives the same 17 results as 10-clean-install.txt. One correction: on Linux the old script would not pass quietly. py would be missing, so the steps would log exit=127. The guard still makes the failure clear.
  2. Patch version pin: verify.py now matches requires a different Python: 3\.10\.\d+ not in '>=3\.11'. Tested: a 3.10.9 log passes, and a 3.11.0 log fails.
  3. Price table: no change. publish.yml checks the age before the build, and the tag is planned for 2026-10-06.

Also from the rerun: the wheel hash is the same (a434b301…). The sdist hash changes with each build, because setuptools stamps build-time mtimes and the gzip header has a time. All 151 members have the same contents. proof/v2.0.0/README.md now says this, so the CI hash comparison does not mislead.

showwork v200-review-fixes-2: VERIFIED. The acceptance check was declared first. One claim pattern was mis-escaped, so I retracted it and recorded it again.

🤖 Addressed by Claude Code

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🤖 Claude review

Review

1. proof/v2.0.0/clean_install.sh — cd not guarded, no set -e

clean_install.sh:18,35 — The two cd calls are not wrapped in step() and there is no set -e:

cd "$work/run"
step "$py311" -I -m agentguard --version
...
cd "$repo"

If cd "$work/run" silently fails (e.g. $work resolves to a path where mkdir -p also failed), all subsequent step calls run from the repo root. Demo/doctor/receipt would write files there, exit 0 anyway, and verify.py would see the expected pass pattern — giving a false green. set -e or cd ... || exit 1 would catch this.

2. .showwork/audit-session-v200-review-fixes-2.md — committed as UNVERIFIED

The audit file is committed with:

**Check verdict: RED**  (3/4 checks passed)
**Outcome: UNVERIFIED**

The session JSONL (sessions/v200-review-fixes-2.jsonl) confirms the session ultimately finished GREEN after re-expressing the file_contains claim as a re.search check. The audit MD is a snapshot from an intermediate refused finish. Committing an artifact that says UNVERIFIED alongside a JSONL that shows VERIFIED is misleading to anyone auditing the ledger — the audit file should reflect the resolved state or include a header noting it is a superseded snapshot.

3. Minor — "not published yet" STALE pattern is broad

proof/v2.0.0/verify.py:4

STALE = re.compile(r"2\.0\.0 candidate|unpublished (AgentGuard )?(\*\*)?2\.0\.0|not published yet"

The not published yet branch matches any occurrence of that phrase in the scanned paths, not just references to the 2.0.0 candidate. It passes today, but will silently reject legitimate future docs (e.g. describing a planned feature) without an obvious error. Scoping it to 2\.0\.0.*not published yet or dropping it (the adjacent patterns already cover the candidate wording) would be safer.


No issues in the SDK source, docs text, or site HTML changes — those are consistent and correct. Items 1 and 2 are the actionable findings; item 3 is low-priority.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🤖 Claude review

PR Review: v2.0.0 Release Prep

Issues Found


1. docs/guides/try-release.md — functional breakage between merge and publish

The walkthrough now fetches agentguard47==2.0.0 directly from pypi.org/simple:

-& .\.venv\Scripts\python.exe … --no-deps agentguard47==1.4.0
+& .\.venv\Scripts\python.exe … --no-deps agentguard47==2.0.0

Any user following this guide after the PR merges but before the PyPI publish completes will get a resolution error (Could not find a version that satisfies the requirement agentguard47==2.0.0). proof/v2.0.0/README.md says "Merge this PR on 2026-10-06, just before the v2.0.0 tag", but that creates a fragile race between merge and publish.yml completing. If CI publish takes more than a few minutes — or is manually triggered after merge — the guide is broken for real users during that window. Consider either (a) keeping the guide at 1.4.0 and updating it in a second commit immediately after the PyPI upload is confirmed, or (b) adding a visible "requires 2.0.0 to be published" note so the guide fails loudly rather than silently misleading users.


2. .showwork/audit-session-v200-review-fixes-2.md — committed in UNVERIFIED/RED state

**Check verdict: RED**  (3/4 checks passed)
**Outcome: UNVERIFIED**

The corresponding session JSONL (v200-review-fixes-2.jsonl) confirms the session did ultimately finish GREEN/VERIFIED, so the underlying data is fine. But this audit file captures an intermediate state from before the claim was retracted and re-recorded. A reader looking only at the audit file gets the wrong picture. Since the repo treats .showwork/ as an append-only ledger, this can't be edited — it's worth at minimum adding a note in the PR description that this file reflects an intermediate run, with the final state in the JSONL.


3. proof/v2.0.0/verify.py — hardcoded test count

check(re.search(r"\b1592 passed, 3 skipped\b", tests) is not None …, "full suite result changed")

If any test is added, removed, or gains a new skip condition (e.g., agents module becomes available in CI), this re-verification check will fail spuriously. Acceptable for a one-time proof artifact, but verify.py is positioned as a rerunnable checker — the hardcoded count makes it a one-shot artifact instead. Lower severity but worth noting.


Not Issues

  • The clean_install.sh step() function not propagating exit codes: intentional — the last step must fail, so set -e is deliberately absent.
  • The CSS-case mismatch in browser_check.py ("Claude Code NEW IN 2.0.0" vs HTML "new in 2.0.0"): explicitly documented as CSS text-transform, confirmed passing in 11-browser.txt.
  • The verify.py \d+ patch-version regex fix: correctly handles any 3.10.x build.

- clean_install.sh: mkdir and both cd calls exit 1 on failure, so a failed
  cd cannot run the first-use steps from the repo root. No set -e: step
  must log the expected failing exit of the Python 3.10 install.
- verify.py: flag 'not published yet' only on lines that mention 2.0.0.
- .showwork audit report for v200-review-fixes-2 rewritten by showwork
  verify; it showed the intermediate RED run. The JSONL keeps the refused
  finish, the retraction and the GREEN finish.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bmdhodl

bmdhodl commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Addressed the third Claude review in 4be010d:

  1. Unguarded setup in clean_install.sh: mkdir and both cd calls now exit 1 on failure. I did not add set -e, because step must log the expected failing exit of the Python 3.10 install. A comment in the script says why. Tests:
    • With cd forced to fail for the run dir, the script exits 1 after the first 5 steps. No first-use step runs, and no files land in the repo root.
    • With the work dir under a file, mkdir fails and the script exits 1 before any step.
    • A full rerun on Git Bash gives the same 17 results as 10-clean-install.txt.
  2. Audit file showed RED: showwork verify rewrites that report on each run, and the finish receipt does not pin it. My diagnostic verify wrote it during the refused close. I let showwork verify write it again: it now shows GREEN, 5/5, VERIFIED. I did not edit it by hand. The append-only JSONL still holds the refused finish, the retraction and the GREEN finish.
  3. Broad pattern: not published yet now matches only on lines that also mention 2.0.0 (2\.0\.0.*not published yet). The old README line also had "2.0.0 candidate", so nothing is lost. Test: in a temporary doc, the 2.0.0 line is flagged and a line about another feature is not.

showwork v200-review-fixes-3: VERIFIED. The acceptance check was declared first. Lint, bash syntax, preflight and review-readiness pass.

🤖 Addressed by Claude Code

@cursor

cursor Bot commented Oct 6, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: 758c867c-104a-4251-82f5-b816b3c93858)

@bmdhodl

bmdhodl commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Reply to the fourth Claude review (it reviewed 563c68a; 4be010d landed three minutes later). No code change this round:

  1. try-release.md gap between merge and publish: no change. docs/RELEASING.md steps 7 and 8 say to merge the release-prep PR, then tag the merge commit. I will push the tag right after the merge, so the gap lasts only while publish.yml runs. During that gap, pip fails loudly with "no matching distribution"; it does not mislead anyone quietly. Option (a) would leave the guide on 1.4.0 after 2.0.0 is live. Option (b) adds a note that goes stale within minutes. Every "new in 2.0.0" label in this PR has the same gap, so the PR body says to merge it just before the tag.
  2. Audit file in RED state: already fixed in 4be010d. showwork verify rewrote .showwork/audit-session-v200-review-fixes-2.md, and it now shows GREEN, 5/5, VERIFIED. That report is rewritten on each verify run, and the finish receipt does not pin it; audit-session-agentguard-scorecard-pin-20260816.md has 11 commits for the same reason. The append-only record is the JSONL, which keeps the refused finish, the retraction and the GREEN finish.
  3. Fixed test count: no change. verify.py:61-62 reads the saved 09-test.txt transcript; it does not run pytest. New tests or skips do not affect it. The check keeps that transcript matched to the "1592 passed, 3 skipped" numbers in the proof README and this PR body.

🤖 Addressed by Claude Code

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🤖 Claude review

Review

Summary

This is a documentation-only release prep PR that removes "candidate/unreleased" wording from docs, site, and README files ahead of the v2.0.0 tag. No SDK code changes. The proof artifacts are well-structured and the showwork audit trail is clean.


Issues

1. Merge timing regression in docs/guides/try-release.md

The walkthrough now installs from PyPI:

... install ... --no-deps agentguard47==2.0.0

If this branch is merged before the PyPI upload completes, the documented walkthrough will produce a pip 404 error for any user who follows it. The proof README acknowledges this ("If it merges earlier, GitHub and the site say 'new in 2.0.0' while PyPI still serves 1.4.0") but there is no automated guard against an early merge. This is a real user-facing breakage window, not just a process note.

2. proof/v2.0.0/verify.py:57 — hardcoded test count is a latent false alarm

check(re.search(r"\b1592 passed, 3 skipped\b", tests) is not None and "exit=0" in tests, "full suite result changed")

Any test addition or deletion after this proof was captured will cause verify.py to fail with "full suite result changed" on the next python proof/v2.0.0/verify.py run, even if nothing is wrong with the release. Since this file ships with the PR and the repo has active test development, this will produce noise immediately after the release. The other step counts in clean_install.sh are computed dynamically; this one is not.

3. proof/v2.0.0/clean_install.sh — rm -rf "$work" with no path safety check

work="$1"
...
rm -rf "$work"

set -u catches the unset case but not a caller passing . or a parent directory. For a proof script run manually this is low risk, but worth noting given the rm -rf with no validation beyond set -u.


Non-issues worth noting (not blocking)

  • The STALE regex in verify.py correctly avoids false-positives from 1.4.1 candidate references in docs/compatibility.md (the pattern anchors on 2\.0\.0).
  • The step function design (no set -e, explicit exit logging, last step expected to fail) is intentional and documented.
  • The showwork ledger is append-only with proper retractions; the retraction-then-reclaim pattern in v200-release-prep.jsonl is messy but valid.

The SDK is unchanged. The doc changes are consistent. The merge-timing risk (#1) is the only concern that could cause real user-visible harm; the others are code-quality observations.

clean_install.sh now refuses to delete a work path that is not empty and
has no marker from an earlier run, so a wrong argument such as . cannot
remove the repo. verify.py notes that 09-test.txt is the saved suite run.

Tested: ., a non-empty dir and a file are refused with exit 1 and left
intact; a new path, an empty dir and a rerun on a marked dir each gave
the same 17 step results as the saved log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bmdhodl

bmdhodl commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Reply to the fifth Claude review, which covered 4be010d. I fixed it in fff367e:

  1. rm -rf "$work" with no path check: fixed. clean_install.sh now deletes the work path only if the path is missing, is an empty directory, or has the .clean-install-workdir marker from an earlier run. In any other case it exits 1. Tests: . from a non-empty directory, a non-empty directory and a regular file were each refused with exit 1, and nothing was changed. A new path, an empty directory and a second run on a marked directory each gave the same 17 step results as 10-clean-install.txt (16 passes, then the Python 3.10 refusal).
  2. Fixed test count: I did not change the check, but I added a comment. verify.py reads 09-test.txt, which is the suite run saved for this release. It does not run pytest, so new tests do not change the result. The comment above the check now says this.
  3. Merge timing: no change; my answer is in the previous reply. Auto-merge is off, and the owner merges by hand right before the tag. That manual step is the guard. A CI check that fails until PyPI has 2.0.0 cannot work, because docs/RELEASING.md tags the merge commit, so the merge must come first.

showwork v200-review-fixes-4: VERIFIED (GREEN 5/5).

🤖 Addressed by Claude Code

@cursor

cursor Bot commented Oct 6, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: 8b53bf0f-d4ff-476a-a18e-100d4204ef0f)

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🤖 Claude review

LGTM - no blocking issues.

The PR is a documentation/text-only release-prep change — converting "2.0.0 candidate" / "unpublished 2.0.0" wording to release wording across docs, site, memory, and proof artifacts. No SDK logic changes.

A few observations that are all intentional and well-documented:

proof/v2.0.0/clean_install.sh — set -u without set -e is deliberate: each step() call records its own exit code and the final Python 3.10 install is expected to fail (exit=1). The script correctly guards against destroying a non-empty pre-existing work dir via the marker-file pattern.

proof/v2.0.0/verify.py — The STALE regex correctly uses re.IGNORECASE and the Python 3.10 refusal check uses r"requires a different Python: 3\.10\.\d+" rather than a pinned patch version, which is the fix the v200-review-fixes-2 session records. The step-count comparison between clean_install.sh and 10-clean-install.txt is sound.

Timing dependency — docs/guides/try-release.md now pins agentguard47==2.0.0 which doesn't exist on PyPI yet. This is explicitly called out in proof/v2.0.0/README.md: merge just before the v2.0.0 tag on 2026-10-06. The price-table age deadline (2026-10-14) is also flagged there.

proof/v2.0.0/browser_check.py — Expects CSS-capitalized text ("Claude Code NEW IN 2.0.0") via Playwright's inner_text(), which renders transforms. The 00-before-browser.txt baseline confirms the pre-PR site failed all 12 checks; 11-browser.txt confirms 12/12 pass post-change.

The showwork sessions all close GREEN with verified acceptance requirements. No security, correctness, or quality concerns.

@bmdhodl
bmdhodl merged commit e4114e0 into main Oct 6, 2026
18 checks passed
@bmdhodl bmdhodl mentioned this pull request Oct 6, 2026
5 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants