Repository navigation
release: prepare 2.0.0 (hold merge until 2026-10-06) - #849
Conversation
Patrick authorized the 2.0.0 release on 2026-10-05 to go live on 2026-10-06. The source was already 2.0.0. This moves the changelog, README, generated PyPI README, docs, site and Agents SDK example from "candidate" to release wording, and records the approval and timing in memory/ and the roadmap. Dated 1.4.1 receipts keep their wording. Merge just before the v2.0.0 tag so GitHub and the site do not say "new in 2.0.0" while PyPI still serves 1.4.0. Proof in proof/v2.0.0/: release gates, 1592 passed / 3 skipped at 93% coverage, clean Python 3.11 install of the built wheel, Python 3.10 refusal, and Playwright checks of four pages at 375/768/1440. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: a445d8af-fe9d-496b-aacf-e42a271607c2) |
🤖 Claude reviewReview1. Business-sensitive outreach detail in a public repo —
|
…e tag-day price check - memory/decisions.md: remove the X/LinkedIn approval line (outreach detail stays out of this public repo). - proof/v2.0.0/verify.py: compare 10-clean-install.txt with the step count in clean_install.sh instead of a fixed 17. A new step without a rerun now fails with both counts in the message. - proof/v2.0.0/README.md: publish.yml runs the price-table age check before the build; run it again on tag day before pushing the tag. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: a1b55c32-8663-4131-b58c-5ba392887ef3) |
|
Addressed the Claude review in 3d2e718:
showwork The Copilot, Codex and Bugbot notices are usage-limit messages with no review content. 🤖 Addressed by Claude Code |
🤖 Claude reviewReviewThis is a documentation release-prep PR: converting "candidate" wording to release language across docs, site, CHANGELOG, and memory files. No SDK code changes. Evidence is extensive and the showwork ledger closed GREEN. No blocking issues. Three minor quality findings: 1.
|
…erify.py - clean_install.sh: exit 1 with a clear message outside Windows Git Bash (it uses the py launcher and Scripts/python.exe), before touching the work dir. - verify.py: match the Python 3.10 refusal with a regex, not 3.10.11. - proof README: a rebuild gives the same wheel hash; the sdist hash changes per build (setuptools mtimes, gzip header) with identical contents. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: c67ae1e9-09df-4024-bd7d-f9b367f61620) |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: a70b281a-aaf1-4611-bf19-058095eed555) |
|
Addressed the second Claude review in 1d72cfa (audit record in 563c68a):
Also from the rerun: the wheel hash is the same ( showwork 🤖 Addressed by Claude Code |
🤖 Claude reviewReview1.
|
🤖 Claude reviewPR Review: v2.0.0 Release PrepIssues Found1. The walkthrough now fetches -& .\.venv\Scripts\python.exe … --no-deps agentguard47==1.4.0
+& .\.venv\Scripts\python.exe … --no-deps agentguard47==2.0.0Any user following this guide after the PR merges but before the PyPI publish completes will get a resolution error ( 2. The corresponding session JSONL ( 3. check(re.search(r"\b1592 passed, 3 skipped\b", tests) is not None …, "full suite result changed")If any test is added, removed, or gains a new skip condition (e.g., Not Issues
|
- clean_install.sh: mkdir and both cd calls exit 1 on failure, so a failed cd cannot run the first-use steps from the repo root. No set -e: step must log the expected failing exit of the Python 3.10 install. - verify.py: flag 'not published yet' only on lines that mention 2.0.0. - .showwork audit report for v200-review-fixes-2 rewritten by showwork verify; it showed the intermediate RED run. The JSONL keeps the refused finish, the retraction and the GREEN finish. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Addressed the third Claude review in 4be010d:
showwork 🤖 Addressed by Claude Code |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: 758c867c-104a-4251-82f5-b816b3c93858) |
|
Reply to the fourth Claude review (it reviewed 563c68a; 4be010d landed three minutes later). No code change this round:
🤖 Addressed by Claude Code |
🤖 Claude reviewReviewSummaryThis is a documentation-only release prep PR that removes "candidate/unreleased" wording from docs, site, and README files ahead of the v2.0.0 tag. No SDK code changes. The proof artifacts are well-structured and the showwork audit trail is clean. Issues1. Merge timing regression in The walkthrough now installs from PyPI: ... install ... --no-deps agentguard47==2.0.0If this branch is merged before the PyPI upload completes, the documented walkthrough will produce a 2. check(re.search(r"\b1592 passed, 3 skipped\b", tests) is not None and "exit=0" in tests, "full suite result changed")Any test addition or deletion after this proof was captured will cause 3. work="$1"
...
rm -rf "$work"
Non-issues worth noting (not blocking)
The SDK is unchanged. The doc changes are consistent. The merge-timing risk (#1) is the only concern that could cause real user-visible harm; the others are code-quality observations. |
clean_install.sh now refuses to delete a work path that is not empty and has no marker from an earlier run, so a wrong argument such as . cannot remove the repo. verify.py notes that 09-test.txt is the saved suite run. Tested: ., a non-empty dir and a file are refused with exit 1 and left intact; a new path, an empty dir and a rerun on a marked dir each gave the same 17 step results as the saved log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Reply to the fifth Claude review, which covered 4be010d. I fixed it in fff367e:
showwork 🤖 Addressed by Claude Code |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: 8b53bf0f-d4ff-476a-a18e-100d4204ef0f) |
🤖 Claude reviewLGTM - no blocking issues. The PR is a documentation/text-only release-prep change — converting "2.0.0 candidate" / "unpublished 2.0.0" wording to release wording across docs, site, memory, and proof artifacts. No SDK logic changes. A few observations that are all intentional and well-documented:
Timing dependency —
The showwork sessions all close GREEN with verified acceptance requirements. No security, correctness, or quality concerns. |
Summary
Release prep for AgentGuard 2.0.0. Patrick authorized the release on 2026-10-05 to go live on 2026-10-06.
sdk/pyproject.toml,agentguard --versionand the release markers were already 2.0.0; this PR changes the public text from "candidate" to release wording.CHANGELOG.md: the 2.0.0 section no longer says "Unreleased candidate".README.mdand the generatedsdk/PYPI_README.md(this is the PyPI page): the Python 3.11 note, and thereceipt,hookandrunsections say "new in 2.0.0". The "not published yet" paragraph is gone.agentguard47==2.0.0, Python 3.11+), free local clients, OpenAI Responses (installsagentguard47==2.0.0from PyPI instead of a local wheel), enforcement boundary, compatibility, Vercel comparison, three discussion drafts, docs index.examples/openai_agents_sdk_budget.pydocstring.memory/state.md,memory/decisions.mdand the roadmap row record the approval and the timing.Not changed: the dated 1.4.1 candidate receipts in
docs/compatibility.md, and theinstall-intent-candidatemetric name on the site (activation tests and reports use it).Goal: the PyPI page, GitHub docs and site describe 2.0.0 as released at tag time. Scope: text only, plus proof. Non-goals: code changes, the tag, PyPI upload, social posts. Done: no candidate wording outside dated receipts, PyPI README in sync, release gates pass, the built wheel installs and runs in a clean Python 3.11 venv, pages pass at 375/768/1440, proof saved (
ops/04-DEFINITION_OF_DONE.md).Merge timing
Hold this PR until 2026-10-06, then merge just before the
v2.0.0tag. If it merges earlier, GitHub and the site say "new in 2.0.0" while PyPI still serves 1.4.0. I did not turn on auto-merge for this reason.After the merge, the runbook in
docs/RELEASING.mdtags the merge commit.publish.ymlpublishes to PyPI with Trusted Publishing, creates the GitHub Release, and dispatchesrelease-content.yml(subscriber email; social drafts only withqueue_social=true, which the dispatch does not set).Related Issues
Release gate left open by #831. Python 3.11 policy in
memory/decisions.md.Proof
Saved under
proof/v2.0.0/(Windows 11, Python 3.13.2 for the suite, Python 3.11 for the clean venv, Git Bash 5.2.37;makeis not installed, so each command ran directly).browser_check.pyon theorigin/mainsitepython scripts/sdk_preflight.pypython scripts/review_readiness_guard.pypython scripts/ci_tools_requirements_guard.pypython scripts/sdk_release_guard.py --check-price-table-ageruff check(thepublish.ymlset plus the changed example)bandit -r sdk/agentguard/ -s B101,B110,B112,B311 -qpytest sdk/tests/test_architecture.pypytest sdk/tests/ --cov=agentguard --cov-fail-under=80clean_install.sh: build likepublish.yml, install the wheel in a new Python 3.11 venv--versionprintsagentguard 2.0.0;doctor,demo,report,receipt,quickstart, raw starter, starterreport,hook claude-code --help,run --helpall exit 0requires a different Python: 3.10.11 not in '>=3.11'python proof/v2.0.0/browser_check.py(Playwright, Chromium)python proof/v2.0.0/verify.pyPrice-table note: the OpenAI and Google rows were checked on 2026-07-15. The
publish.ymlage gate passes until 2026-10-13 and fails from 2026-10-14.Not run here:
actionlintandshellcheckare not installed; no workflow changed.showwork session
v200-release-prep: VERIFIED, withverify.pyas the declared acceptance check.Review follow-up (3d2e718, 1d72cfa, 4be010d, fff367e)
memory/decisions.md(outreach detail stays out of this repo).verify.pycompares10-clean-install.txtwith thestepcount inclean_install.sh, not a fixed 17. Negative test: one extra step fails with both counts.proof/v2.0.0/README.md:publish.ymlruns the price-table age check before the build; run it again on tag day before pushing the tag.v200-review-fixes-outcome: VERIFIED (verify.pyis the declared acceptance check).v200-review-fixesclosed checks-only because I declared its acceptance check late.clean_install.shexits 1 outside Windows Git Bash.verify.pymatches any Python 3.10 patch release in the refusal check. The proof README says the sdist hash changes per build (same contents); the wheel hash is stable. showworkv200-review-fixes-2: VERIFIED.clean_install.shexits 1 ifmkdirorcdfails (noset -e, sostepstill logs the expected 3.10 refusal).verify.pyflags "not published yet" only on 2.0.0 lines. showwork rewrote thev200-review-fixes-2audit report to its final GREEN state. showworkv200-review-fixes-3: VERIFIED.verify.pyreads the saved09-test.txt, not a live suite. Thetry-release.md2.0.0 pin stays: the tag follows the merge (docs/RELEASING.mdsteps 7-8). Reply.clean_install.shdeletes the work path only if it is missing, empty or marked by an earlier run, so.cannot remove the repo.verify.pynotes that09-test.txtis the saved suite run. showworkv200-review-fixes-4: VERIFIED.Review Readiness
CHANGELOG.mdand the existing proof.Risk And Rollback
Low for code: no SDK code changes. The risk is timing (see Merge timing). If the tag publish fails after this merges, revert this commit to restore the candidate wording until the publish succeeds.
Scope
docs/RELEASING.md,ops/03-ROADMAP_NOW_NEXT_LATER.mdChecklist
make checkpasses: 1592 passed, 3 skipped, coverage 93%; ruff exit 0make structuralpasses (9 passed)make securitypasses (bandit, thepublish.ymlgate)verify.py__init__.pyexports changed,ops/02-ARCHITECTURE.mdupdated. N/A.🤖 Generated with Claude Code
Note
Low Risk
Text, proof, and verification metadata only—no SDK runtime changes; main risk is publishing/merge timing and misleading “released” wording before PyPI updates.
Overview
Prepares AgentGuard 2.0.0 for publication by treating it as a shipped release in public copy and recording machine-checkable proof that the tree is ready to tag.
CHANGELOG.mddrops “Unreleased candidate” language and frames 2.0.0 as the major release that supersedes the planned 1.4.1 patch (Python 3.11+ breaking notes unchanged)..showwork/adds claims, session logs, audit reports, and a tree snapshot forv200-release-prepand follow-upv200-review-fixes-*rounds. Declared acceptance ispython proof/v2.0.0/verify.pyexiting with Verified v2.0.0 release prep; sessions finish VERIFIED after review fixes (Git Bash-onlyclean_install.sh, safer workdir handling, dynamic step-count checks, regex-based Python 3.10 wheel refusal, scoped “not published yet” scans, proof README notes).proof/v2.0.0/gains saved transcripts for review-readiness, CI-tools, and release-guard (--check-price-table-age) runs alongside the existing release-prep bundle.Merge timing: hold until tag day so GitHub/site “2.0.0” messaging does not lead PyPI still on 1.4.0.
Reviewed by Cursor Bugbot for commit fff367e. Bugbot is set up for automated code reviews on this repo. Configure here.