Skip to content

docs: fix stale and broken agent instructions from prompt audit - #842

Merged
bmdhodl merged 1 commit into
mainfrom
claude/prompt-audit-checkup-38e6e6
Oct 5, 2026
Merged

bmdhodl merged 1 commit into
mainfrom
claude/prompt-audit-checkup-38e6e6

Conversation

@bmdhodl

@bmdhodl bmdhodl commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

What

Fixes stale and broken agent instructions found by a prompt audit of the Claude Code configuration (target: Claude Opus 5.5). Docs and agent config only; no SDK code changes.

Broken references, now corrected

  • AGENTS.md: agent prompt paths pointed to .Codex/agents/, which does not exist. They now point to .claude/agents/.
  • AGENTS.md: the release comment named PYPI_TOKEN. publish.yml uses Trusted Publishing (OIDC).
  • AGENTS.md: replaced the old module graph and module table with a pointer to ARCHITECTURE.md. The table listed 3 CLI subcommands; there are 13.
  • .claude/agents/sdk-dev.md: fixed the GOLDEN_PRINCIPLES.md link, RateLimitExceeded → RetryLimitExceeded, and MODEL_PRICES → DEFAULT_PRICE_TABLE in price_table.py.
  • .claude/agents/pm.md: removed v1.2.6 as the latest release (state lives in memory/state.md).
  • .agents/skills/next-ticket/SKILL.md: removed the AG-06 hold (AG-06: Add a supported OpenAI Responses and Agents SDK integration #735 and AG-06 remains held: #735 was closed by a release-prep reference #767 are closed).

Conflicts between instruction files, older text rewritten to match newer

  • pm.md and marketing.md: the 3-phase plan and Phase 1 launch table now defer to AgentGuard weekly growth plan: verified limits, portable evidence, and adoption #729 (per SKILL.md:20 and pm.md:49).
  • marketing.md: dropped "expand to full observability" (CLAUDE.md:74 avoids it). Wedge and audience now match memory/distribution.md and docs/enforcement-boundary.md.
  • dashboard-dev.md: removed non-public plan prices (marketing.md:51, site/compare.html:146).
  • AGENTS.md: the staleness check reads the root ARCHITECTURE.md, as CLAUDE.md does.

Roster

  • Added name/description frontmatter to sdk-dev.md and pm.md, so Claude Code registers them as subagents (CLAUDE.md lists them as Claude-side assets).

Wording

  • Removed shouted capitals in the review-loop rules (CLAUDE.md, AGENTS.md) and the contract heading. The rules are unchanged.
  • Dropped the undated 93% coverage figure; the enforced 80% floor stays.

Follow-ups added to ops/FOLLOWUP.md: the root report files, and whether ops-cadence.yml and the PR template should track the root ARCHITECTURE.md.

Proof

Saved under proof/prompt-audit-cleanup/.

Check Result
python scripts/sdk_preflight.py exit 0 (no SDK-relevant changes)
python scripts/ci_tools_requirements_guard.py exit 0
python scripts/review_readiness_guard.py exit 0
ruff check (Makefile lint paths) exit 0
pytest sdk/tests/test_architecture.py 9 passed
bandit -r sdk/agentguard/ -s B101,B110,B112,B311 -q exit 0
python scripts/sdk_release_guard.py Release guard passed
python scripts/generate_pypi_readme.py --check exit 0
pytest sdk/tests/ --cov=agentguard --cov-fail-under=80 1591 passed, 3 skipped; coverage 92.50%
python proof/prompt-audit-cleanup/verify.py Verified prompt-audit cleanup (fails if a fixed stale fact returns)

make is not installed on this Windows host, so each target ran directly. make mcp was not run: no mcp-server/ file changed. CI runs it.

showwork: prompt-audit-cleanup closed checks-only (24/24 green). prompt-audit-acceptance declares verify.py as its acceptance check and closes on the outcome.

The release-guard markers in AGENTS.md, CLAUDE.md and sdk-dev.md are unchanged.

Not in this PR

  • Global ~/.claude conflicts (auto-merge vs the comment loop, pnpm test:e2e). They are outside the repo and reported to the owner.
  • Moving QA_REPORT.md, WORK_PLAN.md, RESEARCH.md out of the root.

🤖 Generated with Claude Code


Note

Low Risk
Documentation and agent-config only; changes how agents are instructed, not product behavior, auth, or data paths.

Overview
Aligns Claude/agent instruction files with current repo facts so automated agents stop following stale paths, wrong SDK names, and superseded launch plans. No runtime SDK or dashboard code changes.

Reference and roster fixes: AGENTS.md and related docs now point at .claude/agents/, Trusted Publishing instead of PYPI_TOKEN, and root ARCHITECTURE.md for staleness. sdk-dev.md corrects guard/price guidance (RetryLimitExceeded, DEFAULT_PRICE_TABLE) and the GOLDEN_PRINCIPLES.md link. pm.md and marketing.md drop the old 3-phase launch tables and defer ordering to GitHub #729; marketing positioning tightens to runtime guardrails (not “full observability”). dashboard-dev.md and the next-ticket skill remove non-public prices and the closed AG-06 hold. sdk-dev and pm subagents get name/description frontmatter.

Proof and follow-ups: Adds proof/prompt-audit-cleanup/ (saved check outputs plus verify.py to re-assert fixes), showwork session/claim receipts, a small CLAUDE.md review-loop wording tweak, and prompt-audit leftovers logged in ops/FOLLOWUP.md.

Reviewed by Cursor Bugbot for commit 1bb18b4. Bugbot is set up for automated code reviews on this repo. Configure here.

A prompt audit of the Claude Code configuration found broken references,
stale release facts, and instruction files that contradict each other.

- AGENTS.md: agent prompts live in .claude/agents/, publishing uses
  Trusted Publishing, the module map points to ARCHITECTURE.md, and the
  staleness check reads the root ARCHITECTURE.md.
- sdk-dev.md: fix the GOLDEN_PRINCIPLES.md link and the SDK symbol names;
  add subagent frontmatter (also to pm.md).
- pm.md, marketing.md: defer work order to #729 instead of the old phase plan.
- dashboard-dev.md: drop non-public plan prices.
- next-ticket skill: drop the AG-06 hold (#735 and #767 are closed).
- CLAUDE.md, AGENTS.md: remove shouted capitals; rules are unchanged.

Proof and a re-check script are in proof/prompt-audit-cleanup/.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 5, 2026 02:22

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-05T02:26:07.756570Z 1bb18b4 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@cursor

cursor Bot commented Oct 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: 609dec97-418a-43a3-b0e0-ef840265e399)

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

🤖 Claude review

Review

This is a docs/config-only cleanup with no SDK code changes. The proof chain is complete and all 24 checks passed. A few items worth noting:

Non-blocking inconsistencies

AGENTS.md — incomplete "Codex" rename (AGENTS.md:152, 164, 168)

The agent-file paths were correctly updated to .claude/agents/, but the surrounding prose still says "Codex agents", "Open Codex in this repo", and "Every Codex session". The rename is half done. These should either all move to "Claude agents" or stay as-is; the current mix confuses readers about which system owns the prompts.

ARCHITECTURE.md path inconsistency (ops/FOLLOWUP.md:52–57)

AGENTS.md now runs the staleness check against ARCHITECTURE.md (root), but ops-cadence.yml and the PR template still reference ops/02-ARCHITECTURE.md. The FOLLOWUP.md entry correctly calls this out, but leaving it unresolved means agents using different entrypoints will check different (possibly stale) files. This should be resolved, not just tracked.

Minor concern in verify.py

require("1591 passed, 3 skipped" in text("proof/prompt-audit-cleanup/09-test.txt"), ...)

This hardcodes a test count that will break as soon as any test is added or removed — even in a completely unrelated PR. A verify.py run on main after any test change would fail. Consider matching on the coverage summary line only (Required test coverage of 80% reached) instead of the exact count, or drop the count assertion entirely since the acceptance check already confirms the proof artifact exists.

Correct fixes confirmed

None of the above are blocking. The path inconsistency is the most likely to cause a real downstream agent error.

@bmdhodl

bmdhodl commented Oct 5, 2026

Copy link
Copy Markdown
Owner Author

Replies to the Claude review above. No code change; reasons per point:

  1. "Codex" prose in AGENTS.md: kept on purpose. AGENTS.md is the instruction file Codex reads, so "Open Codex in this repo" and "Every Codex session" address its reader. The role prompts live in .claude/agents/; Codex reads them by path, and Claude Code loads sdk-dev and pm as subagents now that they have frontmatter. The old .Codex/agents/ path did not exist, which is what this PR fixed.
  2. ARCHITECTURE.md vs ops/02-ARCHITECTURE.md: both files exist and both get updates (root on 2026-09-27, ops/02 on 2026-09-22). Before this PR, AGENTS.md and CLAUDE.md checked different files; now both agent entry points check the root file. Moving ops-cadence.yml and the PR template is a CI and process change outside this PR's scope, so it stays in ops/FOLLOWUP.md for the owner to decide.
  3. Test count in verify.py: verify.py reads the saved file proof/prompt-audit-cleanup/09-test.txt, not a live pytest run. Adding or removing tests does not change that file, so the check stays stable on main. It fails only if someone edits the saved proof.

🤖 Addressed by Claude Code

@bmdhodl
bmdhodl enabled auto-merge (squash) October 5, 2026 14:48
@bmdhodl
bmdhodl merged commit 84b86d1 into main Oct 5, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants