Repository navigation
feat(p3): add the unprefixed-command-list check - #95
Merged
Merged
Conversation
brettdavies
force-pushed
the
feat/p3-unprefixed-command-list-check
branch
from
September 17, 2026 05:00
fd1d3f3 to
279e7df
Compare
brettdavies
added a commit
to brettdavies/agentnative
that referenced
this pull request
Sep 17, 2026
…name (#53) ## Summary Adds a SHOULD to P3 for the shape of the command list in `--help`. Hand-written help often lists commands as `tool status`, `tool server stop`, with the binary name at the head of every entry. A reader or agent scanning the block has to strip that prefix before it can read the command token, and an auditor that reads the block literally sees the binary name repeated instead of the commands. `p3-should-unprefixed-command-list` asks for the command token to sit in the left column on its own; the bare invocation line, which documents what the tool does with no arguments, stays the one legitimate carrier of the binary name alone. The requirement is a SHOULD rather than a MUST: the behavioral layer already grades this surface with warnings rather than failures, and a strict rule would penalize the hand-written help that much of the audited population ships. It applies only to CLIs that use subcommands, matching `p3-must-subcommand-examples`. The Evidence and Anti-Patterns sections gain one bullet each, and `last-revised` moves to 2026-09-17. ## Changelog ### Added - Add `p3-should-unprefixed-command-list` (SHOULD, applies when the CLI uses subcommands): command-list entries name the command directly rather than repeating the binary name as a prefix. ## Linked audit review brettdavies/agentnative-cli#95 (vendors this branch at `d593b3d` and adds the `p3-unprefixed-command-list` audit) ## Human reviewer **Reviewer:** @brettdavies ## AI disclosure The requirement's tier and firing condition were decided by the maintainer during planning; the frontmatter entry, prose bullets, and this PR body were drafted by Claude Code (Fable 5.1) from that plan and are pending the maintainer's review.
Merged
14 of 27 tasks
brettdavies
added this pull request to stack #96
September 17, 2026 15:52
Base automatically changed from
fix/hand-written-help-subcommand-parsing
to
dev
September 17, 2026 15:54
brettdavies
added a commit
that referenced
this pull request
Sep 17, 2026
…their real subcommands (#94) ## Summary A CLI whose `--help` lists its commands under a hand-written heading such as `Common commands:`, with every entry led by the binary name, parsed as having no subcommands. Fifteen behavioral audits read that empty list: ten degraded to skip and left the score denominator, and five answered on no evidence, so anc reported "binary has no subcommands" about herdr, a tool with sixteen. This PR teaches the shared help parser the hand-written shape, collapses the JSON-output audit's private copy of the parser onto it, and guards a substring matcher the fix would otherwise turn into false MUST failures. The parser opens a block on any header whose last word is `commands:` or `subcommands:`, takes the tool's own name from the `Usage:` line and from the binary's file stem, and strips that token from a block's entries only when every entry leads with it. A line indented more deeply than the block's first entry is a continuation of the previous entry (a wrapped description or a nested subcommand), not an entry of its own, so it can neither break the prefix agreement nor become a name. A name is the first token of the invocation part of each entry, so `herdr server stop` contributes `server`, `herdr machine <subcommand>` contributes `machine`, and the bare `herdr` line contributes nothing. Duplicates collapse and `.subcommands()` keeps its type, so none of the fifteen consumers changes. A new `command_blocks()` accessor exposes each block's header, raw entries, and stripped prefix, and `missing_subcommands_reason()` lets an audit say whether it found no block or found one it could not read. Destructive-verb matching moves from substring to segment-prefix so `format`, `transform`, `perform`, `confirm`, and `firmware` stop classifying as destructive while `delete-all`, `dropdb`, `rmdir`, `cleanup`, `reset-keys`, and `force-push` still do. This lands ahead of the parser change because the parser is what makes the substring rule reachable on hand-written-help CLIs. Plan: `docs/plans/2026-09-16-1542-fix-hand-written-help-subcommand-parsing-plan.md` (units U3, U1, U2; U4 is #95). ## Changelog ### Changed - Change scores for CLIs with a hand-written command list: ten audits that skipped on "no subcommands" now evaluate real names, and five that answered on no evidence now answer on real ones, so scores can move in either direction. herdr moves from 77 to 72 and stays badge-eligible. ### Fixed - Fix the help parser so a hand-written `... commands:` block whose entries repeat the binary name yields the real top-level command names instead of no subcommands. - Fix `p5-force-yes` and `p5-read-write-distinction` classifying `format`, `transform`, `perform`, `confirm`, and `firmware` as destructive because they contain `rm`. - Fix `p3-subcommand-examples` and `p6-standard-names` evidence to say whether no command block was found or a block was found but could not be read, instead of asserting the tool has no subcommands. - Fix `p2-json-output` opting out without probing any subcommand on a CLI whose command list is hand-written. ## Type of Change - [x] `fix`: Bug fix (non-breaking change which adds functionality) - [x] `refactor`: Code refactoring (no functional changes) - [ ] `feat`: New feature (non-breaking change which adds functionality) - [ ] `perf`: Performance improvement - [ ] `docs`: Documentation update - [ ] `test`: Adding or updating tests - [ ] `chore`: Maintenance tasks (dependencies, config, etc.) - [ ] `ci`: CI/CD configuration changes - [ ] `style`: Code style/formatting changes - [ ] `build`: Build system changes - [ ] `BREAKING CHANGE`: Breaking API change (requires major version bump) ## Related Issues/Stories - Story: `docs/plans/2026-09-16-1542-fix-hand-written-help-subcommand-parsing-plan.md` - Issue: n/a - Architecture: `docs/solutions/workflow-issues/anc-pager-substring-false-positive-2026-06-02.md` (the recorded sibling defect: a heuristic tuned to one vendor's output shape, applied to third-party text, failing silently) - Related PRs: follows #92 and #93, both merged ahead of this one; #95 (U4) follows, with its requirement from `brettdavies/agentnative` PR brettdavies/agentnative#53 (merged) ## Testing - [x] Unit tests added/updated - [x] Integration tests added/updated - [x] Manual testing completed - [x] All tests passing **Test Summary:** - Unit tests: 763 passing (1 ignored, pre-existing) - Integration tests: 123 passing across the seven integration binaries, including the new `test_handwritten_help_fixture_reports_real_subcommands` - Coverage: n/a (not measured in this repo) Each new test was observed failing against the unfixed code before the fix landed. The destructive-verb guard (U3), against the substring matcher: ```text ---- audits::behavioral::destructive_ops::tests::verb_inside_a_word_is_not_destructive stdout ---- thread '...' panicked at src/audits/behavioral/destructive_ops.rs:105:13: format should not be destructive ``` The hand-written block parse (U1), against the three-header parser: ```text ---- runner::help_probe::tests::hand_written_block_yields_top_level_names_without_the_tool_prefix stdout ---- assertion `left == right` failed left: [] right: ["status", "update", "completion", "server", "channel", "config", "machine", "api"] ``` The JSON-output probe on a hand-written block (U2), against the private parser: ```text ---- audits::behavioral::json_output::tests::json_output_probes_hand_written_command_block stdout ---- assertion `left == right` failed: got OptOut("no --output/--format flag detected — tool does not ship structured output. ...") left: OptOut("...") right: Pass ``` The four parser tests added from code review, against the parser as first written: ```text single_space_bare_entry_in_a_prefixed_block_is_description_only left: ["Launch", "status"] right: ["status"] deeper_indented_nested_lines_are_continuations_of_their_parent_entry left: ["server", "start", "stop"] right: ["server"] wrapped_description_continuation_keeps_the_prefixed_block_intact left: ["herdr", "server"] right: ["status", "update", "completion", "server", "channel", "config", "machine", "api"] probe_strips_the_binary_stem_when_the_usage_line_leads_with_a_launcher left: ["test"] right: ["status", "server"] ``` Manual verification: `anc audit .` on this repo changes no scorecard row against the base branch. `anc audit $(command -v herdr)` moves four rows from skip to real verdicts (`p6-standard-names` warn, `p6-consistent-naming` warn, `p3-subcommand-examples` fail, `p5-read-write-distinction` warn) and every other row keeps its status. `anc audit tests/fixtures/handwritten-help/tally` reports the fixture's seven real commands, passes `p2-json-output` through the `count` subcommand probe, and skips `p5-force-yes` because `format` is no longer destructive. ## Files Modified **Modified:** - `src/runner/help_probe/mod.rs`: command-block parser, `CommandBlock`, `command_blocks()`, `missing_subcommands_reason()`, usage-line tool name, stale allows and doc claims removed - `src/runner/mod.rs`: `BinaryRunner::binary_stem()` for the probe's tool-name fallback - `src/audits/behavioral/destructive_ops.rs`: segment-prefix destructive-verb matching - `src/audits/behavioral/json_output.rs`: shared parser via the project's cached help probe, built-ins filtered at the call site - `src/audits/behavioral/subcommand_help.rs`: `should_skip` shared with the JSON-output audit - `src/audits/behavioral/subcommand_examples.rs`: evidence wording on an empty parse - `src/audits/behavioral/standard_names.rs`: evidence wording on an empty parse - `tests/integration.rs`: end-to-end audit of the hand-written fixture **Created:** - `tests/fixtures/handwritten-help/tally`: shell fixture in the herdr shape (usage line, two prefixed `... commands:` blocks, per-subcommand help, `--output json` on one subcommand) **Renamed:** - None. **Deleted:** - None. ## Key Features - One parser serves every consumer of subcommand names, and it now reads clap, cobra (`Available Commands:`), and hand-written (`Common commands:`) blocks alike. - Evidence tells the truth about what the parser saw: "no command block found" or "a block was found but no names could be parsed", never "the tool has no subcommands". ## Benefits - Tools with hand-written help are graded on what they do. Ten audits re-enter the denominator and five stop answering on no evidence. - No CLI gains a MUST failure that is not a true statement about it: the destructive-verb guard lands in the same PR as the parser that would have exposed it. ## Breaking Changes - [x] No breaking changes - [ ] Breaking changes described below: ## Deployment Notes - [x] No special deployment steps required - [ ] Deployment steps documented below: ## Checklist - [x] Code follows project conventions and style guidelines - [x] Commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) - [x] Self-review of code completed - [x] Tests added/updated and passing - [x] No new warnings or errors introduced - [x] Changes are backward compatible (or breaking changes documented) ## Additional Context ### Post-Deploy Monitoring & Validation No production runtime is deployed by this repo; the artifact is the `anc` binary. After the next release, re-audit every tool on the site registry and record each badge-eligibility crossing: a score that falls because a hand-written-help tool is now graded on real names is the intended change, not a regression. A regression signal would be a clap-shaped tool whose scorecard rows change between the previous release and this one; `anc audit .` on this repo is the committed guard for that and changes no row. ### Review Code review ran through the compound-engineering review skill (run `20260916-220231-1a8ade92`, status complete, verdict "Ready with fixes") with correctness, project-standards, testing, maintainability, learnings, and adversarial lenses. The cross-model adversarial peer did not run: the installed `grok` CLI rejects the runner's `--json-schema` argument at launch, so no content left the machine and the in-process adversarial reviewer covered that lens. Every actionable finding was confirmed by an independent validator and applied in this PR: continuation lines in a prefixed block, the binary stem as a second prefix candidate, the single-space bare-invocation entry, the delimiter-branch destructive-verb test, and the module doc's view list. ### Unapplied review findings Both were demoted to residual risks by the review and are deferred as follow-ups rather than applied here. - [ ] P3, `src/audits/behavioral/destructive_ops.rs:39`: Segment-prefix matching no longer catches fused verbs such as `autoremove` and `autoclean`, which substring matching did. The plan's R9 defines the matcher as start-of-name or start-of-segment; adding the fused forms to the verb list with a failing-first test is the fix if recall matters. - [ ] P3, `src/runner/help_probe/mod.rs:401`: Cobra group headers with a parenthetical before the colon (`Basic Commands (Beginner):`, as kubectl prints) are not recognized, so such tools parse a partial name set. Not a regression, and outside R1 as written; stripping a trailing parenthetical before testing the last word would cover it. Residual risks the review recorded and this PR does not change: every parsed name is spawned as `<bin> <name> --help`, and a hand-written subcommand that ignores `--help` runs until the runner's timeout; the pager audit's substring matcher is the still-open sibling of the destructive-verb defect this PR fixes.
…their real subcommands (#94) ## Summary A CLI whose `--help` lists its commands under a hand-written heading such as `Common commands:`, with every entry led by the binary name, parsed as having no subcommands. Fifteen behavioral audits read that empty list: ten degraded to skip and left the score denominator, and five answered on no evidence, so anc reported "binary has no subcommands" about herdr, a tool with sixteen. This PR teaches the shared help parser the hand-written shape, collapses the JSON-output audit's private copy of the parser onto it, and guards a substring matcher the fix would otherwise turn into false MUST failures. The parser opens a block on any header whose last word is `commands:` or `subcommands:`, takes the tool's own name from the `Usage:` line and from the binary's file stem, and strips that token from a block's entries only when every entry leads with it. A line indented more deeply than the block's first entry is a continuation of the previous entry (a wrapped description or a nested subcommand), not an entry of its own, so it can neither break the prefix agreement nor become a name. A name is the first token of the invocation part of each entry, so `herdr server stop` contributes `server`, `herdr machine <subcommand>` contributes `machine`, and the bare `herdr` line contributes nothing. Duplicates collapse and `.subcommands()` keeps its type, so none of the fifteen consumers changes. A new `command_blocks()` accessor exposes each block's header, raw entries, and stripped prefix, and `missing_subcommands_reason()` lets an audit say whether it found no block or found one it could not read. Destructive-verb matching moves from substring to segment-prefix so `format`, `transform`, `perform`, `confirm`, and `firmware` stop classifying as destructive while `delete-all`, `dropdb`, `rmdir`, `cleanup`, `reset-keys`, and `force-push` still do. This lands ahead of the parser change because the parser is what makes the substring rule reachable on hand-written-help CLIs. Plan: `docs/plans/2026-09-16-1542-fix-hand-written-help-subcommand-parsing-plan.md` (units U3, U1, U2; U4 is #95). ## Changelog ### Changed - Change scores for CLIs with a hand-written command list: ten audits that skipped on "no subcommands" now evaluate real names, and five that answered on no evidence now answer on real ones, so scores can move in either direction. herdr moves from 77 to 72 and stays badge-eligible. ### Fixed - Fix the help parser so a hand-written `... commands:` block whose entries repeat the binary name yields the real top-level command names instead of no subcommands. - Fix `p5-force-yes` and `p5-read-write-distinction` classifying `format`, `transform`, `perform`, `confirm`, and `firmware` as destructive because they contain `rm`. - Fix `p3-subcommand-examples` and `p6-standard-names` evidence to say whether no command block was found or a block was found but could not be read, instead of asserting the tool has no subcommands. - Fix `p2-json-output` opting out without probing any subcommand on a CLI whose command list is hand-written. ## Type of Change - [x] `fix`: Bug fix (non-breaking change which adds functionality) - [x] `refactor`: Code refactoring (no functional changes) - [ ] `feat`: New feature (non-breaking change which adds functionality) - [ ] `perf`: Performance improvement - [ ] `docs`: Documentation update - [ ] `test`: Adding or updating tests - [ ] `chore`: Maintenance tasks (dependencies, config, etc.) - [ ] `ci`: CI/CD configuration changes - [ ] `style`: Code style/formatting changes - [ ] `build`: Build system changes - [ ] `BREAKING CHANGE`: Breaking API change (requires major version bump) ## Related Issues/Stories - Story: `docs/plans/2026-09-16-1542-fix-hand-written-help-subcommand-parsing-plan.md` - Issue: n/a - Architecture: `docs/solutions/workflow-issues/anc-pager-substring-false-positive-2026-06-02.md` (the recorded sibling defect: a heuristic tuned to one vendor's output shape, applied to third-party text, failing silently) - Related PRs: follows #92 and #93, both merged ahead of this one; #95 (U4) follows, with its requirement from `brettdavies/agentnative` PR brettdavies/agentnative#53 (merged) ## Testing - [x] Unit tests added/updated - [x] Integration tests added/updated - [x] Manual testing completed - [x] All tests passing **Test Summary:** - Unit tests: 763 passing (1 ignored, pre-existing) - Integration tests: 123 passing across the seven integration binaries, including the new `test_handwritten_help_fixture_reports_real_subcommands` - Coverage: n/a (not measured in this repo) Each new test was observed failing against the unfixed code before the fix landed. The destructive-verb guard (U3), against the substring matcher: ```text ---- audits::behavioral::destructive_ops::tests::verb_inside_a_word_is_not_destructive stdout ---- thread '...' panicked at src/audits/behavioral/destructive_ops.rs:105:13: format should not be destructive ``` The hand-written block parse (U1), against the three-header parser: ```text ---- runner::help_probe::tests::hand_written_block_yields_top_level_names_without_the_tool_prefix stdout ---- assertion `left == right` failed left: [] right: ["status", "update", "completion", "server", "channel", "config", "machine", "api"] ``` The JSON-output probe on a hand-written block (U2), against the private parser: ```text ---- audits::behavioral::json_output::tests::json_output_probes_hand_written_command_block stdout ---- assertion `left == right` failed: got OptOut("no --output/--format flag detected — tool does not ship structured output. ...") left: OptOut("...") right: Pass ``` The four parser tests added from code review, against the parser as first written: ```text single_space_bare_entry_in_a_prefixed_block_is_description_only left: ["Launch", "status"] right: ["status"] deeper_indented_nested_lines_are_continuations_of_their_parent_entry left: ["server", "start", "stop"] right: ["server"] wrapped_description_continuation_keeps_the_prefixed_block_intact left: ["herdr", "server"] right: ["status", "update", "completion", "server", "channel", "config", "machine", "api"] probe_strips_the_binary_stem_when_the_usage_line_leads_with_a_launcher left: ["test"] right: ["status", "server"] ``` Manual verification: `anc audit .` on this repo changes no scorecard row against the base branch. `anc audit $(command -v herdr)` moves four rows from skip to real verdicts (`p6-standard-names` warn, `p6-consistent-naming` warn, `p3-subcommand-examples` fail, `p5-read-write-distinction` warn) and every other row keeps its status. `anc audit tests/fixtures/handwritten-help/tally` reports the fixture's seven real commands, passes `p2-json-output` through the `count` subcommand probe, and skips `p5-force-yes` because `format` is no longer destructive. ## Files Modified **Modified:** - `src/runner/help_probe/mod.rs`: command-block parser, `CommandBlock`, `command_blocks()`, `missing_subcommands_reason()`, usage-line tool name, stale allows and doc claims removed - `src/runner/mod.rs`: `BinaryRunner::binary_stem()` for the probe's tool-name fallback - `src/audits/behavioral/destructive_ops.rs`: segment-prefix destructive-verb matching - `src/audits/behavioral/json_output.rs`: shared parser via the project's cached help probe, built-ins filtered at the call site - `src/audits/behavioral/subcommand_help.rs`: `should_skip` shared with the JSON-output audit - `src/audits/behavioral/subcommand_examples.rs`: evidence wording on an empty parse - `src/audits/behavioral/standard_names.rs`: evidence wording on an empty parse - `tests/integration.rs`: end-to-end audit of the hand-written fixture **Created:** - `tests/fixtures/handwritten-help/tally`: shell fixture in the herdr shape (usage line, two prefixed `... commands:` blocks, per-subcommand help, `--output json` on one subcommand) **Renamed:** - None. **Deleted:** - None. ## Key Features - One parser serves every consumer of subcommand names, and it now reads clap, cobra (`Available Commands:`), and hand-written (`Common commands:`) blocks alike. - Evidence tells the truth about what the parser saw: "no command block found" or "a block was found but no names could be parsed", never "the tool has no subcommands". ## Benefits - Tools with hand-written help are graded on what they do. Ten audits re-enter the denominator and five stop answering on no evidence. - No CLI gains a MUST failure that is not a true statement about it: the destructive-verb guard lands in the same PR as the parser that would have exposed it. ## Breaking Changes - [x] No breaking changes - [ ] Breaking changes described below: ## Deployment Notes - [x] No special deployment steps required - [ ] Deployment steps documented below: ## Checklist - [x] Code follows project conventions and style guidelines - [x] Commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) - [x] Self-review of code completed - [x] Tests added/updated and passing - [x] No new warnings or errors introduced - [x] Changes are backward compatible (or breaking changes documented) ## Additional Context ### Post-Deploy Monitoring & Validation No production runtime is deployed by this repo; the artifact is the `anc` binary. After the next release, re-audit every tool on the site registry and record each badge-eligibility crossing: a score that falls because a hand-written-help tool is now graded on real names is the intended change, not a regression. A regression signal would be a clap-shaped tool whose scorecard rows change between the previous release and this one; `anc audit .` on this repo is the committed guard for that and changes no row. ### Review Code review ran through the compound-engineering review skill (run `20260916-220231-1a8ade92`, status complete, verdict "Ready with fixes") with correctness, project-standards, testing, maintainability, learnings, and adversarial lenses. The cross-model adversarial peer did not run: the installed `grok` CLI rejects the runner's `--json-schema` argument at launch, so no content left the machine and the in-process adversarial reviewer covered that lens. Every actionable finding was confirmed by an independent validator and applied in this PR: continuation lines in a prefixed block, the binary stem as a second prefix candidate, the single-space bare-invocation entry, the delimiter-branch destructive-verb test, and the module doc's view list. ### Unapplied review findings Both were demoted to residual risks by the review and are deferred as follow-ups rather than applied here. - [ ] P3, `src/audits/behavioral/destructive_ops.rs:39`: Segment-prefix matching no longer catches fused verbs such as `autoremove` and `autoclean`, which substring matching did. The plan's R9 defines the matcher as start-of-name or start-of-segment; adding the fused forms to the verb list with a failing-first test is the fix if recall matters. - [ ] P3, `src/runner/help_probe/mod.rs:401`: Cobra group headers with a parenthetical before the colon (`Basic Commands (Beginner):`, as kubectl prints) are not recognized, so such tools parse a partial name set. Not a regression, and outside R1 as written; stripping a trailing parenthetical before testing the last word would cover it. Residual risks the review recorded and this PR does not change: every parsed name is spawned as `<bin> <name> --help`, and a hand-written subcommand that ignores `--help` runs until the runner's timeout; the pager audit's substring matcher is the still-open sibling of the destructive-verb defect this PR fixes.
`p3-should-unprefixed-command-list` is a new SHOULD in the vendored spec (agentnative d593b3d): command-list entries name the command directly rather than repeating the binary name as a prefix. The audit `p3-unprefixed-command-list` reads the command blocks the help probe already parses. A block whose entries all led with the tool name is prefixed, and every entry that names a command beyond that bare prefix is an offender; the bare-invocation entry, which documents what the tool does with no arguments, is not one, and a block whose header names examples is not graded, since examples are where the spec expects full invocations. It warns and names the offending invocations (the first five, then a count), passes when every graded block is unprefixed, and resolves not-applicable when no subcommand names were parsed from the help, the same gate its P3 and P6 siblings use, with the shared reason saying whether no block was found or a block was found but could not be read. It never fails. The registry grows from 59 to 60 requirements (22 SHOULDs) and the coverage matrix is regenerated to match. Auditing this repository adds one passing row; auditing herdr adds one warning and moves its score from 72 to 71, still badge-eligible.
brettdavies
force-pushed
the
feat/p3-unprefixed-command-list-check
branch
from
September 17, 2026 15:54
279e7df to
49484a8
Compare
brettdavies
added a commit
that referenced
this pull request
Sep 17, 2026
…hipped The UTF-8 evidence-preview plan and the hand-written-help subcommand-parsing plan both carry `status: completed`, and their bodies state what landed: the UTF-8 fix as PR #92 with its guard as PR #93, and the parser work as PR #94 (U3, U1, U2) with the new SHOULD check as PR #95, which vendors spec PR brettdavies/agentnative#53. The parsing plan's requirements and decisions now match the shipped parser: a block strips only when every entry leads with the tool's own name from the Usage line or the binary stem, `subcommands:` headers count, deeper-indented lines continue the previous entry, and a single-space bare entry is description only. Unit file lists name the files each PR touched, the audit is `p3-unprefixed-command-list`, the vendored VERSION stays at 0.5.0, and the follow-up list gains the fused verbs, kubectl's parenthetical headers, the release-time re-audit, and the spec release the scorecard's spec_version pin needs.
brettdavies
added a commit
to brettdavies/agentnative
that referenced
this pull request
Oct 5, 2026
…he release tooling (#56) ## Summary Publishes `p3-should-unprefixed-command-list`, the SHOULD that command-list entries name the command directly (`status`, `server stop`) rather than repeating the binary name on every line (`tool status`, `tool server stop`). The bare invocation line stays the one entry that legitimately carries the binary name alone. `principles/p3-progressive-help-discovery.md` gains the requirement block and the normative text. This closes a lockstep gap rather than adding something new. The published `anc` v0.6.0 already audits this requirement and emits a `p3-should-unprefixed-command-list` row in every scorecard, while declaring `spec_version: 0.5.0` against a published spec that does not define it. Until this merges, the auditor grades tools on a requirement a reader cannot look up. `VERSION` moves to 0.6.0. A new normative SHOULD is an additive spec change, and `publish.yml` resolves its CHANGELOG section from `VERSION`, so leaving it at 0.5.0 would point the release cut at the section v0.5.0 already published. The release also carries the governance and release tooling that accumulated on `dev` since v0.5.0: CODEOWNERS and a Dependabot config (both inert until they reach the default branch), the re-vendored `scripts/release/` helpers, `scripts/sync-dev-after-release.sh`, the consolidated `scripts/generate-changelog.py` replacing the shell version, refreshed guard workflows and the `protect-main` ruleset, and the consumer-facing docs describing them. The prose-check stack stays `dev`-only, as release #46 established. ## Changelog ### Added - Add `p3-should-unprefixed-command-list` (SHOULD, applies when the CLI uses subcommands): command-list entries name the command directly rather than repeating the binary name as a prefix, so a reader or agent takes the command token from the left column instead of reconstructing it from a known prefix. ### Changed - Require code-owner review on `main` and enable Dependabot, so contributor PRs now need an owner approval before merge. - Replace `scripts/generate-changelog.sh` with `scripts/generate-changelog.py`, which keeps the same flags and git-cliff plus PR-body pipeline and adds a duplicate-section guard. ## Linked audit review brettdavies/agentnative-cli#95 adds the `p3-unprefixed-command-list` audit that verifies this requirement. It shipped in `anc` v0.6.0, which is published. ## Human reviewer **Reviewer:** @brettdavies (review pending; this PR is open for it, not approved) ## AI disclosure Claude assembled the release branch, bumped `VERSION`, generated `CHANGELOG.md` from the constituent `dev` PR bodies, and wrote this body; every shipped change to the principle text, workflows, scripts, and docs is human-authored work already reviewed and merged to `dev` through its own PR.
11 of 13 tasks
brettdavies
added a commit
that referenced
this pull request
Oct 5, 2026
## Summary Vendors `agentnative-spec` v0.6.0, so `spec_version` in every scorecard reports 0.6.0 instead of 0.5.0. This closes a coherence gap rather than adopting new requirements. The registry already carried `p3-should-unprefixed-command-list` and `anc` already graded it, because the principle text was vendored from the spec's `dev` branch while the requirement was in flight. The vendored `VERSION` stayed at 0.5.0, so published scorecards declared conformance to a standard that did not define a requirement they carried a row for. Spec v0.6.0 publishes that SHOULD. Only `VERSION` and the vendored `CHANGELOG.md` change. The principle text is byte-identical to what was already vendored, so the generated requirement set, the registry counts, and `docs/coverage-matrix.md` are all untouched. `src/principles/spec/CHANGELOG.md` also joins the markdownlint ignore list, next to the root changelog that is already there for the same reason. The vendored file carries upstream's one-logical-line-per-bullet shape, which trips MD013 at 189 to 379 characters on the new 0.6.0 bullets. Wrapping it would make the vendored copy differ from upstream, which is the one property `scripts/sync-spec.sh` exists to preserve. ## Changelog ### Changed - Report `spec_version: "0.6.0"` in scorecards, matching the published spec that now defines `p3-should-unprefixed-command-list`. ## Type of Change - [x] `fix`: Bug fix (non-breaking change which fixes an issue) ## Related Issues/Stories - Story: n/a - Issue: n/a - Architecture: `scripts/SYNCS.md` (`sync-spec.sh` row) - Related PRs: brettdavies/agentnative#56 (the spec v0.6.0 release), #95 (the audit that verifies the requirement) ## Files Modified **Modified:** - `src/principles/spec/VERSION`: 0.5.0 to 0.6.0. - `src/principles/spec/CHANGELOG.md`: vendored, gains the upstream 0.6.0 section. - `.markdownlint-cli2.yaml`: ignore the vendored spec changelog. **Created:** - None. **Renamed:** - None. **Deleted:** - None. ## Testing - [ ] Unit tests added/updated - [ ] Integration tests added/updated - [x] Manual testing completed - [x] All tests passing **Test Summary:** - Unit tests: 1106 passing, 3 ignored - Integration tests: covered by the above - Coverage: not measured in this repo No new tests: the change is a vendored version string, and nothing in `tests/` asserts a `spec_version` literal, which is what lets the bump pass without edits. Verified in emitted output rather than inferred from the file: `anc audit --command jq --audit-profile posix-utility --output json` reports `spec_version=0.6.0`, and the `p3-should-unprefixed-command-list` row is still present and resolves to `n_a` for `jq`, which is correct for a requirement gated on the CLI having subcommands. Confirmed the new markdownlint ignore takes effect (`Linting: 0 files` for the vendored changelog) while other markdown still lints. ## Breaking Changes - [x] No breaking changes Consumers that pin on `spec_version` see it move from `0.5.0` to `0.6.0`. The site accepts the scorecard on `schema_version`, not `spec_version`, so nothing downstream gates on this. ## Deployment Notes - [x] No special deployment steps required Published `anc` v0.6.0 keeps reporting `spec_version: "0.5.0"`; this reaches users with the next `anc` release. ## Checklist - [x] Code follows project conventions and style guidelines - [x] Commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) - [x] Self-review of code completed - [x] Tests added/updated and passing - [x] No new warnings or errors introduced - [x] Changes are backward compatible (or breaking changes documented) ## Additional Context The spec's `publish.yml` dispatches a `spec-release` event to this repo and to `agentnative-site` on every tagged release, with its loop commented "Non-fatal: downstream consumers opt in when they wire handlers." Neither repo has a handler, so no workflow fires and the dispatch is silently a no-op. That is why this sync is a hand-run script and a PR rather than something that happened on its own, and wiring a handler here would be a reasonable follow-up.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds the P3 SHOULD check for command lists that repeat the binary name. A hand-written
Common commands:block whose entries readtool status,tool server stopmakes a reader or agent strip the prefix before the command token is visible. The newp3-unprefixed-command-listaudit reads the command blocks the help probe already parses (from the parser change in #94): a block whose entries all led with the tool name is prefixed, and every entry that names a command beyond that bare prefix is an offender. The bare-invocation entry, which documents what the tool does with no arguments, is not one, and a block whose header names examples (Example commands:) is not graded, since examples are where the spec expects full invocations. The audit warns and names the offending invocations (the first five, then a count), passes when every graded block is unprefixed, and resolves not-applicable when no subcommand names were parsed, on the same condition its P3 and P6 siblings use, with the shared reason saying whether no block was found or a block was found but could not be read.The requirement
p3-should-unprefixed-command-listis authored upstream in the spec and vendored here from8b84077, the squash-merge of spec PR #53 on the spec'sdevbranch (see Deployment Notes). The registry grows from 59 to 60 requirements (22 SHOULDs) and the coverage matrix is regenerated to match.Plan:
docs/plans/2026-09-16-1542-fix-hand-written-help-subcommand-parsing-plan.md, unit U4.Changelog
Added
p3-unprefixed-command-listaudit for the new SHOULDp3-should-unprefixed-command-list: a--helpcommand list whose entries repeat the binary name warns with the offending entries named; a clap-shaped list passes; a tool whose help yields no subcommand names is not applicable; an examples section is never graded.Changed
--help: the new SHOULD row enters the denominator, as a pass for clap-shaped lists and a warn for prefixed ones. herdr moves from 72 to 71 and stays badge-eligible; this repository gains a passing row.Type of Change
feat: New feature (non-breaking change which adds functionality)fix: Bug fix (non-breaking change which adds functionality)refactor: Code refactoring (no functional changes)perf: Performance improvementdocs: Documentation updatetest: Adding or updating testschore: Maintenance tasks (dependencies, config, etc.)ci: CI/CD configuration changesstyle: Code style/formatting changesbuild: Build system changesBREAKING CHANGE: Breaking API change (requires major version bump)Related Issues/Stories
docs/plans/2026-09-16-1542-fix-hand-written-help-subcommand-parsing-plan.md(U4)CLAUDE.md§ Principle Registry (requirements are generated from the vendored spec frontmatter; the counter tests are bumped deliberately)brettdavies/agentnativePR feat(p3): add SHOULD for command-list entries that repeat the binary name agentnative#53 (merged)Testing
Test Summary:
Each guard was observed failing before the audit and counters landed. With the vendored requirement in place and nothing else changed:
The fixture assertion for the new row, before the audit was registered:
The three tests added from code review, against the audit as first written:
The bare-invocation guards in the unit test and the fixture test were proven live by removing the empty-command filter, observing both fail, and restoring it:
Manual verification:
anc emit coverage-matrix --checkexits 0 after regeneration.anc audit .on this repository adds exactly one row against #94,p3-unprefixed-command-listas pass, and changes no other row.anc audit tests/fixtures/handwritten-help/tallywarns once on the new row and namestally count <path>while leaving the baretallyentry out of the evidence.anc audit $(command -v herdr)warns once, naming its prefixed entries and counting the rest.Files Modified
Modified:
src/principles/spec/principles/p3-progressive-help-discovery.md: vendored spec at8b84077with the new SHOULDsrc/audits/behavioral/mod.rs: audit registeredsrc/runner/mod.rs:CommandBlockre-exported alongsideHelpOutputsrc/principles/registry.rs: counter tests bumped to 60 requirements and 22 SHOULDstests/build_parser.rs: vendored-spec count bumped to 60tests/integration.rs: the hand-written fixture asserts the new warn row and its evidencedocs/coverage-matrix.md,coverage/matrix.json: regeneratedCreated:
src/audits/behavioral/unprefixed_command_list.rs: the audit and its unit testsRenamed:
Deleted:
Key Features
herdr status [server|client]), so the remediation is legible without reopening the help text.Benefits
Breaking Changes
Deployment Notes
The requirement is vendored from the spec's
devbranch at8b84077(the squash-merge of spec PR #53), not from a tagged spec release, because the audit cannot compile until the requirement id resolves. The commit on this branch was vendored from the spec branch headd593b3dbefore that merge;scripts/sync-spec.sh --ref 8b84077run after the merge produces a byte-identical tree, so there is nothing further to commit here. The vendoredVERSIONstays0.5.0, which the scorecard reports asspec_version, while the published v0.5.0 spec has 59 requirements; cut a spec release that includes this SHOULD before the next CLI release, so the pin names a published spec that contains it. The normal tag sync picks it up at that point.Screenshots/Recordings
n/a
Checklist
Additional Context
Post-Deploy Monitoring & Validation
No production runtime is deployed by this repo; the artifact is the
ancbinary. After the next release, every tool on the site registry gains one new SHOULD row. A prefixed hand-written list warning is the intended change. A regression signal would be a clap-shaped tool whose new row is anything but pass, or a tool with no command list whose row is anything but not-applicable;anc audit .on this repository is the committed guard for the pass case.Review
Code review ran through the compound-engineering review skill (run
20260916-231619-bf850931, status complete, verdict "Ready with fixes") with correctness, project-standards, testing, maintainability, learnings, and adversarial lenses; the cross-model adversarial peer was skipped on the route failure observed earlier in the session (the installedgrokCLI rejects the runner's arguments at launch, so nothing left the machine) and the in-process adversarial reviewer covered that lens. All three confirmed findings were applied in this PR: the applicability gate now matches the sibling audits (not applicable when no subcommand names were parsed), example-style blocks are excluded from grading, and the bare-invocation guards in both tests now assert on a rendering the formatter can produce.Unapplied review findings
None of the review's actionable findings were deferred. Residual risks the review recorded and this PR does not change:
devbranch ahead of a spec release; see Deployment Notes for the release ordering, and the site's scorecard regeneration is where badge eligibility shifts for tools sitting at the 70% floor.--audit-profilecategory suppresses the new SHOULD; whether one should is a judgment call the drift tests do not make.