Repository navigation
fix(principles): accept documented equivalents for confirmation and schema, exempt argument-free subcommands from examples, and list real audit IDs - #58
Merged
Conversation
…ples MUST p3-must-subcommand-examples now binds only subcommands that take their own arguments or options. A subcommand whose `--help` shows no positional argument and no option beyond `-h`/`--help` and the inherited global flags is exempt, because its usage line (`tool version`, `tool logout`) is its whole call shape and an example would only repeat it. The requirement's applicability narrows from "CLI uses subcommands" to "CLI has subcommands that take their own arguments or options", the summary and the MUST bullet carry the exemption, and the two Evidence lines that named every subcommand now name the argument-taking ones. last-revised moves to 2026-10-05 for the frontmatter change.
p5-must-force-yes named only `--force` and `--yes`, so a destructive command that confirms through a differently named, documented flag failed the MUST even though an agent can pass it non-interactively. The requirement now asks for an explicit confirmation flag: `--yes` and `--force` stay the recommended names, and a flag under another name satisfies it when the command's `--help` documents it as the flag that confirms the operation without a prompt (`terraform destroy -auto-approve`). The summary, the MUST bullet, and the Evidence line change together. Level and applicability are unchanged; last-revised moves to 2026-10-05 for the summary edit.
p2-must-schema-print named only a `schema` subcommand or a `--schema` flag, so a CLI that exposes its output structure through a documented command under another name failed the MUST. The requirement keeps `schema` and `--schema` as the recommended surface and accepts an introspection command under another name when the tool's `--help` documents it as the command that describes the output's structure (`kubectl explain`, which prints the fields of the resources `kubectl get -o json` returns). The schema still MUST identify its format. The summary and the MUST bullet change together. Level and applicability are unchanged; last-revised moves to 2026-10-05 for the summary edit.
The closing "Measured by audit IDs" line of seven principles named IDs that anc 0.6.0 does not have, and every line listed only part of what `anc audit --principle <n>` reports. Each line now names the audits anc 0.6.0 reports for that principle, in requirement order, with the audits that score no requirement row last. P8 gains the line it lacked. IDs that do not exist in anc 0.6.0, and what replaces them: - `p2-output-json` and `p2-output-format`: `p2-json-output` (behavioral) and `p2-structured-output` (source), the verifiers of p2-must-output-flag. - `p2-stderr-diagnostics`: `p2-output-module` (source), the only verifier of p2-must-stdout-stderr-split. - `p3-after-help`: `p3-subcommand-examples` (per-subcommand examples) beside `p3-help` (top-level examples). - `p4-unwrap`: `code-unwrap`, a cross-cutting code-quality audit that every principle filter reports; its Python counterpart `code-bare-except` is listed with it. - `p5-destructive-guard`: `p5-force-yes`. - `p7-timeout`: `p7-timeout-behavioral`, the verifier of p7-should-timeout. P3's line also gains `p3-unprefixed-command-list`, `p3-subcommand-examples`, `p3-paired-examples`, `p3-about-long-about`, and `p3-examples-subcommand`. P1's line drops its layer labels, one of which named `p1-non-interactive-source` a source audit when anc runs it in the project layer. `p6-agents-md` is left off P6's line: anc groups it under P6, but it scores P8's p8-should-bundle-exists, and `anc audit --principle 8` does not report it, so it cannot sit on P8's line either. `p8-bundle-exists` measures that requirement under P8. Prose only below the frontmatter; no requirement changes.
…mber The four-artifacts list said the spec was "Currently v0.4.0" after v0.5.0 and v0.6.0 shipped. It now points at VERSION, the way the Versioning section already does, so the line holds across releases.
Vale's spelling rule rejects "MUSTs" as a word and blocks the pre-push prose check on any branch that touches README.md. "new or changed MUST requirements" says the same thing and passes.
A command group (`tool server <COMMAND>`) takes the nested command as its argument, so the zero-argument exemption from the examples MUST does not cover it. The text now says so, matching the audit.
13 of 25 tasks
check-last-revised requires a principle whose frontmatter changed to carry the date of the change under review; the three edited principles move from 2026-10-05 to 2026-10-06.
brettdavies
added a commit
to brettdavies/agentnative-cli
that referenced
this pull request
Oct 6, 2026
…ds, a JSON probe, and the schema command (#147) ## Summary `.anc.toml` gains four settings that let a tool's repository tell anc what its help cannot say by name alone, using the same mechanism as `[p6] domain_verbs`: files layer from the target's git root down plus `~/.anc.toml` (relocated by `AGENTNATIVE_HOME_CONFIG`), the nearer file is credited, a wrong type voids the whole chain, and the evidence names the setting and the file it came from. - **`[p5] confirm_flags`** (for example `["-auto-approve"]`): confirmation flags that count beside the built-ins for `p5-force-yes`. A declared flag counts only when the destructive subcommand's own `--help` lists it, so the setting adds a name anc accepts but cannot vouch for a flag the help never shows. Single-dash long flags are matched by their whole name. - **`[p5] not_destructive`** (for example `["clean"]`): subcommands the tool declares are not destructive drop out of `p5-force-yes`; when none remain the row reads `skip`, as for a tool with no destructive subcommands. `p5-read-write-distinction` still treats them as writes. - **`[p2] json_probe`** (for example `["version", "--client", "-o", "json"]`): a safe read-only invocation anc runs with exactly those arguments (no shell, the usual timeout and sandboxing). `p2-json-output` passes when it exits 0 with one JSON value on stdout and fails otherwise; it replaces only the case where anc cannot validate JSON on its own probes. - **`[p2] schema_command`** (for example `["explain"]`): a schema command under another name that counts for `p2-schema-print` when the help lists it. Two changes need no configuration: - **`p5-force-yes` accepts `--auto-approve`, `--assume-yes`, and `--confirm`** beside `--force`, `--yes`, `-y`, and `-f`. - **`p2-json-output` reads `skip`, not `warn`, when an output flag exists but anc's safe probes cannot validate JSON.** The warn recorded anc's inability to check, not the tool's behavior. Antecedent rules carry the skip into `p2-schema-print` and `p2-schema-file`. A passing row's `evidence` lists what the audit matched first, then the `.anc.toml` setting the pass relied on, joined by a semicolon. The scorecard schema stays at 0.9 with no new field. The shared subcommand `--help` probe is one helper (`probe_help`) used by every audit that probes subcommands. The spec's P5 and P2 wording that these settings implement is in brettdavies/agentnative (brettdavies/agentnative#58). ## Changelog ### Added - Add `[p5] confirm_flags` to `.anc.toml`: confirmation flags under other names (`-auto-approve`) count for `p5-force-yes` when the destructive subcommand's own `--help` lists them. - Add `[p5] not_destructive` to `.anc.toml`: subcommands a tool declares are not destructive (a cache `clean`) are left out of `p5-force-yes`. - Add `[p2] json_probe` to `.anc.toml`: a safe read-only invocation anc runs to verify JSON output when its own probes cannot. - Add `[p2] schema_command` to `.anc.toml`: a schema command under another name (`explain`) counts for `p2-schema-print`. ### Changed - Change `p5-force-yes` to accept `--auto-approve`, `--assume-yes`, and `--confirm` as confirmation flags. - Change `p2-json-output` to read `skip` instead of `warn` when an output flag exists but no safe probe can validate JSON. - Change a passing row's evidence to name the `.anc.toml` setting the pass relied on, after what the audit matched. ## Type of Change - [x] `feat`: New feature (non-breaking change which adds functionality) - [ ] `fix`: Bug fix (non-breaking change which fixes an issue) - [ ] `refactor`: Code refactoring (no functional changes) - [ ] `perf`: Performance improvement - [ ] `docs`: Documentation update - [ ] `test`: Adding or updating tests - [ ] `chore`: Maintenance tasks (dependencies, config, etc.) - [ ] `ci`: CI/CD configuration changes - [ ] `style`: Code style/formatting changes - [ ] `build`: Build system changes - [ ] `BREAKING CHANGE`: Breaking API change (requires major version bump) ## Related Issues/Stories - Story: n/a - Issue: n/a - Architecture: P5 `p5-must-force-yes`, P2 `p2-must-output-flag`, `p2-must-schema-print` - Related PRs: #143, #142 (merged; this branch builds on their subcommand probe and pass evidence), brettdavies/agentnative#58 (the spec companion change) ## Testing - [x] Unit tests added/updated - [x] Integration tests added/updated - [x] Manual testing completed - [x] All tests passing **Test Summary:** - Each commit's tests were shown failing with only its fix reverted, for example: built-in aliases `left: Fail("destructive subcommand(s) without --force or --yes: destroy…") right: Pass`; `confirm_flags` precedence `left: [] right: [("-auto-approve", ".anc.toml"), ("--nuke", "~/.anc.toml")]`; unverifiable JSON `expected Skip, got Warn("--output/--format flag detected but could not validate JSON…")`; `schema_command` `left: Fail("…no schema subcommand…") right: Pass`; pass-row evidence `left: Some("delete (--force).") right: Some("delete (--force); destroy accepts -auto-approve via .anc.toml [p5].confirm_flags")`. - `cargo fmt --check`, both clippy runs, `cargo test` (1191 passing), `cargo test -- --ignored`, and the pre-push battery pass; each commit builds and passes clippy and tests on its own. - Real tools in the site's scorer image (the versions the anc.dev corpus uses), dev against this branch: | Tool and setup | Score | Rows | |---|---|---| | terraform 1.16.4, no config or `confirm_flags = ["-auto-approve"]` | 66 to 66 | `p5-force-yes` still fails: `terraform destroy -help` does not list `-auto-approve` (only `apply -help` does); the evidence names the setting | | biome 2.5.15 / with `not_destructive = ["clean"]` | 70 to 71 / 73 | `p2-json-output` skip; `p5-force-yes` skip with the setting | | kubectl 1.37.1, no config | 68 to 68 | `p2-json-output` skip | | kubectl, `json_probe = ["version", "--client", "-o", "json"]` | 67 | `p2-json-output` pass; `p2-schema-print` fail (kubectl has no `schema` command) | | kubectl, that probe plus `schema_command = ["explain"]` | 70 | both pass | | helm 4.3.0, no config / `json_probe = ["repo", "list", "-o", "json"]` | 65 to 67 / 66 | `p2-json-output` skip / pass; `p2-schema-print` skip / fail | | rclone 1.75.1 | 73 to 76 | `p2-json-output` and `p2-schema-print` skip | | ollama, docker, opencode (no config) | +1, +1, +3 | `p2-json-output` skip | A probe that does not work fails the row: kubectl `version -o json` without `--client` exits 1 (it reaches for a cluster), helm `version -o json` exits 1 (unknown flag). ## Files Modified **Modified:** - `src/anc_toml/mod.rs`, `src/types.rs`, `src/scorecard/mod.rs`, `src/runner/help_probe/mod.rs`, `src/audits/behavioral/{force_yes,json_output,schema_print,standard_names,subcommand_help,read_write_distinction}.rs` - `README.md`, `CLAUDE.md`, `schema/scorecard.schema.json` (descriptions only) **Created:** - `src/anc_toml/settings.rs`, `tests/anc_toml_settings_integration.rs` **Renamed:** - None. **Deleted:** - None. ## Breaking Changes - [x] No breaking changes - [ ] Breaking changes described below: ## Deployment Notes - [x] No special deployment steps required - [ ] Deployment steps documented below: ## Checklist - [x] Code follows project conventions and style guidelines - [x] Commit messages follow [Conventional Commits](https://www.conventionalcommits.org/) - [x] Self-review of code completed - [x] Tests added/updated and passing - [x] No new warnings or errors introduced - [x] Changes are backward compatible (or breaking changes documented) ## Additional Context Not changed here: the shared flag parser reads a single-dash long flag such as `-force-copy` by its first letter, so the built-in `-f` can match it for `p5-force-yes`; `HelpOutput::advertises_flag`, added here for declared flags, is where that fix would go. kubectl's grouped command headings (`Basic Commands (Beginner):`) are not read as command lists, so kubectl's destructive and read/write rows skip.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Three MUST requirements read narrower than what they measure, and the principle files named audit IDs anc does not report. Measuring the curated CLI corpus with anc 0.6.0 surfaced each.
p3-must-subcommand-examplesapplies to subcommands that take their own arguments or options. A subcommand whose--helpshows no positional argument (a nested command list counts as one) and no option beyond-h/--helpand the global flags every subcommand inherits is exempt: its usage line is its whole call shape, so an example would repeat it. The frontmatter applicability, the MUST text, and two Evidence lines change to match.p5-must-force-yesaccepts a documented confirmation flag under another name.--yesand--forcestay the recommended names; a flag under another name satisfies the requirement when the command's--helpdocuments it as the flag that confirms without a prompt (terraform destroy -auto-approve).p2-must-schema-printaccepts a documented introspection command under another name (kubectl explain). Aschemasubcommand or--schemaflag stays the recommended surface.anc audit --principle Noutput. Seven listed IDs did not exist (p2-output-json,p2-output-format,p2-stderr-diagnostics,p3-after-help,p4-unwrap,p5-destructive-guard,p7-timeout), P3's line lackedp3-unprefixed-command-list, and P8 had no closing line.VERSIONinstead of a pinned version number, and spells out "MUST requirements" where the prose linter blocked "MUSTs".The three MUST edits change requirement text and applicability without changing any level; they bump
last-revisedon P2, P3, and P5.Changelog
Changed
p3-must-subcommand-examplesto apply to subcommands that take their own arguments or options; a subcommand that takes none (no positional argument, no nested commands, no option beyond help and the global flags) is exempt.p5-must-force-yesto accept a confirmation flag under another name when the command's--helpdocuments it as confirming without a prompt, with--yesand--forceas the recommended names.p2-must-schema-printto accept a documented introspection command under another name as the schema surface, withschema/--schemaas the recommended surface.Fixed
Linked audit review
p3-subcommand-examples, with the same rule..anc.toml[p5] confirm_flags(a declared flag counts only when the destructive subcommand's own--helplists it),[p2] schema_command, and the built-in--auto-approve,--assume-yes,--confirm.Human reviewer
Reviewer: @brettdavies (review pending before merge)
AI disclosure
An AI coding agent drafted every edit from the anc 0.6.0 corpus measurements and anc's source; the maintainer reviews before merge.