Skip to content

fix(audit): pass p7-quiet only when the help declares --quiet or -q - #171

Open
brettdavies wants to merge 1 commit into
fix/help-flags-u9-declared-flagsfrom
fix/help-flags-u10-quiet
Open

brettdavies wants to merge 1 commit into
fix/help-flags-u9-declared-flagsfrom
fix/help-flags-u10-quiet

Conversation

@brettdavies

@brettdavies brettdavies commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Summary

p7-quiet passed whenever the --help text contained the characters --quiet or -q anywhere. Four registry tools passed that way without a quiet flag, because -q sits inside another flag's name:

  • eza: --no-quotes
  • helm: --qps
  • scc: --file-list-queue-size
  • yq: --ini-preserve-quotes

A -q shown only in a usage line (Usage: quill [-q] [-o FILE] <input>) passed too.

The audit now asks the definition query whether the help declares --quiet or -q, and reads the help through the shared probe the other flag audits use. A warn says what was searched: no option definition in --help declares --quiet or -q; usage lines are not read. When the help declares -quiet beside double-dash names, the warn adds why that spelling does not count.

p7-quiet is one of the four audits behind the scorecard's audience, so a pass that becomes a warn is one more warn in that count.

Changelog

Fixed

  • Fix p7-must-quiet passing for tools with no quiet flag. The row passed when -q appeared anywhere in --help, including inside --no-quotes or --qps and in a usage line. It now passes only when an option definition declares --quiet or -q.

Changed

  • p7-must-quiet warn evidence now reads no option definition in --help declares --quiet or -q; usage lines are not read.

Type of Change

  • fix: Bug fix (non-breaking change which fixes an issue)

Related Issues/Stories

Testing

  • Unit tests added/updated
  • Integration tests added/updated
  • Manual testing completed
  • All tests passing

Test Summary:

  • Unit tests: five in quiet.rs. The four tools' real help lines do not pass. -q, --quiet, --quiet alone and -q alone pass. A -q shown only in a usage line warns with the search-scope evidence. -quiet beside double-dash names warns and says why, and -quiet in a help with no double-dash name passes.
  • Integration: tests/fixtures/handwritten-help/quill shows -q only in its usage line. anc audit <path> --output json gives a p7-quiet row that warns with the search-scope evidence, and the audience label equals the label for the number of warns among the four signal rows with that row counted.
  • 1302 passing across the suite. Self-audit: cargo test --test dogfood passes.
  • Snapshots: none change.

Negatives observed failing. The new tests, run against the base (#170) with the audit reading the shared probe and the substring match still in place:

---- audits::behavioral::quiet::tests::a_quiet_flag_shown_only_in_a_usage_line_warns_and_names_what_was_searched stdout ----
assertion `left == right` failed
  left: Pass
 right: Warn("no option definition in --help declares --quiet or -q; usage lines are not read.")
---- audits::behavioral::quiet::tests::a_single_dash_quiet_beside_double_dash_names_warns_and_says_why stdout ----
expected Warn, got Pass
---- audits::behavioral::quiet::tests::a_flag_whose_name_contains_q_is_not_a_quiet_flag stdout ----
["eza 0.23.5", "helm 4.3.0", "scc 4.1.0", "yq 4.54.1"]
test result: FAILED. 5 passed; 3 failed; 0 ignored; 0 measured; 1029 filtered out; finished in 0.01s

---- test_usage_line_flag_is_not_a_declared_quiet_flag stdout ----
assertion `left == right` failed: {"id":"p7-must-quiet","label":"Quiet mode available","group":"P7","layer":"behavioral","status":"pass","evidence":null,"confidence":"high","tier":"must","audit_id":"p7-quiet"}
  left: String("pass")
 right: "warn"
test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 71 filtered out; finished in 0.08s

Expected moves. Written before the corpus run, from the base and head builds replayed over the full registry capture set. 66 rows, all p7-must-quiet (p7-quiet):

tools before after why
eza, helm, scc, yq (4) pass warn -q appears only inside another flag's name (the lines in the summary); no definition declares --quiet or -q
actionlint, age, ast-grep, atuin, aws-cli, bat, biome, bottom, broot, bun, claude-code, cmake, codex, cosign, curl, dasel, datasette, delta, direnv, docker, dust, ffmpeg, files-to-prompt, flyctl, gemini-cli, gh, git, git-cliff, gitleaks, gitui, glow, goose, gum, hyperfine, jnv, jq, kubectl, lazygit, llm, lsd, miller, nushell, ollama, opencode, pastel, procs, qmd, rclone, sd, shell-gpt, shellcheck, sqlite-utils, starship, supabase, terraform, tmux, tokei, typst, wrangler, xsv, yazi, zoxide (62) warn warn, evidence only every existing warn is reworded from no --quiet/-q flag detected in --help output to the search-scope sentence

The other 29 scored tools pass before and after: each declares --quiet or -q on a definition line. No row in the corpus carries the dash-rule note.

Derived, computed from the base scorecards:

tool badge.score_pct band badge.eligible audience
eza 85 to 83 85-100 to 80-84 true, unchanged agent-optimized, unchanged
helm 67 to 66 50-69, unchanged false, unchanged agent-optimized, unchanged
scc 78 to 77 75-79, unchanged true, unchanged agent-optimized, unchanged
yq 77 to 76 75-79, unchanged true, unchanged agent-optimized, unchanged

Each of the four also moves summary.pass down one and summary.warn up one. No propagated row moves: p7-must-quiet is no row's antecedent.

The plan's seed for this PR expected helm, scc and yq to move from agent-optimized to mixed. The base scorecards say otherwise: for all three, p2-json-output is skip and p1-non-interactive and p6-no-color-behavioral pass, so the quiet warn is their only warn among the four signal rows, and one warn keeps the label at agent-optimized.

Corpus before/after. Run after the tables above were written. Base 175c634d343a, head 682c69ed3ac3. Image sha256:05db16cf226cd14a93a2682aea41b72d3e6b4d240cadba006f2881d3ffa996b3, network bridge. 98 tools selected, 95 scored under both builds, 67 rerun three times.

Moved rows (66), grouped by what changed

tools id audit_id status evidence
eza, helm, scc, yq (4) p7-must-quiet p7-quiet pass → warn base: null
head: no option definition in --help declares --quiet or -q; usage lines are not read.
actionlint, age, ast-grep, atuin, aws-cli, bat, biome, bottom, broot, bun, claude-code, cmake, codex, cosign, curl, dasel, datasette, delta, direnv, docker, dust, ffmpeg, files-to-prompt, flyctl, gemini-cli, gh, git, git-cliff, gitleaks, gitui, glow, goose, gum, hyperfine, jnv, jq, kubectl, lazygit, llm, lsd, miller, nushell, ollama, opencode, pastel, procs, qmd, rclone, sd, shell-gpt, shellcheck, sqlite-utils, starship, supabase, terraform, tmux, tokei, typst, wrangler, xsv, yazi, zoxide (62) p7-must-quiet p7-quiet warn base: no --quiet/-q flag detected in --help output
head: no option definition in --help declares --quiet or -q; usage lines are not read.

Derived fields moved (13)

tool field base head
eza badge.score_pct 85 83
eza band 85-100 80-84
eza summary.pass 19 18
eza summary.warn 6 7
helm badge.score_pct 67 66
helm summary.pass 16 15
helm summary.warn 15 16
scc badge.score_pct 78 77
scc summary.pass 18 17
scc summary.warn 11 12
yq badge.score_pct 77 76
yq summary.pass 19 18
yq summary.warn 13 14

The 66 moved rows are the 66 expected, row for row with the same before and after evidence, and the derived moves equal the predictions: no audience or badge.eligible value moved. Every moved row returned the same result in all four runs of each build. No row is unstable, no harness bug is flagged, and no tool failed to score under one build only. mods p3-should-about-long-about, the one row on the A/A noise list, did not move (base runs pass, pass, pass, pass; head runs warn, warn, pass, pass). cursor, nvidia-smi and xai-grok-build are absent from the image and ran under neither build.

Files Modified

Modified:

  • src/audits/behavioral/quiet.rs: audit_quiet asks the definition query; run() reads the shared help probe.
  • tests/integration.rs: the usage-line case.

Created:

  • tests/fixtures/handwritten-help/quill

Renamed:

  • None.

Deleted:

  • None.

Breaking Changes

  • No breaking changes

Four registry tools lose a p7-must-quiet pass they did not earn, and every p7-must-quiet warn carries new evidence text. The scorecard schema and JSON shape are unchanged.

Deployment Notes

  • No special deployment steps required

Checklist

  • Code follows project conventions and style guidelines
  • Commit messages follow Conventional Commits
  • Self-review of code completed
  • Tests added/updated and passing
  • No new warnings or errors introduced
  • Changes are backward compatible (or breaking changes documented)

`p7-quiet` passed whenever the `--help` text contained the characters `--quiet` or `-q` anywhere. eza's `--no-quotes`, helm's `--qps`, scc's `--file-list-queue-size` and yq's `--ini-preserve-quotes` each carry `-q` inside another flag's name, so all four passed without a quiet flag. A `-q` shown only in a usage line passed too.

The audit now asks the definition query for `--quiet` or `-q` and reads the help through the shared probe the other flag audits use. A warn says what was searched: `no option definition in --help declares --quiet or -q; usage lines are not read.` When the help declares `-quiet` beside double-dash names, the warn adds why that spelling does not count.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant