Skip to content

fix(audit): detect an output flag from definitions and probe it as spelled - #178

Open
brettdavies wants to merge 2 commits into
fix/help-flags-u10-non-interactivefrom
fix/help-flags-u10-json-output
Open

brettdavies wants to merge 2 commits into
fix/help-flags-u10-non-interactivefrom
fix/help-flags-u10-json-output

Conversation

@brettdavies

@brettdavies brettdavies commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Summary

p2-json-output decided whether a tool has an output flag by looking for the characters --output or --format anywhere in its lowercased --help, at the top level and then in each subcommand. Then it probed with --output json or --format json, whatever the help printed.

  • A longer flag counted: claude's --output-format, make's --output-sync, cosign's --output-file, codex's --output-schema, biome's --formatter. The probe then passed --output json to a tool that has no --output.
  • A usage line and a sentence counted.
  • A Go flag help that declares -format was not detected, and would have been probed with --format if it had been.

Detection now asks the definition query, per help, and each safe probe passes the flag as the help spells it: --help -format json for a help that declares -format and no double-dash name. A longer flag that starts with --output or --format is a different name and triggers no probe.

The evidence names what was found or searched:

  • skip: `--output` is declared in `export --help`, but no safe probe printed JSON (--help and --version override output flags in most CLIs)
  • opt_out: no option definition in --help declares --output or --format, the flags this requirement accepts; usage lines are not read. It gives the number of subcommand helps read when there were any, and adds the dash-rule note when a help declares -format beside double-dash names. It no longer says the tool does not ship structured output: a tool whose help declares --output-format would read that beside its own JSON mode.

The flag list is unchanged (--output, --format), and so are the safe probes (--help and --version only) and the [p2] json_probe hand-off. A --help that times out or crashes still leaves the row a skip.

For the reviewer: this PR moves scores. Nine tools were a skip only because a longer flag was mistaken for --output or --format. With no such flag declared they are an opt_out, which counts in the score at zero credit. Five of the nine (claude-code, dust, gemini-cli, ruff, yq) do select a format, through --output-format. Whether --output-format should satisfy p2-must-output-flag is a question about the requirement's flag list, which this series leaves as it is.

Changelog

Fixed

  • Fix p2-must-output-flag treating a longer flag such as --output-format, --output-file or --formatter as --output or --format, and then probing with a flag the tool does not have. An output flag counts when an option definition declares --output or --format.
  • Fix p2-must-output-flag missing Go flag help that declares -format or -output with one dash. The safe probes now pass the flag as the help spells it.

Changed

  • p2-must-output-flag skip evidence names the flag and the help that declares it, and opt_out evidence names what was searched. p2-must-schema-print quotes the same text when it follows that row.

Documentation

  • The README's json_probe section says the output flag is read from option definitions, is passed as the help spells it, and that a longer flag such as --output-format is a different name.

Type of Change

  • fix: Bug fix (non-breaking change which fixes an issue)

Related Issues/Stories

Testing

  • Unit tests added/updated
  • Integration tests added/updated
  • Manual testing completed
  • All tests passing

Test Summary:

  • Unit tests: five new in json_output.rs. A table of eight helps pairs each with the flags it declares as the probe must spell them. A stand-in that prints JSON only for --help -format json passes. A stand-in that would print JSON for --output json, and declares only --output-format, is an opt_out and is not probed. The skip and opt_out evidence is asserted whole, at the top level and for a subcommand, with the dash-rule note.
  • Seven existing tests declared their flag only in a usage line (Usage: test [--output FORMAT]). They now declare it on a definition line, and two assert the new evidence text.
  • 1311 passing across the suite. Self-audit: cargo test --test dogfood passes, and anc audited with this build keeps its own p2-must-output-flag pass. It now gets there through audit --help --output json: its top-level help mentions --output only in prose.
  • Snapshots: none change.

Negatives observed failing. The new tests, run with declared_output_flags put back to the substring match and everything else at this PR's head:

---- audits::behavioral::json_output::tests::a_single_dash_format_flag_is_probed_as_the_help_spells_it stdout ----
assertion `left == right` failed: got OptOut("no option definition in --help declares --output or --format, the flags this requirement accepts; usage lines are not read. The row is opt_out, and the schema-discovery requirements (p2-must-schema-print, p2-should-schema-file) collapse to n/a via antecedent propagation.")
  left: OptOut("no option definition in --help declares --output or --format, the flags this requirement accepts; usage lines are not read. The row is opt_out, and the schema-discovery requirements (p2-must-schema-print, p2-should-schema-file) collapse to n/a via antecedent propagation.")
 right: Pass
---- audits::behavioral::json_output::tests::output_flags_are_the_ones_the_help_declares_spelled_as_printed stdout ----
[
    "Go flag: one dash, and no double-dash name in the help: read [], declares [\"-format\"]",
    "a longer flag that starts with the name: read [\"--output\", \"--format\"], declares []",
    "a usage line: read [\"--output\"], declares []",
    "a sentence: read [\"--format\"], declares []",
]
---- audits::behavioral::json_output::tests::a_longer_flag_that_starts_with_output_does_not_trigger_the_probe stdout ----
assertion `left == right` failed
  left: Pass
 right: OptOut("no option definition in --help declares --output or --format, the flags this requirement accepts; usage lines are not read. The row is opt_out, and the schema-discovery requirements (p2-must-schema-print, p2-should-schema-file) collapse to n/a via antecedent propagation.")
test result: FAILED. 19 passed; 3 failed; 0 ignored; 0 measured; 1024 filtered out; finished in 1.08s

The third case is the one to read twice: with the substring match, a tool that declares only --output-format passes, because the probe sends it --output json.

The two evidence tests assert text this PR introduces. Before the rewording they failed against the old strings, --output/--format flag detected but could not validate JSON via safe probes (--help/--version override output flags in most CLIs) and no --output/--format flag detected in any subcommand — tool does not ship structured output.

Expected moves. Written before the corpus run. 184 rows: every p2-must-output-flag row that is not a pass, and the p2-must-schema-print row that follows each one.

id audit_id before after tools
p2-must-output-flag p2-json-output skip opt_out biome, claude-code, codex, cosign, dust, gemini-cli, make, ruff, yq (9)
p2-must-output-flag p2-json-output opt_out skip actionlint, terraform (2)
p2-must-output-flag p2-json-output skip skip, evidence only age, ast-grep, atuin, curl, deno, docker, fd, files-to-prompt, git-cliff, gum, helm, hyperfine, kubectl, llm, mise, mods, ollama, opencode, pandoc, pixi, qmd, rclone, scc, shellcheck, sqlite-utils, starship, supabase, tokei, typst, uv, vhs, xh, xr, xsv (34)
p2-must-output-flag p2-json-output opt_out opt_out, evidence only act, aws-cli, bandwhich, bat, bottom, broot, bun, cargo-binstall, cf, cmake, dasel, datasette, delta, direnv, doggo, eza, ffmpeg, flyctl, fzf, gh, git, gitleaks, gitui, glow, goose, jj, jnv, jq, just, lazygit, lsd, miller, miniserve, navi, nushell, pastel, procs, ripgrep, rsync, sd, shell-gpt, tealdeer, tmux, watchexec, wrangler, yazi, zoxide (47)
p2-must-schema-print p2-schema-print skip n_a the same 9
p2-must-schema-print p2-schema-print n_a skip actionlint, terraform (2)
p2-must-schema-print p2-schema-print skip skip, evidence only the same 34
p2-must-schema-print p2-schema-print n_a n_a, evidence only the same 47

Why each status moves:

tool what the substring matched what the help declares
claude-code --output-format no --output or --format, at the top level or in any subcommand
gemini-cli --output-format neither
yq --output-format neither
dust --output-format, --output-json neither
ruff --output-format and --output-file on check neither, in any subcommand
codex --output-schema and --output-last-message on exec neither, in any subcommand
cosign --output-file neither
make --output-sync neither
biome --formatter on rage neither, in any subcommand
actionlint nothing: -format string has one dash -format, in a help with no double-dash name; probed as -format json, which prints no JSON
terraform nothing -format=FORMAT on graph; probed as graph --help -format json, which prints no JSON

The three passes hold: anc, bird and trivy are probed with a flag their help declares, and their rows and the rows that follow them do not move.

Derived, computed from the base scorecards. An opt_out joins the score's denominator at zero credit, and a skip leaves it:

tool badge.score_pct band badge.eligible
biome 71 to 69 70-74 to 50-69 true to false
claude-code 70 to 68 70-74 to 50-69 true to false
codex 77 to 74 75-79 to 70-74 true, unchanged
cosign 65 to 63 50-69, unchanged false, unchanged
dust 78 to 75 75-79, unchanged true, unchanged
gemini-cli 73 to 71 70-74, unchanged true, unchanged
make 81 to 79 80-84 to 75-79 true, unchanged
ruff 80 to 78 80-84 to 75-79 true, unchanged
yq 76 to 74 75-79 to 70-74 true, unchanged
actionlint 79 to 81 75-79 to 80-84 true, unchanged
terraform 65 to 67 50-69, unchanged false, unchanged

For the nine, summary.skip falls by two, and summary.opt_out and summary.n_a rise by one each. For actionlint and terraform it is the reverse. No audience moves: the label counts warns among its four signal rows, and no p2-json-output row is a warn before or after.

How the prediction was made:

  • Statuses and evidence come from the base and head builds replayed over the full registry capture set. The stand-ins answer no JSON probe, so the replay shows which flag each build detects and where.
  • A second replay logged every call each build makes. The probes differ for the eleven tools above and for four that stay a skip: gum (choose to log, where --format is declared), uv (version to export), xh (--output alone, without --format-options) and anc.
  • For gum, uv, actionlint and terraform, the new probes were run against the real tools in the scorer image. None prints JSON, so each row is a skip. xh's new probes are a subset of its old ones, which printed none.

Corpus before/after. Run after the tables above were written. Base 58497bf5f381, head 4697a3c24879. Image sha256:05db16cf226cd14a93a2682aea41b72d3e6b4d240cadba006f2881d3ffa996b3, network bridge. 98 tools selected, 95 scored under both builds, 92 rerun three times.

Moved rows (184), by row and status

id audit_id status tools
p2-must-output-flag p2-json-output opt_out → skip actionlint, terraform (2)
p2-must-output-flag p2-json-output skip → opt_out biome, claude-code, codex, cosign, dust, gemini-cli, make, ruff, yq (9)
p2-must-output-flag p2-json-output opt_out, evidence only act, aws-cli, bandwhich, bat, bottom, broot, bun, cargo-binstall, cf, cmake, dasel, datasette, delta, direnv, doggo, eza, ffmpeg, flyctl, fzf, gh, git, gitleaks, gitui, glow, goose, jj, jnv, jq, just, lazygit, lsd, miller, miniserve, navi, nushell, pastel, procs, ripgrep, rsync, sd, shell-gpt, tealdeer, tmux, watchexec, wrangler, yazi, zoxide (47)
p2-must-output-flag p2-json-output skip, evidence only age, ast-grep, atuin, curl, deno, docker, fd, files-to-prompt, git-cliff, gum, helm, hyperfine, kubectl, llm, mise, mods, ollama, opencode, pandoc, pixi, qmd, rclone, scc, shellcheck, sqlite-utils, starship, supabase, tokei, typst, uv, vhs, xh, xr, xsv (34)
p2-must-schema-print p2-schema-print n_a → skip actionlint, terraform (2)
p2-must-schema-print p2-schema-print skip → n_a biome, claude-code, codex, cosign, dust, gemini-cli, make, ruff, yq (9)
p2-must-schema-print p2-schema-print n_a, evidence only act, aws-cli, bandwhich, bat, bottom, broot, bun, cargo-binstall, cf, cmake, dasel, datasette, delta, direnv, doggo, eza, ffmpeg, flyctl, fzf, gh, git, gitleaks, gitui, glow, goose, jj, jnv, jq, just, lazygit, lsd, miller, miniserve, navi, nushell, pastel, procs, ripgrep, rsync, sd, shell-gpt, tealdeer, tmux, watchexec, wrangler, yazi, zoxide (47)
p2-must-schema-print p2-schema-print skip, evidence only age, ast-grep, atuin, curl, deno, docker, fd, files-to-prompt, git-cliff, gum, helm, hyperfine, kubectl, llm, mise, mods, ollama, opencode, pandoc, pixi, qmd, rclone, scc, shellcheck, sqlite-utils, starship, supabase, tokei, typst, uv, vhs, xh, xr, xsv (34)

Derived fields moved (53), the 33 summary.* counts left out of the table

tool field base head
actionlint badge.score_pct 79 81
actionlint band 75-79 80-84
biome badge.score_pct 71 69
biome badge.eligible true false
biome band 70-74 50-69
claude-code badge.score_pct 70 68
claude-code badge.eligible true false
claude-code band 70-74 50-69
codex badge.score_pct 77 74
codex band 75-79 70-74
cosign badge.score_pct 65 63
dust badge.score_pct 78 75
gemini-cli badge.score_pct 73 71
make badge.score_pct 81 79
make band 80-84 75-79
ruff badge.score_pct 80 78
ruff band 80-84 75-79
terraform badge.score_pct 65 67
yq badge.score_pct 76 74
yq band 75-79 70-74

The 184 moved rows are the 184 expected. Compared row for row against the replay, each has the same status and the same evidence text before and after, subcommand counts included. The derived moves equal the predictions: the 20 fields above, and three summary counts for each of the eleven tools. No audience moved. Every moved row returned the same result in all four runs of each build. No row is unstable, no harness bug is flagged, and no tool failed to score under one build only. mods p3-should-about-long-about, the one row on the A/A noise list, did not move (base runs pass, warn, pass, pass; head runs pass, pass, pass, pass). cursor, nvidia-smi and xai-grok-build are absent from the image and ran under neither build.

Files Modified

Modified:

  • src/audits/behavioral/json_output.rs: declared_output_flags asks the definition query and returns the declared spellings; the probes take them; the unverified outcome is its own variant; evidence reworded.
  • README.md: the json_probe paragraph.

Created:

  • None.

Renamed:

  • None.

Deleted:

  • None.

Breaking Changes

  • No breaking changes

Eleven registry tools change status on p2-must-output-flag, and two of them (biome, claude-code) drop below the badge floor. Every other non-pass row carries new evidence text. The scorecard schema and JSON shape are unchanged.

Deployment Notes

  • No special deployment steps required

Checklist

  • Code follows project conventions and style guidelines
  • Commit messages follow Conventional Commits
  • Self-review of code completed
  • Tests added/updated and passing
  • No new warnings or errors introduced
  • Changes are backward compatible (or breaking changes documented)

…elled

`p2-json-output` decided whether a tool has an output flag by looking for the characters `--output` or `--format` anywhere in its lowercased `--help`, at the top level and then in each subcommand. A longer flag (`--output-format`, `--output-dir`, `--format-version`), a usage line and a sentence all counted, and the probe that followed always passed `--output json` or `--format json`, whatever the help printed. A Go `flag` help that declares `-format` was not detected at all.

Detection now asks the definition query, per help, and the safe probes pass each flag as the help spells it: `--help -format json` for a help that declares `-format` and no double-dash name. A longer flag that starts with `--output` or `--format` is a different name and triggers no probe.

The evidence names what was found or searched:

- skip: `` `--output` is declared in `export --help`, but no safe probe printed JSON (--help and --version override output flags in most CLIs) ``
- opt_out: `no option definition in --help declares --output or --format; usage lines are not read. The tool is scored as shipping no structured output, ...`, with the count of subcommand helps read when there were any, and the dash-rule note when a help declares `-format` beside double-dash names.

The unverified outcome is a variant of its own, so the declared `[p2] json_probe` takes over from it without comparing evidence strings. `--help` that times out or crashes still leaves the row a skip.
The opt-out evidence closed with `The tool is scored as shipping no structured output`. A tool whose help declares `--output-format` reads that sentence beside its own JSON mode, and the sentence is then false on its face.

The evidence now says which flags the requirement accepts and stops there: `no option definition in --help declares --output or --format, the flags this requirement accepts; usage lines are not read.`
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant