Skip to content

fix(audit): pass p1-flag-existence only when the help declares a gate flag - #172

Open
brettdavies wants to merge 1 commit into
fix/help-flags-u10-quietfrom
fix/help-flags-u10-flag-existence
Open

brettdavies wants to merge 1 commit into
fix/help-flags-u10-quietfrom
fix/help-flags-u10-flag-existence

Conversation

@brettdavies

@brettdavies brettdavies commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Summary

p1-flag-existence asks whether a tool that blocks on a bare call advertises a non-interactive gate flag (--no-interactive, --batch, -y, --yes, -p, --print and four more). It searched the whole --help text for each one, bounded by non-name characters. The bound stopped --print-json from answering for --print, but the search still passed a flag that the help only mentions:

  • in a usage line: Usage: tool [-y] [--batch] <file>
  • in a sentence: Run with --no-input in CI.
  • in another flag's wrapped description

It also missed Go flag help, where -batch is the only spelling printed.

The audit now asks the definition query for its gate flags, and contains_flag is removed. -p and -y answer only when declared as those letters. A single-dash word answers for its double-dash spelling in a help that declares no double-dash name, and beside double-dash names the warn says why it does not count. A warn names what was searched and where.

The audit's flag list is unchanged. So is the skip for a target that prints usage or exits on a bare call, which is why most tools never reach the lookup.

Changelog

Fixed

  • Fix p1-flag-existence passing when a non-interactive flag such as -y or --batch appears only in a usage line, a sentence, or another flag's description. It passes when an option definition declares one.
  • Fix p1-flag-existence missing a gate flag in Go flag help, where -batch is printed with one dash.

Changed

  • p1-flag-existence warn evidence now reads no option definition in --help declares a non-interactive flag (one of: ...); usage lines are not read.

Type of Change

  • fix: Bug fix (non-breaking change which fixes an issue)

Related Issues/Stories

Testing

  • Unit tests added/updated
  • Integration tests added/updated
  • Manual testing completed
  • All tests passing

Test Summary:

  • Unit tests: seven in flag_existence.rs, all through the helper run() calls. contains_flag's boundary cases carry over as negatives against the query (--print-json and --batching are not --print and --batch; -pr is not -p), beside the usage-line, sentence and wrapped-description cases. A Go-style -batch passes in a help with no double-dash name. A find-style help that declares --help beside -print warns and says -print does not count as --print. The warn evidence is asserted whole.
  • 1301 passing across the suite: the eight tests of the old matcher and its test-only twin of the audit are replaced by seven. Self-audit: cargo test --test dogfood passes.
  • Snapshots: none change.

Negatives observed failing. The new tests, run against the base (#171) with the audit restructured around one helper and contains_flag still doing the match:

---- audits::behavioral::flag_existence::tests::a_single_dash_word_beside_double_dash_names_warns_and_says_why stdout ----
no non-interactive flag found in --help; expected one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes
---- audits::behavioral::flag_existence::tests::a_single_dash_word_passes_in_a_help_without_double_dash_names stdout ----
assertion `left == right` failed
  left: Warn("no non-interactive flag found in --help; expected one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes")
 right: Pass
---- audits::behavioral::flag_existence::tests::a_gate_flag_named_but_not_declared_does_not_pass stdout ----
["a usage line", "another flag's wrapped description", "a sentence"]
---- audits::behavioral::flag_existence::tests::warn_names_the_flags_searched_and_where stdout ----
assertion `left == right` failed
  left: Warn("no non-interactive flag found in --help; expected one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes")
 right: Warn("no option definition in --help declares a non-interactive flag (one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes); usage lines are not read.")
test result: FAILED. 3 passed; 4 failed; 0 ignored; 0 measured; 1029 filtered out; finished in 0.00s

The boundary cases (--print-json, --batching, -pr) passed at the base, as the old matcher already rejected them.

Expected moves. Written before the corpus run. Two rows, evidence only:

tool id audit_id before after why
datasette p1-must-no-interactive p1-flag-existence warn warn, evidence only reworded from no non-interactive flag found in --help; expected one of: ... to the search-scope sentence
opencode p1-must-no-interactive p1-flag-existence warn warn, evidence only the same rewording

No status moves, so no derived field moves. How the prediction was made:

  • Of the 95 scored tools, 82 skip this row because the bare call prints usage or exits, and 11 skip it under the human-tui profile. The flag lookup runs for datasette and opencode alone, whose bare calls block.
  • The capture replay cannot reproduce a bare call that blocks, and it moves no row. The two predictions come from stand-ins that serve each tool's captured --help and hang on a bare call, audited with the base and head builds: both warn under both builds, with the old and new evidence.
  • p1-flag-existence is not one of the four audience signals and is no row's antecedent.

Corpus before/after. Run after the table above was written. Base 682c69ed3ac3, head 972b5b34f5aa. Image sha256:05db16cf226cd14a93a2682aea41b72d3e6b4d240cadba006f2881d3ffa996b3, network bridge. 98 tools selected, 95 scored under both builds, 3 rerun three times (datasette, mods, opencode).

Moved rows (2)

tool id audit_id status confidence evidence note
datasette p1-must-no-interactive p1-flag-existence warn high base: no non-interactive flag found in --help; expected one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes
head: no option definition in --help declares a non-interactive flag (one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes); usage lines are not read.
opencode p1-must-no-interactive p1-flag-existence warn high base: no non-interactive flag found in --help; expected one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes
head: no option definition in --help declares a non-interactive flag (one of: --no-interactive, --non-interactive, -p, --print, --no-input, --batch, --headless, -y, --yes, --assume-yes); usage lines are not read.

Derived fields moved (0)

None.

The two moved rows are the two expected, evidence only, and no derived field moved. Each returned the same result in all four runs of each build. No row is unstable, no harness bug is flagged, and no tool failed to score under one build only. mods p3-should-about-long-about, the one row on the A/A noise list, is reported as noise: both builds returned pass and warn for it (base runs pass, warn, warn, warn; head runs pass, warn, pass, pass). The row compares the tool's -h and --help output and reads no flag. cursor, nvidia-smi and xai-grok-build are absent from the image and ran under neither build.

Files Modified

Modified:

  • src/audits/behavioral/flag_existence.rs: audit_flag_existence asks the definition query and is the helper run() calls; contains_flag and the test-only copy of the audit removed.

Created:

  • None.

Renamed:

  • None.

Deleted:

  • None.

Breaking Changes

  • No breaking changes

Two registry rows carry new evidence text. The scorecard schema and JSON shape are unchanged.

Deployment Notes

  • No special deployment steps required

Checklist

  • Code follows project conventions and style guidelines
  • Commit messages follow Conventional Commits
  • Self-review of code completed
  • Tests added/updated and passing
  • No new warnings or errors introduced
  • Changes are backward compatible (or breaking changes documented)

… flag

`p1-flag-existence` searched the whole `--help` text for each non-interactive gate flag, bounded by non-name characters. That stopped `--print-json` from answering for `--print`, but it still passed a flag shown only in a usage line (`Usage: tool [-y] <file>`), in a sentence, or in another flag's wrapped description. It also missed Go `flag` help, where `-batch` is the only spelling printed.

The audit now asks the definition query for its gate flags, and `contains_flag` is removed. A single letter such as `-p` or `-y` answers only when declared as that letter, and a single-dash word answers for its double-dash spelling in a help that declares no double-dash name. A warn says what was searched and where: `no option definition in --help declares a non-interactive flag (one of: ...); usage lines are not read.`

`run()` and the tests share one core helper. The test-only copy of the audit's logic is gone.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant