Skip to content

fix(audit): credit a non-interactive flag only when the help declares it - #173

Open
brettdavies wants to merge 1 commit into
fix/help-flags-u10-flag-existencefrom
fix/help-flags-u10-non-interactive
Open

brettdavies wants to merge 1 commit into
fix/help-flags-u10-flag-existencefrom
fix/help-flags-u10-non-interactive

Conversation

@brettdavies

@brettdavies brettdavies commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Summary

p1-non-interactive passes a tool whose bare call blocks or crashes when its --help shows a headless flag. It looked for eleven flag spellings and four spacing variants (-y,, -y , -p,, -p) as substrings of the help text, so the help only had to contain the characters:

  • opencode's bare call starts a TUI, and it passed on --print-logs, which contains --print.
  • --batch-size and --yes-i-really-mean-it contain --batch and --yes.
  • A usage line (Usage: foo [--no-interactive] <file>), an example (foo -p "summarize this") and a sentence (Pass -y to skip the confirmation.) each matched.

The audit now asks the definition query for the same flags. The spacing variants are gone: -y and -p answer when declared as those letters. A single-dash word answers for its double-dash spelling in a help that declares no double-dash name, and beside double-dash names the warn says why it does not count.

Each warn says what the bare call did and what was searched. The flag list is unchanged, and so are the other two ways to pass: a bare call that exits, or one that prints usage.

p1-non-interactive is one of the four audits behind the scorecard's audience, so a pass that becomes a warn is one more warn in that count.

Changelog

Fixed

  • Fix p1-non-interactive passing a tool whose bare call blocks because its help contains a headless flag's characters: inside a longer flag (--print-logs, --batch-size), in a usage line, in an example, or in a sentence. A headless flag counts when an option definition declares it.
  • Fix p1-non-interactive missing a headless flag in Go flag help, where -batch is printed with one dash.

Changed

  • p1-non-interactive warn evidence now says what the bare call did and what was searched, for example bare invocation timed out, so the binary may be waiting for interactive input, and no option definition in --help declares a non-interactive flag (one of: ...); usage lines are not read.

Type of Change

  • fix: Bug fix (non-breaking change which fixes an issue)

Related Issues/Stories

Testing

  • Unit tests added/updated
  • Integration tests added/updated
  • Manual testing completed
  • All tests passing

Test Summary:

  • Unit tests: six new in non_interactive.rs, through the core helper run() calls. -p, --print, -y, --yes and --no-input pass a bare call that timed out. Five helps that name a flag without declaring it do not, opencode 1.18.34's real --print-logs line among them. A Go-style -batch passes in a help with no double-dash name. A find-style help that declares --help beside -print warns and says -print does not count as --print. All three warn sentences are asserted whole. The run() test for a declared flag now declares it on a definition line and crashes on the bare call.
  • 1306 passing across the suite. Self-audit: cargo test --test dogfood passes.
  • Snapshots: none change.

Negatives observed failing. The new tests, run against the base (#172) with the audit restructured around the core helper and the substring markers still doing the match:

---- audits::behavioral::non_interactive::tests::a_single_dash_word_passes_in_a_help_without_double_dash_names stdout ----
assertion `left == right` failed
  left: Warn("binary may be waiting for interactive input")
 right: Pass
---- audits::behavioral::non_interactive::tests::a_single_dash_print_beside_double_dash_names_warns_and_says_why stdout ----
binary may be waiting for interactive input
---- audits::behavioral::non_interactive::tests::each_warn_says_what_the_bare_call_did_and_what_was_searched stdout ----
assertion `left == right` failed
  left: "binary may be waiting for interactive input"
 right: "bare invocation timed out, so the binary may be waiting for interactive input, and no option definition in --help declares a non-interactive flag (one of: --no-interactive, --non-interactive, --batch, --headless, --yes, --no-input, --no-browser, --device-code, -y, -p, --print); usage lines are not read."
---- audits::behavioral::non_interactive::tests::an_agentic_flag_named_but_not_declared_does_not_pass stdout ----
["opencode 1.18.34: a longer flag that starts with the name", "longer flags that start with --batch and --yes", "a usage line", "an example", "a sentence"]
test result: FAILED. 7 passed; 4 failed; 0 ignored; 0 measured; 1030 filtered out; finished in 3.17s

Expected moves. Written before the corpus run. Two rows:

tool id audit_id before after why
opencode p1-must-no-interactive p1-non-interactive pass warn its bare call starts a TUI and times out; the only match was --print-logs containing --print, and no definition declares a headless flag
datasette p1-must-no-interactive p1-non-interactive warn warn, evidence only reworded from binary may be waiting for interactive input to the sentence that also says what was searched

Derived, computed from the base scorecard:

tool field before after
opencode audience agent-optimized mixed
opencode badge.score_pct 77 75
opencode summary.pass 17 16
opencode summary.warn 15 16

opencode's band (75-79) and badge.eligible (true) hold. Its audience moves because the new warn joins its existing p7-quiet warn: two warns among the four signal rows is mixed. No propagated row moves: p1-must-no-interactive is no row's antecedent.

How the prediction was made:

  • The flag lookup decides this row only for a tool whose bare call times out or crashes. Of the 95 scored tools, 83 pass because the bare call exits, and 11 are suppressed under the human-tui profile. That leaves datasette and opencode.
  • The capture replay cannot reproduce a bare call that blocks, and it moves no row. The two predictions come from stand-ins that serve each tool's captured --help and hang on a bare call, audited with the base and head builds.
  • This row warns on a timeout, so it is the timing-sensitive kind the harness reruns. The result counts when all four runs of each build agree.

Corpus before/after. Run after the table above was written. Base 972b5b34f5aa, head 58497bf5f381. Image sha256:05db16cf226cd14a93a2682aea41b72d3e6b4d240cadba006f2881d3ffa996b3, network bridge. 98 tools selected, 95 scored under both builds, 3 rerun three times (datasette, mods, opencode).

Moved rows (2)

tool id audit_id status confidence evidence note
datasette p1-must-no-interactive p1-non-interactive warn high base: binary may be waiting for interactive input
head: bare invocation timed out, so the binary may be waiting for interactive input, and no option definition in --help declares a non-interactive flag (one of: --no-interactive, --non-interactive, --batch, --headless, --yes, --no-input, --no-browser, --device-code, -y, -p, --print); usage lines are not read.
opencode p1-must-no-interactive p1-non-interactive pass → warn high base: null
head: bare invocation timed out, so the binary may be waiting for interactive input, and no option definition in --help declares a non-interactive flag (one of: --no-interactive, --non-interactive, --batch, --headless, --yes, --no-input, --no-browser, --device-code, -y, -p, --print); usage lines are not read.

Derived fields moved (4)

tool field base head
opencode audience agent-optimized mixed
opencode badge.score_pct 77 75
opencode summary.pass 17 16
opencode summary.warn 15 16

The two moved rows and the four derived moves equal the expected moves. Each returned the same result in all four runs of each build, so the timeout behind both rows was stable. No row is unstable, no harness bug is flagged, and no tool failed to score under one build only. mods p3-should-about-long-about, the one row on the A/A noise list, did not move (base runs pass, warn, pass, pass; head runs warn, warn, pass, pass). cursor, nvidia-smi and xai-grok-build are absent from the image and ran under neither build.

Files Modified

Modified:

  • src/audits/behavioral/non_interactive.rs: audit_non_interactive is the core helper; the flag list is matched through the definition query; the spacing markers are removed.

Created:

  • None.

Renamed:

  • None.

Deleted:

  • None.

Breaking Changes

  • No breaking changes

One registry tool loses a p1-non-interactive pass it got from --print-logs, and its audience label moves with it. The scorecard schema and JSON shape are unchanged.

Deployment Notes

  • No special deployment steps required

Checklist

  • Code follows project conventions and style guidelines
  • Commit messages follow Conventional Commits
  • Self-review of code completed
  • Tests added/updated and passing
  • No new warnings or errors introduced
  • Changes are backward compatible (or breaking changes documented)

`p1-non-interactive` passes a tool whose bare call blocks or crashes when its `--help` shows a headless flag. It looked for eleven flag spellings and four spacing variants (`-y,`, `-y `, ` -p,`, ` -p `) as substrings of the help text, so opencode, whose bare call starts a TUI, passed on `--print-logs` containing `--print`. `--batch-size`, `--yes-i-really-mean-it`, a usage line, an example such as `foo -p "..."` and a sentence such as `Pass -y to skip` all passed the same way.

The audit now asks the definition query for the same flags. The spacing variants are gone: `-y` and `-p` answer when declared as those letters. A single-dash word answers for its double-dash spelling in a help that declares no double-dash name.

Each warn says what the bare call did and what was searched, for example `bare invocation timed out, so the binary may be waiting for interactive input, and no option definition in --help declares a non-interactive flag (one of: ...); usage lines are not read.`

`run()` hands the bare call's outcome and the shared help probe to one core helper, which the tests call.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant