Repository navigation
Outcome bench on Claude Code, and the Claude fixes it surfaced - #2
Merged
Merged
Conversation
3dgiordano
force-pushed
the
claude/amazing-dijkstra-g938ep
branch
from
September 27, 2026 13:48
b25a253 to
0cd66e5
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Runs the outcome bench on Claude Code (Sonnet 5, Opus 5, Opus 5.5, all High), then fixes what those runs showed: host bugs in the Claude adapters, lexicon misses on Claude's closes, one plugin message that ended Opus 5's runs, a new anchored signal for executive, and grader/audit defects that penalised Claude's formats.
Tooling (bench, not CI)
scripts/claude-bench.js: outcome driver for Claude (driverclaudeinbench/suite.json). Each invocation gets a scratch HOME and an allowlisted environment, so the baseline has no installed plugin, no user hooks and no inherited session id. Grants live in that HOME,nodeis fenced to the workspace, and the account's 5h/7d quota is read from the stream (the run waits or stops before exhausting it). Options:--rescore(re-audit and re-grade stored runs, no calls),--plugin-env(measure an opt-in strict gate).scripts/claude-hooks.jsreads it to show what each hook said, when, after which tool, plus skill loads, blocks and Stop verdicts.bench.jspicks the driver per suite row and re-reads the session before storing a slot. A case that is run again starts from a clean directory (both drivers).Plugins (Claude Code)
<task-notification>prompt no longer counts as a user turn (measured: 20 of 23 prompts in one cloud session were notifications).SessionEndonly sweeps by age now, and a newSessionStart(resume|compact)reloads the discipline, as Cursor'ssessionStartdoes. Before this, the retrospective the Stop hook parked (and promised to the user) was lost on resume.SECURITY.md.[reasoning_extraction]in 11 of 12 runs; without it, 0 of 3.Grader and audit
epistemic/the-cause-i-named-first: joins hard-wrapped prose, reads**Status:** conjecture, and treats quotes, asides and decision-table arrows as non-claims. Fixtures were added first. Every changed stored verdict was read by hand.find /) still void a run.Measured (n=3 per arm per model, two sessions pooled, n=18 per arm):
Details and numbers are in
CHANGELOG.mdunder Unreleased.Open, left for a maintainer decision:
Checklist
node scripts/test.jspasses locally (149/149; alsohosts.js,samples.js,corpus.js,bench-check.js)node scripts/version.js --checkpasses; plugin version bumped if behaviour changedCHANGELOG.mdhas a line under Unreleased (skip for docs-only)hosts.js --checkpasses, and the Cursor driver got the same clean-case fix.