Skip to content

eval(bakeoff): shared driver for the #1269 model bake-off - #1282

Merged
jasonssdev merged 3 commits into
mainfrom
eval/1269-bakeoff-driver
Oct 2, 2026
Merged

jasonssdev merged 3 commits into
mainfrom
eval/1269-bakeoff-driver

Conversation

@jasonssdev

Copy link
Copy Markdown
Owner

Summary

The shared driver the #1269 pre-registration requires before the first bake-off run. evals/ only; no model was run.

1. Latency and identity stamping (all eight harnesses in the run plan)

  • query_attribution stores elapsed_s per answer and run_latencies_s per run; decision_revisions wraps its client in a timer and stores call_latencies_s / run_latencies_s; edge_typing, contradictions and adjudication persist the per-run latencies they already collected. Old stored rows still load.
  • evals/harness_stamp.py adds a "stamp" to every runs JSON: harness commit and dirty flag, model name and Ollama digest (loopback /api/tags only), and sha256(prompt)[:16] per prompt sent (ids like contradiction/system), read after any treatment swap. The definition matches feat: keep LLM prompts in one folder, one file each, versioned by content hash #1277's prompt hash and production's SUBJECT_PROMPT_VERSION / JUDGE_PROMPT_VERSION (asserted in the self-test).
  • query_sufficiency gains --arms quote, so the bake-off pays only for the shipped prompt.

2. evals/bakeoff/

  • bakeoff_spec.py: candidates and every bar as data, each citing the pre-registration comment.
  • bakeoff_bars.py: "no worse than baseline", vetoes, the waiver for bars the baseline already fails, the no-headroom rule, the knockout.
  • bakeoff_metrics.py: one extractor per harness; on stored results they reproduce the pre-registered baselines (edge accuracy 0.357, attribution 0.933, B1–B8 pass, auto_merge FAIL).
  • bakeoff_ollama.py: license, native-context and ollama ps memory gates, loopback only, never pulls.
  • run_bakeoff.py: --plan, --run, --evaluate-only, --self-test; resumable (done cells skipped, failed retried); the second baseline session only runs in a later invocation; an unworked adjudication queue blocks rather than drops a candidate; Write takes the top 2 Generate survivors.
  • Interpretation choices are listed in evals/bakeoff/README.md.

Related issue

Refs #1269

Type of change

  • feat — new feature
  • fix — bug fix
  • docs — documentation only
  • refactor — no behavior change
  • test — tests only
  • chore / ci — tooling, build, or CI
  • Breaking change

How was this tested?

Checklist

  • My commits follow Conventional Commits.
  • I added or updated tests for the change.
  • I updated docs where behavior, interfaces, or the knowledge model changed.
  • Lint, format, type check, and tests pass locally (ruff, mypy, pytest).
  • Output remains OKF-conformant and derived stores stay reconstructible from the bundle + sources.
  • The change is consistent with the project's guiding principles (local-first, provenance, freshness, human-in-the-loop).

… runs; record latency for attribution and revisions (Refs #1269)
…bility gates, bars as data and a resumable knockout (Refs #1269)
@jasonssdev
jasonssdev force-pushed the eval/1269-bakeoff-driver branch from 438f950 to 60fcbef Compare October 2, 2026 22:59
@jasonssdev
jasonssdev merged commit 29a3fba into main Oct 2, 2026
9 checks passed
@jasonssdev
jasonssdev deleted the eval/1269-bakeoff-driver branch October 2, 2026 23:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant