Skip to content

eval(resolution): reproduce #1223 on its field shapes; refute a one-sentence treatment - #1275

Merged
jasonssdev merged 1 commit into
mainfrom
fix/1223-contradiction-judge
Oct 2, 2026
Merged

jasonssdev merged 1 commit into
mainfrom
fix/1223-contradiction-judge

Conversation

@jasonssdev

Copy link
Copy Markdown
Owner

Summary

Measurement only: reproduces #1223 on the field shapes the issue quotes and refutes the smallest prompt treatment. No production prompt changes; resolution/contradiction.py is untouched.

  • New merged-content fixture classes in evals/contradictions/: 5 merged-scope-guidance and 5 merged-narrower-use cases (expected consistent, including the two verbatim pairs from the issue), plus 2 merged-contradiction guards in the same shapes.
  • New pure scoring functions field_shape_wrong and per_case_wrong with self-test coverage; the report prints n of TOTAL for the field-shape pool and a per-case table.
  • README section "Third attempt" with the pre-registered rule, results and verdict. The refuted sentence stays in contradiction_prompts.py, marked REFUTED, so --arm treatment reproduces.

Pre-registered rule (written before any treatment run)

qwen3:8b, production client settings, n=15 per arm, both arms on the full fixture in one session. Primary metric: wrong verdicts on field-shape probes (150 cells).

  • R0 reproduction: baseline wrong ≥ 15/150.
  • R1 effect: treatment ≤ 50% of baseline and an absolute drop ≥ 15 cells.
  • R2 retention: guards missed ≤ baseline + 2; typed TP and evaluative retention ≥ 0.95 and ≥ baseline − 0.05.
  • R3 no regression on older compatible classes and typed FP rates.
  • R4 latency ≤ 1.25× baseline.

Results

metric baseline treatment
field-shape wrong 44 of 150 43 of 150
older merged compatible wrong 0 of 120 0 of 120
guards missed 0 of 60 0 of 60
typed accuracy / TP / evaluative retention 1.00 1.00
typed FP rates 0.00 0.00
mean run latency 121.3s 120.6s

R0 holds (reproduced); R1 fails (1 cell, not ≥ 15); R2–R4 hold. All failures sit in three cases: both verbatim pairs are 15/15 wrong in both arms (contradicts at 0.95), and the synthetic Spanish narrower-use case is 14/15 → 13/15. A larger judge model is the next hypothesis and belongs to the #1269 bake-off.

Related issue

Refs #1223
Refs #1269

Type of change

  • feat — new feature
  • fix — bug fix
  • docs — documentation only
  • refactor — no behavior change
  • test — tests only
  • chore / ci — tooling, build, or CI
  • Breaking change

How was this tested?

  • Harness --self-test passes; 5 mutants on the new scoring and fixture guards are all killed.
  • ruff check, ruff format --check, mypy ., pytest --cov (96.46%), evals/run_self_tests.py (46/46) pass locally.

Checklist

  • My commits follow Conventional Commits.
  • I added or updated tests for the change.
  • I updated docs where behavior, interfaces, or the knowledge model changed.
  • Lint, format, type check, and tests pass locally (ruff, mypy, pytest).
  • Output remains OKF-conformant and derived stores stay reconstructible from the bundle + sources.
  • The change is consistent with the project's guiding principles (local-first, provenance, freshness, human-in-the-loop).

… narrower-statement treatment is refuted

Add merged-scope-guidance and merged-narrower-use cases (two verbatim public field pairs) plus two guards. The shipped judge is wrong 44 of 150 on them over 15 runs, concentrated in three cases, always contradicts at 0.95. The candidate sentence scored 43 of 150, missing the pre-registered bar (at least 15 fewer), so no prompt change ships.

Refs #1223
@jasonssdev
jasonssdev force-pushed the fix/1223-contradiction-judge branch from bcdaea4 to 2eee266 Compare October 2, 2026 19:26
@jasonssdev
jasonssdev merged commit bf57889 into main Oct 2, 2026
9 checks passed
@jasonssdev
jasonssdev deleted the fix/1223-contradiction-judge branch October 2, 2026 19:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant