Skip to content

eval: #1258 hard negatives for adjudication and a direction metric for edge typing - #1272

Merged
jasonssdev merged 2 commits into
mainfrom
eval/1269-prereqs
Oct 2, 2026
Merged

jasonssdev merged 2 commits into
mainfrom
eval/1269-prereqs

Conversation

@jasonssdev

Copy link
Copy Markdown
Owner

Summary

The two prerequisites #1269 lists before any bake-off run. Only evals/ changes; no engine code, and no model was run, so no recall or accuracy numbers are reported.

A. #1258-shaped hard negatives in the identity adjudication fixture (evals/adjudication/)

  • New probe class procedure-about (3 pairs): a Concept against a Procedure about it (e.g. "Lumen Toolkit" vs "Installing Lumen Toolkit"). Each procedure carries type_alternative="Concept" so the production cross-type bridge nominates it.
  • New probe class ui-component (2 pairs): a framework against its own UI, both Concepts (e.g. "Tessera" vs "Tessera Web UI") — the part-whole exclusion the rubric already states.
  • Every new pair is expected different; all content is synthetic. Stored arms predate these classes and are not comparable with runs that include them; the README says so.

B. A direction metric for edge_typing (evals/edge_typing/)

  • direction_pairs() pairs each forward edge with an asymmetric label to its reversed twin whose trap_type is that label: 9 pairs covering 18 of 29 labelled edges. Symmetric labels, abstentions and forward edges with no reversed probe are excluded, not counted as passes.
  • score_direction() classifies each pair-run as discriminated, blind (same asymmetric type both ways) or abstained (no correct forward claim). The existing trap-hit count only looks at reversed probes, so it cannot tell a direction-blind model from one that is right forwards.
  • The report prints "Defined on 18 of 29 labelled edges" and each outcome as n of total pair-runs; counts go into runs-*.json.

Judgement calls worth a look: the new classes are small (rates move in steps of 1/3 and 1/2); they are negatives only, relying on the existing *-same controls; and a wrong asymmetric forward type lands in abstained rather than being charged to direction.

Related issue

Refs #1269

Type of change

  • feat — new feature
  • fix — bug fix
  • docs — documentation only
  • refactor — no behavior change
  • test — tests only
  • chore / ci — tooling, build, or CI
  • Breaking change

How was this tested?

  • adjudication --self-test: 27 labelled pairs across 11 probe classes materialize into exactly that many candidate groups.
  • edge_typing --self-test: 11/11, including oracle, direction-blind and all-related_to synthetic models with known expected counts. Mutation: flipping the reversed-answer comparison in score_direction turns it red (9/11).
  • ruff check, ruff format --check, mypy ., pytest --cov (96.46%), evals/run_self_tests.py (46/46) pass locally.

Checklist

  • My commits follow Conventional Commits.
  • I added or updated tests for the change.
  • I updated docs where behavior, interfaces, or the knowledge model changed.
  • Lint, format, type check, and tests pass locally (ruff, mypy, pytest).
  • Output remains OKF-conformant and derived stores stay reconstructible from the bundle + sources.
  • The change is consistent with the project's guiding principles (local-first, provenance, freshness, human-in-the-loop).

@jasonssdev
jasonssdev merged commit d191b27 into main Oct 2, 2026
9 checks passed
@jasonssdev
jasonssdev deleted the eval/1269-prereqs branch October 2, 2026 19:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant