Skip to content

fix(eval): bake-off memory gate reloads the embedder and estimates gemma4 instead of trusting a bad snapshot - #1285

Merged
jasonssdev merged 2 commits into
mainfrom
fix/bakeoff-memory-gate
Oct 3, 2026
Merged

jasonssdev merged 2 commits into
mainfrom
fix/bakeoff-memory-gate

Conversation

@jasonssdev

Copy link
Copy Markdown
Owner

Summary

Fixes two measurement defects in the #1269 bake-off's memory eligibility gate (rule 6.3: 24 GB must hold the chat model, its KV cache at production settings and bge-m3 together), found in the first live run before any candidate cell ran.

  1. bge-m3 missing from the ollama ps snapshot. For qwen3.6:35b-a3b at num_ctx 12288 Ollama had evicted the embedder, and the gate summed only what was listed. Now a snapshot without bge-m3 is an invalid reading: the driver reloads the embedder (a minimal embed call) and re-snapshots, at most twice. If it still does not co-reside, the embedder's own measured size is added and the figure is marked estimated: bge-m3 not co-resident. Eligibility is decided by the sum against 24 GB at each num_ctx. On a 48 GB machine an eviction at ~23 GB reflects load order, not memory pressure, so it is not treated as ineligibility.
  2. gemma4 reads ~1 GB resident against 8.0 GB / 18.7 GB on disk (Ollama 0.32.9 appears to report only non-weight allocations for gemma4; unproven without a load). A candidate reading under 90% of its /api/tags size is now invalid; its footprint is estimated as on-disk size + the reported size + bge-m3, marked estimated (disk + reported), and judged against 24 GB. With no on-disk size, no bge-m3 size, or no resident candidate, the model stays pending — never eligible.

Estimates are visible everywhere: the eligibility JSON carries memory_method / memory_estimated and the full per-context check, the plan note says (memory estimated (...)), and the report prefixes estimated GB with ~ and adds a memory-method column. Eligibility records are now version 3, so every earlier record is re-measured on the next --run; nothing needs deleting by hand. The baseline-pair figure subtracts the embedder once.

These rules were decided by the maintainer before any candidate run.

Related issue

Refs #1269

Type of change

  • feat — new feature
  • fix — bug fix
  • docs — documentation only
  • refactor — no behavior change
  • test — tests only
  • chore / ci — tooling, build, or CI
  • Breaking change

How was this tested?

  • Model-free self-tests built from the exact live readings: qwen3.6:35b-a3b at 12288 without bge-m3 → estimated 22.93 GB, eligible; gemma4:12b → ~9.7 GB estimated; gemma4:26b-a4b → ~20.4 GB estimated; a correct mistral-small3.2:24b reading stays measured; a reading with no disk size stays pending. Against a fake server: one reload fixes a transient eviction; a sticky one stops after exactly two reloads and estimates; an estimate over 24 GB is ineligible with the reason saying so.
  • 14 mutations (reload bound both ways, disk+reported → reported only, embedder addition dropped or unmarked, 90% floor off, invalid branch off, budget gate off for estimates, missing-disk guard off, version check ignored in both places, plan note hidden, disk size not passed, embedder double-counted in the pairing) are all killed.
  • ruff check, ruff format --check, mypy ., evals/run_self_tests.py (52/52) and pytest --cov (96.46%) pass; no model was loaded.

Checklist

  • My commits follow Conventional Commits.
  • I added or updated tests for the change.
  • I updated docs where behavior, interfaces, or the knowledge model changed.
  • Lint, format, type check, and tests pass locally (ruff, mypy, pytest).
  • Output remains OKF-conformant and derived stores stay reconstructible from the bundle + sources.
  • The change is consistent with the project's guiding principles (local-first, provenance, freshness, human-in-the-loop).

The 24 GB budget holds the chat model, its KV and bge-m3 together, so a
snapshot without bge-m3 resident is ineligible, and a candidate reading
under 90% of its on-disk weights is an invalid measurement (pending),
never eligible. Eligibility records carry a version so records written
under the old rules are re-measured on the next run.

Refs #1269
…bad snapshot

A snapshot without bge-m3 is an invalid reading, not a failure: reload
the embedder (bounded, twice) and re-snapshot, else add bge-m3's own
measured size and mark the figure estimated. A candidate reading under
90% of its on-disk weights (gemma4) is estimated as disk + reported.
Estimates are judged against 24 GB and shown in the eligibility JSON,
the plan and the report; a missing disk size stays pending. Eligibility
records move to version 3 so all current records re-measure.

Refs #1269
@jasonssdev
jasonssdev merged commit 2118514 into main Oct 3, 2026
9 checks passed
@jasonssdev
jasonssdev deleted the fix/bakeoff-memory-gate branch October 3, 2026 03:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant