fix(eval): bake-off memory gate reloads the embedder and estimates gemma4 instead of trusting a bad snapshot - #1285
Merged
Merged
Conversation
The 24 GB budget holds the chat model, its KV and bge-m3 together, so a snapshot without bge-m3 resident is ineligible, and a candidate reading under 90% of its on-disk weights is an invalid measurement (pending), never eligible. Eligibility records carry a version so records written under the old rules are re-measured on the next run. Refs #1269
…bad snapshot A snapshot without bge-m3 is an invalid reading, not a failure: reload the embedder (bounded, twice) and re-snapshot, else add bge-m3's own measured size and mark the figure estimated. A candidate reading under 90% of its on-disk weights (gemma4) is estimated as disk + reported. Estimates are judged against 24 GB and shown in the eligibility JSON, the plan and the report; a missing disk size stays pending. Eligibility records move to version 3 so all current records re-measure. Refs #1269
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes two measurement defects in the #1269 bake-off's memory eligibility gate (rule 6.3: 24 GB must hold the chat model, its KV cache at production settings and
bge-m3together), found in the first live run before any candidate cell ran.bge-m3missing from theollama pssnapshot. Forqwen3.6:35b-a3batnum_ctx12288 Ollama had evicted the embedder, and the gate summed only what was listed. Now a snapshot withoutbge-m3is an invalid reading: the driver reloads the embedder (a minimal embed call) and re-snapshots, at most twice. If it still does not co-reside, the embedder's own measured size is added and the figure is markedestimated: bge-m3 not co-resident. Eligibility is decided by the sum against 24 GB at eachnum_ctx. On a 48 GB machine an eviction at ~23 GB reflects load order, not memory pressure, so it is not treated as ineligibility./api/tagssize is now invalid; its footprint is estimated as on-disk size + the reported size +bge-m3, markedestimated (disk + reported), and judged against 24 GB. With no on-disk size, nobge-m3size, or no resident candidate, the model stays pending — never eligible.Estimates are visible everywhere: the eligibility JSON carries
memory_method/memory_estimatedand the full per-context check, the plan note says(memory estimated (...)), and the report prefixes estimated GB with~and adds a memory-method column. Eligibility records are now version 3, so every earlier record is re-measured on the next--run; nothing needs deleting by hand. The baseline-pair figure subtracts the embedder once.These rules were decided by the maintainer before any candidate run.
Related issue
Refs #1269
Type of change
feat— new featurefix— bug fixdocs— documentation onlyrefactor— no behavior changetest— tests onlychore/ci— tooling, build, or CIHow was this tested?
qwen3.6:35b-a3bat 12288 withoutbge-m3→ estimated 22.93 GB, eligible;gemma4:12b→ ~9.7 GB estimated;gemma4:26b-a4b→ ~20.4 GB estimated; a correctmistral-small3.2:24breading stays measured; a reading with no disk size stays pending. Against a fake server: one reload fixes a transient eviction; a sticky one stops after exactly two reloads and estimates; an estimate over 24 GB is ineligible with the reason saying so.ruff check,ruff format --check,mypy .,evals/run_self_tests.py(52/52) andpytest --cov(96.46%) pass; no model was loaded.Checklist
ruff,mypy,pytest).