Repository navigation
fix: repair flat shortlist scores within the existing inference budget - #48
Merged
Merged
Conversation
3 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A live perspective request returned the same ordinal fit score for every candidate. The ranker rejected it after the managed adapter had already accepted the response, so the adapter's existing one-time invalid-output repair never ran.
Opt shortlist ordinal calls into the same flat-score validation inside the managed adapter. Invalid flat output can use the existing second call under the same deadline; a second flat answer is still rejected. Final selected winners may legitimately tie and remain exempt. Safety decisions, generic classification, model/provider automation, token limits, and attempt ceilings are preserved. The flat-output diagnostic includes only allowlisted termination metadata.
Validation: real ranker-to-adapter regressions cover recovery, exhausted repair, equal expected fit from different distributions, legitimate final ties, safety, and cancellation. The prior live 7/8 sample is preserved as a failure, not counted as healthy. Full CI and new live qualification remain required.