Keep valid perspective fit ties and explain strict probe failures - #45
Merged
Merged
Conversation
…t-ranking-20261005
…t-ranking-20261005
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The final perspective-fit check rejected a valid batch when all five independently selected memes had equal fit scores, degrading the response despite non-flat upstream perspective rankings. Allow ties only in this final verification step; retain flat-score rejection for initial full-shortlist and perspective-choice ranking, malformed-score validation, safety behavior and confidence thresholds. This source defect is reproduced by a regression; the specific preceding hosted probe failure could not be conclusively attributed because the diagnostic tail helper stopped on its error.
Add allowlisted confidence, ranking mode, fallback and perspective-completeness fields to production probe receipts, without changing strict failure criteria or logging private input. No additional model call, budget increase, provider/model pin, production dependency or configuration change.
Validation: 76 focused tests pass, including equal strong final fits, preserved full-shortlist flat rejection, serious-input handling and prompt-free probe diagnostics. The shared Fleet disk cap refused the new full local run before executing it; existing exact-source CI runs the full suite, package validation and build before deployment; hosted strict monitoring follows release. Refs #30 and sass-maker/free-ai#95.