The evaluation harness for the Mixley dual-model pipeline.
In development. The Mixley docs cite relative improvement figures (+16% with Claude as the final model, +11% with GPT) as illustrative example values, not as published benchmark results. This repository is the framework that produces those figures; results will be published once the harness is run against settled model versions.
The dual-model pipeline runs three stages: reasoning, critique and synthesis. The evaluation measures whether the final synthesized answer is more reliable, more nuanced and more transparent than either model alone.
- A set of prompts is run through each single model and through the dual-model pipeline
- Each output is scored on correctness, completeness, reasoning quality and hallucination rate
- The relative improvement is the score of the dual-model answer over the score of the single-model answer, on the same prompt
eval/harness.ts- the evaluation script (a skeleton; not yet run against live models)eval/metrics.md- the metrics the harness records
MIT - see LICENSE.