Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Mixley Eval

The evaluation harness for the Mixley dual-model pipeline.

Status

In development. The Mixley docs cite relative improvement figures (+16% with Claude as the final model, +11% with GPT) as illustrative example values, not as published benchmark results. This repository is the framework that produces those figures; results will be published once the harness is run against settled model versions.

What we evaluate

The dual-model pipeline runs three stages: reasoning, critique and synthesis. The evaluation measures whether the final synthesized answer is more reliable, more nuanced and more transparent than either model alone.

Methodology

  • A set of prompts is run through each single model and through the dual-model pipeline
  • Each output is scored on correctness, completeness, reasoning quality and hallucination rate
  • The relative improvement is the score of the dual-model answer over the score of the single-model answer, on the same prompt

Files

  • eval/harness.ts - the evaluation script (a skeleton; not yet run against live models)
  • eval/metrics.md - the metrics the harness records

License

MIT - see LICENSE.

About

Evaluation harness for the Mixley dual-model pipeline

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages