Hi — I’m Peter Kruger from autobench.org. We’re building a cost-aware router and would like to submit it to RouterArena.
The pool we plan to route is:
deepseek/deepseek-v4-flash-0731
openai/gpt-5.6-luna
openai/gpt-5.6-sol
On current main (cff9659d, 2026-08-29) luna and sol aren’t in universal_model_names.py or model_cost.json. deepseek/deepseek-v4-flash is there (priced $0.14 / $0.28); the dated -0731 slug is not.
Could you add luna and sol, and tell us whether -0731 should be treated as the existing flash row or as its own key?
Prices we’d like you to pin (standard list, with a source, per #193):
| Key |
$/MTok in / out |
Source |
openai/gpt-5.6-luna |
0.20 / 1.20 |
OpenAI pricing, OpenRouter |
openai/gpt-5.6-sol |
2.00 / 10.00 |
OpenRouter. OpenAI’s page currently shows a promotional $4 / $20 through at least 2026-11-21 — we’d rather pin the non-promo list if you agree. |
deepseek/deepseek-v4-flash-0731 |
0.05 / 0.16 |
OpenRouter as of 2026-09-07. Only needed if you treat -0731 as its own artifact. |
Two process questions: should we add those rows in our submission PR, or do you prefer to pre-register? And we’ll populate generated_result ourselves; we can add model_inference.py endpoints if that’s useful.
On a separate note: while building the router we’ve also been running our own benchmark, earlier-stage than yours. In doing so we are using the same methodology we've always used for our benchmarks: we write open-ended questions, have a committee of models judge answers from a fixed 20-model pool, and then score a router by looking up the cell it picked (no generation at score time). So far we’ve run Not Diamond, OpenRouter Auto, Orca, Azure Model Router, NadirClaw, and Brick. Happy to share notes if that’s ever useful.
Thanks for the work you’ve put into the arena. Happy to jump on a call if that’s easier.
Peter Kruger
autobench.org
Hi — I’m Peter Kruger from autobench.org. We’re building a cost-aware router and would like to submit it to RouterArena.
The pool we plan to route is:
deepseek/deepseek-v4-flash-0731openai/gpt-5.6-lunaopenai/gpt-5.6-solOn current
main(cff9659d, 2026-08-29) luna and sol aren’t inuniversal_model_names.pyormodel_cost.json.deepseek/deepseek-v4-flashis there (priced $0.14 / $0.28); the dated-0731slug is not.Could you add luna and sol, and tell us whether
-0731should be treated as the existing flash row or as its own key?Prices we’d like you to pin (standard list, with a source, per #193):
openai/gpt-5.6-lunaopenai/gpt-5.6-soldeepseek/deepseek-v4-flash-0731-0731as its own artifact.Two process questions: should we add those rows in our submission PR, or do you prefer to pre-register? And we’ll populate
generated_resultourselves; we can addmodel_inference.pyendpoints if that’s useful.On a separate note: while building the router we’ve also been running our own benchmark, earlier-stage than yours. In doing so we are using the same methodology we've always used for our benchmarks: we write open-ended questions, have a committee of models judge answers from a fixed 20-model pool, and then score a router by looking up the cell it picked (no generation at score time). So far we’ve run Not Diamond, OpenRouter Auto, Orca, Azure Model Router, NadirClaw, and Brick. Happy to share notes if that’s ever useful.
Thanks for the work you’ve put into the arena. Happy to jump on a call if that’s easier.
Peter Kruger
autobench.org