Not sure which model to use?
Ori Eval is live on OpenRouter. If you already route models through OR, this is the systematic version of "should we switch."
Public benches and Twitter recs measure a fixed task set. They tell you how a model behaved on someone else's prompts, someone else's tools, someone else's budget. With 500+ models on OpenRouter, "just switch to the new one" is expensive guesswork. Same leak as picking a CMS because a roundup said it was number one.
OpenRouter shipped Ori Eval for that. It is not a leaderboard. It is a way to score models on your app.
What it actually does:
- Scans the repo for every place a model gets called
- Asks what you care about: accuracy, latency, cost, tool-call correctness, whatever you name
- Writes a *.eval.ts file. You do not have to be an eval person
- Runs your agent against candidate models through OpenRouter, so the comparison can cross labs
- Pins the harness and the model for the run. If the score moves, it is the model, not the furniture shifting
- Checks whether tools were called or not, and grades open-ended answers with an LLM judge
- Can sit in CI so a worse model does not ship. Re-run when a new one drops
The sample they show is a table: catch rate, p50 latency, dollars per PR, pass/fail against your criteria. Then a recommendation, plus a cheaper "value" pick if volume grows.
I have not run this on my agents yet. Treating it as garage until someone in here drops a real table.
Cost is real. It calls live models. OpenRouter's own skill says 10 to 30 minutes and it can spend more than the credit on the key. First pass lands in a throwaway workspace. The eval is not yours until you decide to keep the file.
Try it
Tell your coding agent:
and follow the instructions in its output to get started
Already on the OpenRouter MCP? /spawn-ori-eval
By hand:
Then ori login. Evals need Bun.
If you already pay OpenRouter and you keep asking whether to switch models, this is that question with a file you can open. If you run it, post the table. I want the boring numbers, not the rec.
21
7 comments
David Vogel
7
Not sure which model to use?
Clief Notes
skool.com/cliefnotes
What we give away free beats most paid courses. Build durable AI systems with a Marine vet and Edinburgh researcher. 40+ lessons, growing.
Leaderboard (30-day)
Powered by