Pragmatikos
Benchmarks score models. You run setups.
A benchmark is a clean task, a hidden test and a single model. Real work is a messy repo, your tools, your steering — and increasingly one model planning while another builds. Pragmatikos scores that: every model and every planner → builder pairing, on real sessions, by what actually shipped.
Rankings below are pooled from real developer sessions.
claude-fable → muse-spark · 54 / 100
the rankings
Models and pairings, ranked by what ships
Did the work land in a commit, how many nudges did it take, how clean was it, what did it cost. Scored on real sessions, not lab tasks. Thin evidence gets pulled toward the middle so one lucky run can't top the chart.
join in
This is a first answer, not the final one. It's built from the sessions developers have shared so far. The more histories in the pool, the harder the ranking gets to argue with.
why it exists
Three questions. Three instruments.
The first two are useful and stay useful. Developers choosing a model and a workflow are missing the third answer.
question 01 · the lab
Benchmarks ask: how capable is this model?
SWE-bench, Terminal-bench, Aider, LMArena. Controlled tasks, hidden tests, preference votes — the best way to compare models on equal footing and know each model’s ceiling.
- signal
- pass rate · votes
- blind spot
- One model, a clean task, nobody steering.
question 02 · the crowd
Usage rankings ask: what is the market using?
opencode.ai/data, OpenRouter. Tokens, users, retention, dollars per session — the best way to see adoption and where the mix is shifting.
- signal
- tokens · users · $/session
- blind spot
- Popular is not productive — and still one model at a time.
question 03 · the recorder
Pragmatikos asks: what ships in real work?
Turns, phases, edits and errors from real sessions, correlated with what shipped. The one instrument that scores the setups developers actually run.
- signal
- session record × shipped outcomes
- blind spot
- Anything it never saw. Observational, not a benchmark.
Capable. Popular. Effective.
Only one of them has a row for the setup you actually run.
the blind spot both miss
Same builder. Different planner. Opposite results.
The build model is identical. One planner got it to a commit in a few turns; the other burned dozens and shipped nothing. A model ranking can't see this. A pairing ranking can.
claude-fable-5 → grok-4.6
0 turns
grok-4.6 → grok-4.6
0 turns
Real sessions. Unshipped work counts against the model.
how the score works · scroll to explode
Six tiers, ten axes, one weighted mean
Outcome
Ship rate and one-shot rate (prompt plus approval).
Cost of a ship
Turns, hours and dollars per shipped cycle.
Precision
Edit errors and human interventions.
Discipline
Share of cycles verified after the last edit.
Efficiency
Edits per human turn.
Latency
Median seconds per assistant step.
Pool-relative
Every axis is a log distance to the pooled average — log odds for rates — on a fixed ×4 span, so ×2 and ÷2 sit the same distance from the line.
Evidence-weighted
Small groups shrink toward the pool with k = 10 pseudo-cycles. A two-cycle fluke nearly vanishes.
Failure lowers the rate
Unshipped cycles stay in the ship-rate denominator, so few-turn dead ends can never look good. Turns, hours and dollars are priced per shipped cycle.
Pool-relative
Every axis is a log distance to the pooled average — log odds for rates — on a fixed ×4 span, so ×2 and ÷2 sit the same distance from the line.
Evidence-weighted
Small groups shrink toward the pool with k = 10 pseudo-cycles. A two-cycle fluke nearly vanishes.
Failure lowers the rate
Unshipped cycles stay in the ship-rate denominator, so few-turn dead ends can never look good. Turns, hours and dollars are priced per shipped cycle.
two records, correlated
The session record comes first
Every turn, plan and build phase, edit, tool error and dollar is read from the agent's own record — that is where every axis is measured. Commits only decide which cycles count as shipped: a window from the previous commit, file overlap decides credit, active-but-absent sessions advised only.
contribute
Make the ranking yours to trust
- 1
Run it locally
Clone ocInsights, make install-plugin, open the deck — your card is computed where the work happened.
- 2
Check your card
One number per model and per planner → builder pair.
- 3
Share the numbers
One history is anecdote; pooled histories are evidence. Run make contribute — twenty fields per cycle, nothing else.
Twenty fields per cycle — day, models, turns, edits, cost, shipping. No paths, prompts, session ids or projects. Preview with make contribute-preview before anything leaves your machine.
▍