Custom Suite
9 hand-written tasks · cross-judged 0–5 · the original bench
Full leaderboard
Every model that completed at least 4 trials across the 9 custom tasks, sorted by mean judge score. Pass % is the share of trials where the judge returned passes_key_criteria: true. Wall is end-to-end latency including the model's reasoning trace. 1P-EST is the apples-to-apples cost at each provider's public-API list price (see the pricing footnote on the home page). SUBSCRIPTION shows the actual wire cost: free / sub if the user is on a metered tier (skynet, ChatGPT Plus, Anthropic Max), else the real per-call cost. Use 1P-EST to compare across providers; use SUBSCRIPTION to see what you actually paid.
| # | MODEL | PROVIDER | KIND | SCORE | PASS % | MED WALL | 1P-EST | SUBSCRIPTION | TOKENS | REASON |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-fable | claude-cli | long-context | 4.44 / 5 | 88% | 18.1s | $9.6823 | $0.7480 wire | 585,055 in / 12,087 out | — |
| 2 | claude-opus-5-high | claude-cli | reasoning | 4.44 / 5 | 87% | 23.1s | — | $5.1991 wire | 380 in / 265,080 out | — |
| 3 | claude-opus-5 | claude-cli | reasoning | 4.39 / 5 | 89% | 15.9s | $1.3818 | $2.1607 wire | 180 in / 55,236 out | — |
| 4 | ollama-cl/gemma4 | skynet | general | 4.38 / 5 | 85% | 3.8s | — | free / sub | 20,170 in / 19,675 out | — |
| 5 | grok-4.5-direct | xai-direct | reasoning | 4.33 / 5 | 81% | 12.2s | $4.0035 | $2.3398 wire | 75,150 in / 41,675 out | — |
| 6 | mistral/magistral-small | skynet | reasoning | 4.32 / 5 | 82% | 1.7s | $0.0220 | free / sub | 13,140 in / 10,262 out | — |
| 7 | grok-4.6-direct | xai-direct | reasoning | 4.31 / 5 | 80% | 27.3s | $0.1281 | $0.7684 wire | 24,365 in / 13,222 out | — |
| 8 | ollama-cl/deepseek-v4-flash | skynet | reasoning | 4.29 / 5 | 82% | 6.5s | $0.0092 | free / sub | 11,854 in / 43,164 out | 3 × thinking-only |
| 9 | gemini-3.1-pro-high | antigravity-cli | reasoning | 4.29 / 5 | 82% | 25.5s | $2.5499 | free / sub | 413,967 in / 143,495 out | — |
| 10 | muse-glimmer-30b-q4_k_xl-local | llama.cpp-local | reasoning | 4.22 / 5 | 80% | 34.2s | — | free / sub | 15,690 in / 107,826 out | — |
| 11 | ollama-cl/qwen3-coder-next | skynet | coding | 4.19 / 5 | 79% | 4.4s | — | free / sub | 19,149 in / 18,869 out | — |
| 12 | claude-sonnet | claude-cli | general | 4.18 / 5 | 83% | 9.6s | $2.9467 | $1.2044 wire | 882,467 in / 19,953 out | — |
| 13 | codex-gpt-5.6-sol | codex-cli | reasoning | 4.16 / 5 | 80% | 11.9s | $7.2342 | free / sub | 1,120,466 in / 65,273 out | — |
| 14 | mistral/mistral-small-latest | skynet | fast | 4.15 / 5 | 76% | 1.7s | — | free / sub | 13,140 in / 10,203 out | — |
| 15 | gemini-3.6-flash-high | antigravity-cli | fast | 4.13 / 5 | 82% | 25.7s | $1.6449 | free / sub | 407,237 in / 137,870 out | — |
| 16 | claude-opus | claude-cli | reasoning | 4.11 / 5 | 78% | 14.0s | $13.3107 | $1.2231 wire | 784,297 in / 20,617 out | — |
| 17 | ollama-cl/kimi-k2.7-code | skynet | coding | 4.06 / 5 | 79% | 11.6s | $0.1197 | free / sub | 11,954 in / 51,139 out | 7 × thinking-only |
| 18 | mistral/codestral | skynet | coding | 4.03 / 5 | 74% | 3.9s | $0.0574 | free / sub | 12,732 in / 14,879 out | — |
| 19 | grok-4.20-non-reasoning-direct | xai-direct | general | 4.02 / 5 | 74% | 2.1s | $0.4730 | $0.1703 wire | 67,235 in / 56,424 out | — |
| 20 | grok-4-fast-direct | xai-direct | fast | 4.00 / 5 | 76% | 6.5s | — | $0.1064 wire | 23,100 in / 10,472 out | — |
| 21 | zai-coding/glm-5.3 | zai-coding | coding | 4.00 / 5 | 76% | 26.3s | — | free / sub | 15,130 in / 91,340 out | 7 × thinking-only |
| 22 | grok-4.20-reasoning-direct | xai-direct | reasoning | 3.97 / 5 | 74% | 20.8s | $0.3654 | $1.6737 wire | 67,201 in / 38,501 out | — |
| 23 | minimax-m3 | skynet | reasoning | 3.94 / 5 | 70% | 16.8s | $0.2393 | free / sub | 49,602 in / 187,041 out | 21 × thinking-only |
| 24 | mistral-medium-2508 | skynet | general | 3.93 / 5 | 71% | 4.2s | — | free / sub | 15,760 in / 19,633 out | — |
| 25 | ollama-cl/nemotron-3-super | skynet | general | 3.92 / 5 | 69% | 14.5s | — | free / sub | 19,660 in / 79,490 out | 5 × thinking-only |
| 26 | ollama-cl/nemotron-3-nano:30b | skynet | fast | 3.91 / 5 | 71% | 9.4s | — | free / sub | 13,140 in / 48,910 out | 2 × thinking-only |
| 27 | ollama-local/qwen3.6:35b-mlx | skynet | general | 3.89 / 5 | 78% | 48.8s | — | free / sub | 2,939 in / 41,646 out | — |
| 28 | claude-haiku | claude-cli | fast | 3.86 / 5 | 74% | 40.4s | $22.4783 | $27.3860 wire | 854,856 in / 4,324,693 out | — |
| 29 | ollama-local/gemma4:26b | skynet | general | 3.85 / 5 | 77% | 53.6s | — | free / sub | 5,197 in / 59,285 out | 4 × thinking-only |
| 30 | zai-coding/glm-5.1 | skynet | general | 3.84 / 5 | 73% | 20.0s | $0.2659 | free / sub | 15,130 in / 78,356 out | 8 × thinking-only |
| 31 | grok-3-mini-direct | xai-direct | fast | 3.82 / 5 | 68% | 8.8s | $0.0370 | $0.3583 wire | 68,395 in / 32,991 out | — |
| 32 | grok-4.3-latest-direct | xai-direct | general | 3.82 / 5 | 70% | 8.0s | $0.6910 | $0.3514 wire | 68,395 in / 32,388 out | — |
| 33 | codex-gpt-5.4 | codex-cli | general | 3.79 / 5 | 66% | 12.4s | $4.7321 | free / sub | 1,491,152 in / 100,426 out | — |
| 34 | ollama-cl/gpt-oss:120b | skynet | general | 3.78 / 5 | 67% | 10.5s | — | free / sub | 3,297 in / 6,897 out | — |
| 35 | codex-gpt-5.3-spark | codex-cli | fast | 3.77 / 5 | 64% | 7.7s | $1.8118 | free / sub | 2,657,050 in / 845,573 out | — |
| 36 | ollama-cl/minimax-m2.5 | skynet | general | 3.74 / 5 | 69% | 13.6s | — | free / sub | 22,710 in / 79,280 out | 2 × thinking-only |
| 37 | mistral/devstral | skynet | coding | 3.71 / 5 | 70% | 3.6s | — | free / sub | 24,542 in / 25,237 out | — |
| 38 | codex-gpt-5.5 | codex-cli | reasoning | 3.71 / 5 | 61% | 13.2s | $7.3264 | free / sub | 1,287,508 in / 44,442 out | — |
| 39 | kimi-k3 | kimi | reasoning | 3.67 / 5 | 78% | 17.6s | — | free / sub | 6,028 in / 54,037 out | 2 × thinking-only |
| 40 | codex-gpt-5.6-luna | codex-cli | fast | 3.63 / 5 | 62% | 20.3s | $0.9295 | free / sub | 1,313,199 in / 136,435 out | — |
| 41 | mistral/mistral-medium | skynet | general | 3.62 / 5 | 62% | 3.4s | $0.3179 | free / sub | 25,418 in / 30,780 out | — |
| 42 | zai-coding/glm-5-turbo | skynet | fast | 3.60 / 5 | 60% | 60.5s | $0.2823 | free / sub | 14,815 in / 83,577 out | 9 × thinking-only |
| 43 | ollama-cl/gpt-oss:20b | skynet | fast | 3.56 / 5 | 63% | 25.2s | — | free / sub | 10,719 in / 42,780 out | 4 × thinking-only |
| 44 | codex-gpt-5.6-terra | codex-cli | general | 3.54 / 5 | 63% | 10.7s | $5.2195 | free / sub | 1,700,849 in / 80,617 out | — |
| 45 | minimax-m2.7-highspeed | skynet | reasoning | 3.52 / 5 | 66% | 26.1s | $0.0621 | free / sub | 27,170 in / 148,529 out | 20 × thinking-only |
| 46 | xai/grok-4-1-fast-reasoning | skynet | reasoning | 3.51 / 5 | 54% | 6.2s | $0.0374 | free / sub | 30,685 in / 62,513 out | — |
| 47 | minimax-m2.7 | skynet | reasoning | 3.48 / 5 | 64% | 26.0s | $0.1176 | free / sub | 27,184 in / 140,248 out | 21 × thinking-only |
| 48 | xai/grok-3-mini | skynet | reasoning | 3.46 / 5 | 59% | 6.2s | $0.0407 | free / sub | 30,367 in / 63,245 out | — |
| 49 | ollama-cl/nemotron-3-ultra | skynet | general | 3.44 / 5 | 59% | 73.1s | — | free / sub | 24,011 in / 115,227 out | 18 × thinking-only |
| 50 | xai/grok-4.3-latest | skynet | general | 3.44 / 5 | 57% | 6.7s | $1.0429 | free / sub | 30,685 in / 63,387 out | — |
| 51 | zai-coding/glm-5 | skynet | general | 3.39 / 5 | 61% | 31.5s | $0.3373 | free / sub | 16,771 in / 100,180 out | 15 × thinking-only |
| 52 | codex-gpt-5.4-mini | codex-cli | fast | 3.37 / 5 | 60% | 12.4s | $0.4561 | free / sub | 1,239,449 in / 70,180 out | — |
| 53 | zai-coding/glm-4.7 | skynet | general | 3.21 / 5 | 50% | 64.8s | $0.5048 | free / sub | 16,792 in / 224,891 out | — |
| 54 | kimi-for-coding | kimi | coding | 3.17 / 5 | 61% | 20.9s | — | free / sub | 6,028 in / 68,274 out | 7 × thinking-only |
| 55 | zai-coding/glm-5.2 | skynet | coding | 3.00 / 5 | 50% | 25.2s | $0.2921 | free / sub | 15,031 in / 86,587 out | 10 × thinking-only |
| 56 | grok-4-1-fast-direct | xai-direct | reasoning | 2.98 / 5 | 54% | 5.8s | $0.0266 | $0.3035 wire | 60,085 in / 29,087 out | — |
| 57 | grok-build-0.1 | xai-direct | general | 2.95 / 5 | 54% | 18.6s | $0.6801 | $0.8359 wire | 37,600 in / 19,684 out | — |
| 58 | ollama-cl/kimi-k2.6 | skynet | general | 2.85 / 5 | 48% | 27.3s | $0.3018 | free / sub | 17,982 in / 132,264 out | 27 × thinking-only |
| 59 | grok-4.3-direct | xai-direct | general | 2.70 / 5 | 46% | 4.8s | $0.3809 | $0.1875 wire | 38,080 in / 17,774 out | — |
| 60 | ollama-cl/kimi-k2.5 | skynet | general | 2.67 / 5 | 42% | 19.8s | $0.5073 | free / sub | 27,434 in / 223,120 out | 45 × thinking-only |
| 61 | mistral/mistral-large | skynet | general | 2.27 / 5 | 40% | 5.5s | $0.0305 | free / sub | 17,186 in / 14,604 out | — |
| 62 | ollama-cl/qwen3.5:397b | skynet | general | 2.04 / 5 | 36% | 48.7s | — | free / sub | 19,720 in / 147,794 out | 30 × thinking-only |
| 63 | ollama-cl/mistral-large-3:675b | skynet | general | 1.92 / 5 | 33% | 4.6s | — | free / sub | 9,456 in / 10,533 out | — |
| 64 | ollama-local/qwen3.6:27b | skynet | general | 0.73 / 5 | 9% | 20.8s | $0.0063 | free / sub | 1,049 in / 10,111 out | — |
Quality per dollar
Same models, sorted by score then 1P-EST cost (apples-to-apples public-API rate). The diagonal: top of the list is the bargain tier (cheap, high score). Bottom: you pay 10–50× for 0.3 score points. Three rules of thumb from the data:
- For most agent tasks (structured output, summarization, simple reasoning), the cheap tier matches or beats the expensive tier at a fraction of the cost.
- For long-context creative prose and dense code generation, the more expensive models do pull ahead — but usually by less than the cost ratio suggests.
- Reasoning model "reasoning-only" answers (model thought, didn't print content) are a hidden cost: the trial timed out, the user got no answer, the score is 0. Cheap tier has fewer of these because they have less reasoning budget to burn.
| MODEL | SCORE | 1P-EST | $ / 1.0 SCORE POINT (1P) |
|---|---|---|---|
| claude-fable | 4.44 | $9.6823 | $2.1807 |
| claude-opus-5-high | 4.44 | — | — |
| claude-opus-5 | 4.39 | $1.3818 | $0.3148 |
| grok-4.5-direct | 4.33 | $4.0035 | $0.9246 |
| grok-4.6-direct | 4.31 | $0.1281 | $0.0297 |
| claude-sonnet | 4.18 | $2.9467 | $0.7050 |
| claude-opus | 4.11 | $13.3107 | $3.2386 |
| grok-4.20-non-reasoning-direct | 4.02 | $0.4730 | $0.1177 |
| grok-4-fast-direct | 4 | — | — |
| grok-4.20-reasoning-direct | 3.97 | $0.3654 | $0.0920 |
| claude-haiku | 3.86 | $22.4783 | $5.8234 |
| grok-4.3-latest-direct | 3.82 | $0.6910 | $0.1809 |
| grok-3-mini-direct | 3.82 | $0.0370 | $0.0097 |
| grok-4-1-fast-direct | 2.98 | $0.0266 | $0.0089 |
| grok-build-0.1 | 2.95 | $0.6801 | $0.2305 |
| grok-4.3-direct | 2.7 | $0.3809 | $0.1411 |
Cross-judge validation (Fable vs minimax-m3)
Custom-suite trials were originally scored by minimax-m3. To test judge bias, 1041 paired records were also scored by claude-fable-5 (a different model family). Across those pairs, Fable differs by +0.39 points on average; the direction and size vary by task:
| TASK | FABLE - M3 MEAN Δ | PATTERN |
|---|---|---|
| agentic_prompt | -0.28 | minimax-m3 more generous |
| agentic_tool_use | -0.23 | Close agreement |
| code_debug | -0.16 | Close agreement |
| code_gen_long | +0.42 | Fable more generous |
| creative_write | +1.93 | Fable substantially more generous |
| json_strict | +0.03 | Close agreement |
| reasoning_multistep | +0.09 | Close agreement |
| refactor_existing_code | +0.90 | Fable more generous |
| summarize | +0.22 | Close agreement |
Agreement stats: 59% exact match, 82% within 1 point, 91% within 2 points, and 85% pass agreement. This is a sensitivity analysis, not a replacement judge: task-level deltas show where rank claims are more dependent on the rubric model.
Full per-cell cross-judge data: data/full_scorecard.json under cross_judge_fable.per_cell.
Per-task leaderboards
The leader changes by task. Click any task on the /tasks/ page to inspect responses from the 68 custom-suite model IDs.
- 1 ollama-cl/minimax-m2.5 5.0
- 2 claude-opus-5 5.0
- 3 claude-opus-5-high 5.0
- 1 xai/grok-4-1-fast-reasoning 5.0
- 2 ollama-cl/qwen3-coder-next 5.0
- 3 ollama-cl/qwen3.5:397b 5.0
- 1 grok-4.5-direct 5.0
- 2 grok-4.6-direct 5.0
- 3 kimi-k3 5.0
- 1 zai-coding/glm-5.2 5.0
- 2 zai-coding/glm-4.7 5.0
- 3 ollama-cl/deepseek-v4-pro 5.0
- 1 ollama-cl/gemma4 3.0
- 2 zai-coding/glm-5-turbo 2.8
- 3 muse-glimmer-30b-q4_k_xl-local 2.6
- 1 minimax-m3 5.0
- 2 ollama-cl/minimax-m2.5 5.0
- 3 minimax-m2.7 5.0
- 1 minimax-m3 5.0
- 2 ollama-cl/minimax-m2.5 5.0
- 3 minimax-m2.7 5.0
- 1 ollama-cl/minimax-m2.5 4.4
- 2 minimax-m2.7-highspeed 4.3
- 3 ollama-cl/nemotron-3-ultra 4.3
- 1 ollama-cl/qwen3-coder:480b 5.0
- 2 mistral-medium-2508 5.0
- 3 zai-coding/glm-5.1 5.0
Blind head-to-head arena (top 5)
80 pairwise bouts across 8 harder-than-grid prompts. Each bout judged by an independent 3-judge panel — none of the contenders judges itself. A/B orientation is hash-randomized per bout to cancel position bias. Consensus = majority vote; split = tie.
| # | MODEL | ELO | W | L | T |
|---|---|---|---|---|---|
| claude-fable | 1610.9 | 20 | 7 | 5 | |
| claude-opus | 1516.6 | 17 | 10 | 5 | |
| claude-sonnet | 1479.5 | 11 | 17 | 4 | |
| codex-gpt-5.6-sol | 1414.5 | 6 | 18 | 8 | |
| grok-4.5-direct | 1478.5 | 12 | 14 | 6 |
How this differs from the grid leaderboard above
The grid rewards breadth (acing all 9 structured tasks). The arena rewards depth on harder multi-axis prompts (bug hunting, ADR writing, ruthless code review, literary fiction, agent planning, math chains, SQL fan-trap fixes). A model that aces structured tasks but loses on genuinely hard reasoning drops in the arena — this is why Grok 4.5 ranks #1 on the grid but #4 in the arena, and why Claude Fable wins the arena.
Use-case pairing
Which model to use for which kind of work, distilled from the data. This is the "what pairs with what" cheat sheet.
- 1 ollama-cl/minimax-m2.5 — 5.0/5
- 2 claude-opus-5 — 5.0/5
- 3 claude-opus-5-high — 5.0/5
- 1 grok-4.5-direct — 5.0/5
- 2 grok-4.6-direct — 5.0/5
- 3 kimi-k3 — 5.0/5
- 1 zai-coding/glm-5.2 — 5.0/5
- 2 zai-coding/glm-4.7 — 5.0/5
- 3 ollama-cl/deepseek-v4-pro — 5.0/5
- 1 ollama-cl/minimax-m2.5 — 4.4/5
- 2 minimax-m2.7-highspeed — 4.3/5
- 3 ollama-cl/nemotron-3-ultra — 4.3/5
- 1 ollama-cl/gemma4 — 3.0/5
- 2 zai-coding/glm-5-turbo — 2.8/5
- 3 muse-glimmer-30b-q4_k_xl-local — 2.6/5
- 1 minimax-m3 — 5.0/5
- 2 ollama-cl/minimax-m2.5 — 5.0/5
- 3 minimax-m2.7 — 5.0/5
- 1 minimax-m3 — 5.0/5
- 2 ollama-cl/minimax-m2.5 — 5.0/5
- 3 minimax-m2.7 — 5.0/5
- 1 ollama-cl/qwen3-coder:480b — 5.0/5
- 2 mistral-medium-2508 — 5.0/5
- 3 zai-coding/glm-5.1 — 5.0/5
- 1 xai/grok-4-1-fast-reasoning — 5.0/5
- 2 ollama-cl/qwen3-coder-next — 5.0/5
- 3 ollama-cl/qwen3.5:397b — 5.0/5