Empirical cross-provider AI model benchmark. Task-level quality, latency, coverage, and cost-equivalent measurements for practical model routing.

Results (custom suite)

Deprecated path · use /custom/ for the new canonical route

Full leaderboard

Every model that completed at least 4 trials across the 9 custom tasks, sorted by mean judge score. Pass % is the share of trials where the judge returned passes_key_criteria: true. Wall is end-to-end latency including the model's reasoning trace. 1P-EST is the apples-to-apples cost at each provider's public-API list price (see the pricing footnote on the home page). SUBSCRIPTION shows the actual wire cost: free / sub if the user is on a metered tier (skynet, ChatGPT Plus, Anthropic Max), else the real per-call cost. Use 1P-EST to compare across providers; use SUBSCRIPTION to see what you actually paid.

#MODELPROVIDERKINDSCOREPASS %MED WALL1P-ESTSUBSCRIPTIONTOKENSREASON
1claude-fableclaude-clilong-context4.44 / 588%18.1s$9.6823$0.7480 wire585,055 in / 12,087 out
2claude-opus-5-highclaude-clireasoning4.44 / 587%23.1s$5.1991 wire380 in / 265,080 out
3claude-opus-5claude-clireasoning4.39 / 589%15.9s$1.3818$2.1607 wire180 in / 55,236 out
4ollama-cl/gemma4skynetgeneral4.38 / 585%3.8sfree / sub20,170 in / 19,675 out
5grok-4.5-directxai-directreasoning4.33 / 581%12.2s$4.0035$2.3398 wire75,150 in / 41,675 out
6mistral/magistral-smallskynetreasoning4.32 / 582%1.7s$0.0220free / sub13,140 in / 10,262 out
7grok-4.6-directxai-directreasoning4.31 / 580%27.3s$0.1281$0.7684 wire24,365 in / 13,222 out
8ollama-cl/deepseek-v4-flashskynetreasoning4.29 / 582%6.5s$0.0092free / sub11,854 in / 43,164 out3 × thinking-only
9gemini-3.1-pro-highantigravity-clireasoning4.29 / 582%25.5s$2.5499free / sub413,967 in / 143,495 out
10muse-glimmer-30b-q4_k_xl-localllama.cpp-localreasoning4.22 / 580%34.2sfree / sub15,690 in / 107,826 out
11ollama-cl/qwen3-coder-nextskynetcoding4.19 / 579%4.4sfree / sub19,149 in / 18,869 out
12claude-sonnetclaude-cligeneral4.18 / 583%9.6s$2.9467$1.2044 wire882,467 in / 19,953 out
13codex-gpt-5.6-solcodex-clireasoning4.16 / 580%11.9s$7.2342free / sub1,120,466 in / 65,273 out
14mistral/mistral-small-latestskynetfast4.15 / 576%1.7sfree / sub13,140 in / 10,203 out
15gemini-3.6-flash-highantigravity-clifast4.13 / 582%25.7s$1.6449free / sub407,237 in / 137,870 out
16claude-opusclaude-clireasoning4.11 / 578%14.0s$13.3107$1.2231 wire784,297 in / 20,617 out
17ollama-cl/kimi-k2.7-codeskynetcoding4.06 / 579%11.6s$0.1197free / sub11,954 in / 51,139 out7 × thinking-only
18mistral/codestralskynetcoding4.03 / 574%3.9s$0.0574free / sub12,732 in / 14,879 out
19grok-4.20-non-reasoning-directxai-directgeneral4.02 / 574%2.1s$0.4730$0.1703 wire67,235 in / 56,424 out
20grok-4-fast-directxai-directfast4.00 / 576%6.5s$0.1064 wire23,100 in / 10,472 out
21zai-coding/glm-5.3zai-codingcoding4.00 / 576%26.3sfree / sub15,130 in / 91,340 out7 × thinking-only
22grok-4.20-reasoning-directxai-directreasoning3.97 / 574%20.8s$0.3654$1.6737 wire67,201 in / 38,501 out
23minimax-m3skynetreasoning3.94 / 570%16.8s$0.2393free / sub49,602 in / 187,041 out21 × thinking-only
24mistral-medium-2508skynetgeneral3.93 / 571%4.2sfree / sub15,760 in / 19,633 out
25ollama-cl/nemotron-3-superskynetgeneral3.92 / 569%14.5sfree / sub19,660 in / 79,490 out5 × thinking-only
26ollama-cl/nemotron-3-nano:30bskynetfast3.91 / 571%9.4sfree / sub13,140 in / 48,910 out2 × thinking-only
27ollama-local/qwen3.6:35b-mlxskynetgeneral3.89 / 578%48.8sfree / sub2,939 in / 41,646 out
28claude-haikuclaude-clifast3.86 / 574%40.4s$22.4783$27.3860 wire854,856 in / 4,324,693 out
29ollama-local/gemma4:26bskynetgeneral3.85 / 577%53.6sfree / sub5,197 in / 59,285 out4 × thinking-only
30zai-coding/glm-5.1skynetgeneral3.84 / 573%20.0s$0.2659free / sub15,130 in / 78,356 out8 × thinking-only
31grok-3-mini-directxai-directfast3.82 / 568%8.8s$0.0370$0.3583 wire68,395 in / 32,991 out
32grok-4.3-latest-directxai-directgeneral3.82 / 570%8.0s$0.6910$0.3514 wire68,395 in / 32,388 out
33codex-gpt-5.4codex-cligeneral3.79 / 566%12.4s$4.7321free / sub1,491,152 in / 100,426 out
34ollama-cl/gpt-oss:120bskynetgeneral3.78 / 567%10.5sfree / sub3,297 in / 6,897 out
35codex-gpt-5.3-sparkcodex-clifast3.77 / 564%7.7s$1.8118free / sub2,657,050 in / 845,573 out
36ollama-cl/minimax-m2.5skynetgeneral3.74 / 569%13.6sfree / sub22,710 in / 79,280 out2 × thinking-only
37mistral/devstralskynetcoding3.71 / 570%3.6sfree / sub24,542 in / 25,237 out
38codex-gpt-5.5codex-clireasoning3.71 / 561%13.2s$7.3264free / sub1,287,508 in / 44,442 out
39kimi-k3kimireasoning3.67 / 578%17.6sfree / sub6,028 in / 54,037 out2 × thinking-only
40codex-gpt-5.6-lunacodex-clifast3.63 / 562%20.3s$0.9295free / sub1,313,199 in / 136,435 out
41mistral/mistral-mediumskynetgeneral3.62 / 562%3.4s$0.3179free / sub25,418 in / 30,780 out
42zai-coding/glm-5-turboskynetfast3.60 / 560%60.5s$0.2823free / sub14,815 in / 83,577 out9 × thinking-only
43ollama-cl/gpt-oss:20bskynetfast3.56 / 563%25.2sfree / sub10,719 in / 42,780 out4 × thinking-only
44codex-gpt-5.6-terracodex-cligeneral3.54 / 563%10.7s$5.2195free / sub1,700,849 in / 80,617 out
45minimax-m2.7-highspeedskynetreasoning3.52 / 566%26.1s$0.0621free / sub27,170 in / 148,529 out20 × thinking-only
46xai/grok-4-1-fast-reasoningskynetreasoning3.51 / 554%6.2s$0.0374free / sub30,685 in / 62,513 out
47minimax-m2.7skynetreasoning3.48 / 564%26.0s$0.1176free / sub27,184 in / 140,248 out21 × thinking-only
48xai/grok-3-miniskynetreasoning3.46 / 559%6.2s$0.0407free / sub30,367 in / 63,245 out
49ollama-cl/nemotron-3-ultraskynetgeneral3.44 / 559%73.1sfree / sub24,011 in / 115,227 out18 × thinking-only
50xai/grok-4.3-latestskynetgeneral3.44 / 557%6.7s$1.0429free / sub30,685 in / 63,387 out
51zai-coding/glm-5skynetgeneral3.39 / 561%31.5s$0.3373free / sub16,771 in / 100,180 out15 × thinking-only
52codex-gpt-5.4-minicodex-clifast3.37 / 560%12.4s$0.4561free / sub1,239,449 in / 70,180 out
53zai-coding/glm-4.7skynetgeneral3.21 / 550%64.8s$0.5048free / sub16,792 in / 224,891 out
54kimi-for-codingkimicoding3.17 / 561%20.9sfree / sub6,028 in / 68,274 out7 × thinking-only
55zai-coding/glm-5.2skynetcoding3.00 / 550%25.2s$0.2921free / sub15,031 in / 86,587 out10 × thinking-only
56grok-4-1-fast-directxai-directreasoning2.98 / 554%5.8s$0.0266$0.3035 wire60,085 in / 29,087 out
57grok-build-0.1xai-directgeneral2.95 / 554%18.6s$0.6801$0.8359 wire37,600 in / 19,684 out
58ollama-cl/kimi-k2.6skynetgeneral2.85 / 548%27.3s$0.3018free / sub17,982 in / 132,264 out27 × thinking-only
59grok-4.3-directxai-directgeneral2.70 / 546%4.8s$0.3809$0.1875 wire38,080 in / 17,774 out
60ollama-cl/kimi-k2.5skynetgeneral2.67 / 542%19.8s$0.5073free / sub27,434 in / 223,120 out45 × thinking-only
61mistral/mistral-largeskynetgeneral2.27 / 540%5.5s$0.0305free / sub17,186 in / 14,604 out
62ollama-cl/qwen3.5:397bskynetgeneral2.04 / 536%48.7sfree / sub19,720 in / 147,794 out30 × thinking-only
63ollama-cl/mistral-large-3:675bskynetgeneral1.92 / 533%4.6sfree / sub9,456 in / 10,533 out
64ollama-local/qwen3.6:27bskynetgeneral0.73 / 59%20.8s$0.0063free / sub1,049 in / 10,111 out

Quality per dollar

Same models, sorted by score then 1P-EST cost (apples-to-apples public-API rate). The diagonal: top of the list is the bargain tier (cheap, high score). Bottom: you pay 10–50× for 0.3 score points. Three rules of thumb from the data:

  1. For most agent tasks (structured output, summarization, simple reasoning), the cheap tier matches or beats the expensive tier at a fraction of the cost.
  2. For long-context creative prose and dense code generation, the more expensive models do pull ahead — but usually by less than the cost ratio suggests.
  3. Reasoning model "reasoning-only" answers (model thought, didn't print content) are a hidden cost: the trial timed out, the user got no answer, the score is 0. Cheap tier has fewer of these because they have less reasoning budget to burn.
MODELSCORE1P-EST$ / 1.0 SCORE POINT (1P)
claude-fable4.44$9.6823$2.1807
claude-opus-5-high4.44
claude-opus-54.39$1.3818$0.3148
grok-4.5-direct4.33$4.0035$0.9246
grok-4.6-direct4.31$0.1281$0.0297
claude-sonnet4.18$2.9467$0.7050
claude-opus4.11$13.3107$3.2386
grok-4.20-non-reasoning-direct4.02$0.4730$0.1177
grok-4-fast-direct4
grok-4.20-reasoning-direct3.97$0.3654$0.0920
claude-haiku3.86$22.4783$5.8234
grok-4.3-latest-direct3.82$0.6910$0.1809
grok-3-mini-direct3.82$0.0370$0.0097
grok-4-1-fast-direct2.98$0.0266$0.0089
grok-build-0.12.95$0.6801$0.2305
grok-4.3-direct2.7$0.3809$0.1411

Cross-judge validation (Fable vs minimax-m3)

Custom-suite trials were originally scored by minimax-m3. To test judge bias, 1041 paired records were also scored by claude-fable-5 (a different model family). Across those pairs, Fable differs by +0.39 points on average; the direction and size vary by task:

TASKFABLE - M3 MEAN ΔPATTERN
agentic_prompt-0.28minimax-m3 more generous
agentic_tool_use-0.23Close agreement
code_debug-0.16Close agreement
code_gen_long+0.42Fable more generous
creative_write+1.93Fable substantially more generous
json_strict+0.03Close agreement
reasoning_multistep+0.09Close agreement
refactor_existing_code+0.90Fable more generous
summarize+0.22Close agreement

Agreement stats: 59% exact match, 82% within 1 point, 91% within 2 points, and 85% pass agreement. This is a sensitivity analysis, not a replacement judge: task-level deltas show where rank claims are more dependent on the rubric model.

Full per-cell cross-judge data: data/full_scorecard.json under cross_judge_fable.per_cell.

Per-task leaderboards

The leader changes by task. Click any task on the /tasks/ page to inspect responses from the 68 custom-suite model IDs.

agentic_prompt
  1. 1 ollama-cl/minimax-m2.5 5.0
  2. 2 claude-opus-5 5.0
  3. 3 claude-opus-5-high 5.0
agentic_tool_use
  1. 1 xai/grok-4-1-fast-reasoning 5.0
  2. 2 ollama-cl/qwen3-coder-next 5.0
  3. 3 ollama-cl/qwen3.5:397b 5.0
code_debug
  1. 1 grok-4.5-direct 5.0
  2. 2 grok-4.6-direct 5.0
  3. 3 kimi-k3 5.0
code_gen_long
  1. 1 zai-coding/glm-5.2 5.0
  2. 2 zai-coding/glm-4.7 5.0
  3. 3 ollama-cl/deepseek-v4-pro 5.0
creative_write
  1. 1 ollama-cl/gemma4 3.0
  2. 2 zai-coding/glm-5-turbo 2.8
  3. 3 muse-glimmer-30b-q4_k_xl-local 2.6
json_strict
  1. 1 minimax-m3 5.0
  2. 2 ollama-cl/minimax-m2.5 5.0
  3. 3 minimax-m2.7 5.0
reasoning_multistep
  1. 1 minimax-m3 5.0
  2. 2 ollama-cl/minimax-m2.5 5.0
  3. 3 minimax-m2.7 5.0
refactor_existing_code
  1. 1 ollama-cl/minimax-m2.5 4.4
  2. 2 minimax-m2.7-highspeed 4.3
  3. 3 ollama-cl/nemotron-3-ultra 4.3
summarize
  1. 1 ollama-cl/qwen3-coder:480b 5.0
  2. 2 mistral-medium-2508 5.0
  3. 3 zai-coding/glm-5.1 5.0

Blind head-to-head arena (top 5)

80 pairwise bouts across 8 harder-than-grid prompts. Each bout judged by an independent 3-judge panel — none of the contenders judges itself. A/B orientation is hash-randomized per bout to cancel position bias. Consensus = majority vote; split = tie.

Judges: minimax-m3 + claude-haiku + grok-4.3-direct 80 bouts
#MODELELOWLT
claude-fable1610.92075
claude-opus1516.617105
claude-sonnet1479.511174
codex-gpt-5.6-sol1414.56188
grok-4.5-direct1478.512146
How this differs from the grid leaderboard above

The grid rewards breadth (acing all 9 structured tasks). The arena rewards depth on harder multi-axis prompts (bug hunting, ADR writing, ruthless code review, literary fiction, agent planning, math chains, SQL fan-trap fixes). A model that aces structured tasks but loses on genuinely hard reasoning drops in the arena — this is why Grok 4.5 ranks #1 on the grid but #4 in the arena, and why Claude Fable wins the arena.

Use-case pairing

Which model to use for which kind of work, distilled from the data. This is the "what pairs with what" cheat sheet.

agentic-planning
  1. 1 ollama-cl/minimax-m2.5 — 5.0/5
  2. 2 claude-opus-5 — 5.0/5
  3. 3 claude-opus-5-high — 5.0/5
code-bug-finding
  1. 1 grok-4.5-direct — 5.0/5
  2. 2 grok-4.6-direct — 5.0/5
  3. 3 kimi-k3 — 5.0/5
code-generation
  1. 1 zai-coding/glm-5.2 — 5.0/5
  2. 2 zai-coding/glm-4.7 — 5.0/5
  3. 3 ollama-cl/deepseek-v4-pro — 5.0/5
code-refactoring
  1. 1 ollama-cl/minimax-m2.5 — 4.4/5
  2. 2 minimax-m2.7-highspeed — 4.3/5
  3. 3 ollama-cl/nemotron-3-ultra — 4.3/5
creative-prose
  1. 1 ollama-cl/gemma4 — 3.0/5
  2. 2 zai-coding/glm-5-turbo — 2.8/5
  3. 3 muse-glimmer-30b-q4_k_xl-local — 2.6/5
reasoning-math
  1. 1 minimax-m3 — 5.0/5
  2. 2 ollama-cl/minimax-m2.5 — 5.0/5
  3. 3 minimax-m2.7 — 5.0/5
structured-output
  1. 1 minimax-m3 — 5.0/5
  2. 2 ollama-cl/minimax-m2.5 — 5.0/5
  3. 3 minimax-m2.7 — 5.0/5
summarization
  1. 1 ollama-cl/qwen3-coder:480b — 5.0/5
  2. 2 mistral-medium-2508 — 5.0/5
  3. 3 zai-coding/glm-5.1 — 5.0/5
tool-use
  1. 1 xai/grok-4-1-fast-reasoning — 5.0/5
  2. 2 ollama-cl/qwen3-coder-next — 5.0/5
  3. 3 ollama-cl/qwen3.5:397b — 5.0/5