Empirical cross-provider AI model benchmark. Task-level quality, latency, coverage, and cost-equivalent measurements for practical model routing.
Tasks
The 9 custom task types represented in the current scorecard
Tasks in detail
Click a task to see every model's response. Each task card shows the top 3 models by mean score. The full set of trials for each task is in /raw/.
code_debug- 1 grok-4.5-direct 5.0
- 2 grok-4.6-direct 5.0
- 3 kimi-k3 5.0
1081 trials
3.2 avg score
473 pass
json_strict- 1 minimax-m3 5.0
- 2 ollama-cl/minimax-m2.5 5.0
- 3 minimax-m2.7 5.0
853 trials
4.8 avg score
822 pass
code_gen_long- 1 zai-coding/glm-5.2 5.0
- 2 zai-coding/glm-4.7 5.0
- 3 ollama-cl/deepseek-v4-pro 5.0
1050 trials
4.3 avg score
872 pass
summarize- 1 ollama-cl/qwen3-coder:480b 5.0
- 2 mistral-medium-2508 5.0
- 3 zai-coding/glm-5.1 5.0
1009 trials
4.4 avg score
864 pass
creative_write- 1 ollama-cl/gemma4 3.0
- 2 zai-coding/glm-5-turbo 2.8
- 3 muse-glimmer-30b-q4_k_xl-local 2.6
899 trials
1.3 avg score
194 pass
reasoning_multistep- 1 minimax-m3 5.0
- 2 ollama-cl/minimax-m2.5 5.0
- 3 minimax-m2.7 5.0
800 trials
4.8 avg score
776 pass
agentic_prompt- 1 ollama-cl/minimax-m2.5 5.0
- 2 claude-opus-5 5.0
- 3 claude-opus-5-high 5.0
903 trials
4.2 avg score
712 pass
refactor_existing_code- 1 ollama-cl/minimax-m2.5 4.4
- 2 minimax-m2.7-highspeed 4.3
- 3 ollama-cl/nemotron-3-ultra 4.3
774 trials
3.3 avg score
446 pass
agentic_tool_use- 1 xai/grok-4-1-fast-reasoning 5.0
- 2 ollama-cl/qwen3-coder-next 5.0
- 3 ollama-cl/qwen3.5:397b 5.0
664 trials
4.3 avg score
598 pass