Empirical cross-provider AI model benchmark. Task-level quality, latency, coverage, and cost-equivalent measurements for practical model routing.

Tasks

The 9 custom task types represented in the current scorecard

Tasks in detail

Click a task to see every model's response. Each task card shows the top 3 models by mean score. The full set of trials for each task is in /raw/.

code_debug
  1. 1 grok-4.5-direct 5.0
  2. 2 grok-4.6-direct 5.0
  3. 3 kimi-k3 5.0
1081 trials 3.2 avg score 473 pass
json_strict
  1. 1 minimax-m3 5.0
  2. 2 ollama-cl/minimax-m2.5 5.0
  3. 3 minimax-m2.7 5.0
853 trials 4.8 avg score 822 pass
code_gen_long
  1. 1 zai-coding/glm-5.2 5.0
  2. 2 zai-coding/glm-4.7 5.0
  3. 3 ollama-cl/deepseek-v4-pro 5.0
1050 trials 4.3 avg score 872 pass
summarize
  1. 1 ollama-cl/qwen3-coder:480b 5.0
  2. 2 mistral-medium-2508 5.0
  3. 3 zai-coding/glm-5.1 5.0
1009 trials 4.4 avg score 864 pass
creative_write
  1. 1 ollama-cl/gemma4 3.0
  2. 2 zai-coding/glm-5-turbo 2.8
  3. 3 muse-glimmer-30b-q4_k_xl-local 2.6
899 trials 1.3 avg score 194 pass
reasoning_multistep
  1. 1 minimax-m3 5.0
  2. 2 ollama-cl/minimax-m2.5 5.0
  3. 3 minimax-m2.7 5.0
800 trials 4.8 avg score 776 pass
agentic_prompt
  1. 1 ollama-cl/minimax-m2.5 5.0
  2. 2 claude-opus-5 5.0
  3. 3 claude-opus-5-high 5.0
903 trials 4.2 avg score 712 pass
refactor_existing_code
  1. 1 ollama-cl/minimax-m2.5 4.4
  2. 2 minimax-m2.7-highspeed 4.3
  3. 3 ollama-cl/nemotron-3-ultra 4.3
774 trials 3.3 avg score 446 pass
agentic_tool_use
  1. 1 xai/grok-4-1-fast-reasoning 5.0
  2. 2 ollama-cl/qwen3-coder-next 5.0
  3. 3 ollama-cl/qwen3.5:397b 5.0
664 trials 4.3 avg score 598 pass