industry leaderboard · benchmark-specific graders

How every model stacks up against the canonical evals.

GPQA Diamond, MMLU-Pro, IFEval and HumanEval+ run through the same harness. Each family keeps its own prompt format and deterministic grader; coverage and N are shown instead of implying every model completed every suite.

Academic

2 benchmarks
ACADEMIC SATURATED
as of 2026-08-13T20:37:55+00:00

GPQA Diamond

Graduate-level Q&A (Bio/Chem/Physics)

67 models tested · 198 source items · 4,418 observed attempts

Saturation warning. Saturated above ~93% for frontier models; score deltas at top are within noise.

TOP MODELS
  1. 1 claude-fable 96.7%
  2. 2 claude-opus 93.3%
  3. 3 codex-gpt-5.6-sol 93.3%
VIEW DETAILS
  1. 1 claude-fable 96.7%
  2. 2 claude-opus 93.3%
  3. 3 codex-gpt-5.6-sol 93.3%
  4. 4 codex-gpt-5.6-terra 93.3%
  5. 5 gemini-3.1-pro-high 92.0%
  6. 6 grok-4.6-direct 92.0%
  7. 7 grok-4-fast-direct 92.0%
  8. 8 grok-4.3-latest-direct 90.9%
  9. 9 gemini-3.6-flash-high 90.0%
  10. 10 claude-sonnet 90.0%
  11. 11 grok-4-1-fast-direct 90.0%
  12. 12 grok-4.20-reasoning-direct 89.1%
  13. 13 claude-opus-5-high 88.0%
  14. 14 grok-4.5-direct 87.3%
  15. 15 codex-gpt-5.5 86.7%
  16. 16 claude-opus-5 84.8%
  17. 17 grok-4.3-direct 83.3%
  18. 18 grok-3-mini-direct 83.3%
  19. 19 codex-gpt-5.6-luna 83.3%
  20. 20 grok-build-0.1 80.0%
  21. 21 codex-gpt-5.4 73.3%
  22. 22 ollama-cl/gemma4 73.3%
  23. 23 grok-4.20-non-reasoning-direct 70.0%
  24. 24 codex-gpt-5.4-mini 66.7%
UNAVAILABLE · NO USABLE RESPONSE
  • claude-cli/claude-fable 50/50 attempts
  • claude-cli/claude-fable-5 50/50 attempts
  • claude-cli/claude-opus-4-8 50/50 attempts
  • claude-cli/claude-sonnet-4-6 50/50 attempts
  • kimi-k3 10/10 attempts
  • minimax-medium 30/30 attempts
  • oai/grok-4-1-fast-reasoning 30/30 attempts
  • ollama-cl/minimax-m2.5 310/310 attempts
  • xai/grok-3-mini 30/30 attempts
  • xai/grok-4-1-fast-reasoning 30/30 attempts
  • xai/grok-4.3-latest 30/30 attempts
  • xai/grok-4.5 30/30 attempts
ACADEMIC SATURATED
as of 2026-07-30T04:26:47+00:00

MMLU-Pro

Professional-level MCQ across 14 disciplines (math, physics, chemistry, biology, law, philosophy, etc.)

32 models tested · 12,032 source items · 4,546 observed attempts

Saturation warning. Saturated above ~80% for frontier models on most disciplines; per-discipline ranking is the real signal.

TOP MODELS
  1. 1 codex-gpt-5.6-sol 92.0%
  2. 2 claude-opus-5 88.4%
  3. 3 ollama-cl/kimi-k2.5 63.0%
VIEW DETAILS
  1. 1 codex-gpt-5.6-sol 92.0%
  2. 2 claude-opus-5 88.4%
  3. 3 ollama-cl/kimi-k2.5 63.0%
  4. 4 zai-coding/glm-4.6 58.3%
  5. 5 minimax-m3 57.0%
  6. 6 zai-coding/glm-5-turbo 50.0%
  7. 7 ollama-cl/kimi-k2.7-code 46.2%
  8. 8 minimax-m2.7-highspeed 44.0%
  9. 9 zai-coding/glm-5.1 41.7%
  10. 10 minimax-m2.7 38.0%
  11. 11 mistral/mistral-medium 36.5%
  12. 12 ollama-cl/gpt-oss:120b 35.0%
  13. 13 ollama-cl/nemotron-3-ultra 34.4%
  14. 14 ollama-cl/deepseek-v4-pro 29.3%
  15. 15 gemini-3.6-flash-high 27.0%
  16. 16 gemini-3.1-pro-high 26.0%
  17. 17 zai-coding/glm-5.2 19.2%
  18. 18 zai-coding/glm-4.7 18.0%
  19. 19 mistral/devstral 18.0%
  20. 20 zai-coding/glm-5 16.2%
  21. 21 mistral/mistral-large 12.0%
UNAVAILABLE · NO USABLE RESPONSE
  • claude-cli/claude-fable 140/140 attempts
  • claude-cli/claude-fable-4 80/80 attempts
  • claude-cli/claude-fable-5 140/140 attempts
  • claude-cli/claude-fable-5-t 80/80 attempts
  • claude-cli/claude-haiku 140/140 attempts
  • claude-cli/claude-opus 80/80 attempts
  • claude-cli/claude-opus-4-8 140/140 attempts
  • claude-cli/claude-sonnet 140/140 attempts
  • grok-4.5-direct 60/60 attempts
  • kimi-k3 10/10 attempts
  • ollama-cl/minimax-m2.5 200/200 attempts

Instruction Following

1 benchmark
INSTR-FOLLOWING SNAPSHOT
as of 2026-08-13T20:19:42+00:00

IFEval

Verifiable instruction-following constraints

61 models tested · 541 source items · 4,611 observed attempts
VIEW DETAILS
  1. 1 grok-build-0.1 92.2%
  2. 2 grok-4.20-reasoning-direct 91.5%
  3. 3 grok-4.20-non-reasoning-direct 91.1%
  4. 4 grok-4.6-direct 90.7%
  5. 5 grok-4.3-latest-direct 89.7%
  6. 6 grok-4.3-direct 88.9%
  7. 7 grok-3-mini-direct 88.9%
  8. 8 grok-4-fast-direct 86.7%
  9. 9 grok-4.5-direct 86.1%
  10. 10 grok-4-1-fast-direct 85.6%
  11. 11 claude-haiku 78.3%
  12. 12 claude-sonnet 77.8%
  13. 13 ollama-cl/nemotron-3-super 77.8%
  14. 14 claude-fable 77.4%
  15. 15 codex-gpt-5.6-sol 77.3%
  16. 16 zai-coding/glm-4.6 77.2%
  17. 17 ollama-cl/qwen3-coder:480b 77.2%
  18. 18 ollama-cl/qwen3.5:397b 75.6%
  19. 19 gemini-3.1-pro-high 75.0%
  20. 20 ollama-cl/kimi-k2.7-code 74.4%
  21. 21 mistral/mistral-medium 73.1%
  22. 22 mistral-medium-2508 72.8%
  23. 23 claude-opus 72.8%
  24. 24 ollama-cl/deepseek-v4-flash 72.8%
UNAVAILABLE · NO USABLE RESPONSE
  • claude-cli/claude-fable 110/110 attempts
  • claude-cli/claude-fable-5 110/110 attempts
  • claude-cli/claude-haiku 60/60 attempts
  • claude-cli/claude-opus 60/60 attempts
  • claude-cli/claude-opus-4-8 50/50 attempts
  • claude-cli/claude-sonnet 60/60 attempts
  • claude-cli/claude-sonnet-4-6 50/50 attempts
  • kimi-k3 10/10 attempts
  • minimax-medium 30/30 attempts
  • ollama-cl/gemma4:26b 60/60 attempts
  • ollama-cl/gpt-oss-120b 60/60 attempts
  • ollama-cl/minimax-m2.5 190/190 attempts
  • ollama-cl/mistral-large-3-675b 60/60 attempts
  • ollama-cl/nemotron-3-nano-30b 60/60 attempts
  • ollama-cl/qwen3.5 50/50 attempts
  • ollama-cl/qwen3.5-code 60/60 attempts
  • ollama-cl/qwen3.5:27b 60/60 attempts

Coding

1 benchmark
CODING SNAPSHOT
as of 2026-07-30T14:20:48+00:00

HumanEval+

Python function-completion problems with hidden tests (evalplus hard-set)

27 models tested · 164 source items · 1,958 observed attempts
TOP MODELS
  1. 1 mistral/mistral-medium 97.5%
  2. 2 codex-gpt-5.6-luna 96.0%
  3. 3 mistral/devstral 95.0%
VIEW DETAILS
  1. 1 mistral/mistral-medium 97.5%
  2. 2 codex-gpt-5.6-luna 96.0%
  3. 3 mistral/devstral 95.0%
  4. 4 gemini-3.6-flash-high 93.9%
  5. 5 codex-gpt-5.6-sol 92.0%
  6. 6 gemini-3.1-pro-high 91.5%
  7. 7 mistral/mistral-large 90.0%
  8. 8 codex-gpt-5.3-spark 88.0%
  9. 9 ollama-cl/nemotron-3-super 84.0%
  10. 10 claude-opus-5-high 82.5%
  11. 11 minimax-m2.7 82.5%
  12. 12 minimax-m2.7-highspeed 80.0%
  13. 13 codex-gpt-5.6-terra 80.0%
  14. 14 claude-opus-5 78.0%
  15. 15 minimax-m3 73.3%
  16. 16 zai-coding/glm-5.2 68.0%
  17. 17 ollama-cl/deepseek-v4-pro 55.0%
  18. 18 claude-haiku 48.0%
  19. 19 claude-fable 47.0%
  20. 20 claude-sonnet 44.0%
  21. 21 claude-opus 43.0%
  22. 22 ollama-cl/kimi-k2.5 36.1%
  23. 23 zai-coding/glm-5 32.5%
  24. 24 zai-coding/glm-4.7 25.0%
UNAVAILABLE · NO USABLE RESPONSE
  • grok-4.5-direct 80/80 attempts
  • kimi-k3 10/10 attempts
  • ollama-cl/minimax-m2.5 40/40 attempts

Fleet scatter

accuracy vs wire cost across 93 plotted model results
GPQA Diamond IFEval MMLU-Pro HumanEval+ x = total observed wire USD · y = accuracy % · bubble = attempts

Caveats

01

Saturation

GPQA Diamond: Saturated above ~93% for frontier models; score deltas at top are within noise.
MMLU-Pro: Saturated above ~80% for frontier models on most disciplines; per-discipline ranking is the real signal.

02

Contamination

These are public test sets. Training-data exposure cannot be ruled out from score alone, and this harness does not claim a retrospective decontamination pass.

03

Prompt specificity

Prompt shape differs by family: GPQA uses 5-shot CoT; MMLU-Pro uses question + options; IFEval uses the verbatim constraint prompt; HumanEval+ uses signature + docstring.

04

Coverage

Family coverage ranges from 27 to 67 model IDs in this snapshot. Read N before comparing ranks across cards.