Herman · Model Bench
DATA SNAPSHOT · 23,566 published benchmark trials · 81 industry-tested model IDs · 4 industry families
updated 2026-08-15
empirical evaluation · task-aware routing

Testing AI on real work, not one global rank.

9 custom tasks sit beside GPQA Diamond, MMLU-Pro, IFEval and HumanEval+. Industry suites use benchmark-specific deterministic graders; custom tasks use rubric scoring with a partial second-judge validation layer.

Corpus accounting: 23,680 source records = 23,566 published benchmark trials + 114 quarantined harness-control records. See methodology.

how to read this board

Five columns of signal, no magic number.

FULL METHODOLOGY →
ACC

Accuracy

Percent passed by that benchmark's reference verifier. The homepage mean is an unweighted average across only the industry families a model has completed, so always read its coverage count.

WALL

Wall time

Median end-to-end seconds observed by the harness, including any reasoning trace. Lower is faster; it is measured latency, not provider-advertised speed.

1P$

First-party equivalent

Estimated public-API cost for the recorded tokens using first_party_pricing.json. It normalizes the workload; it is not a subscription charge or a claim about the operator's bill.

SAT

Saturation and exposure

2 of 4 industry families are flagged saturated in the generated metadata. Small gaps near the top can be noise, and public test exposure may inflate results.

ROUTE

Why one rank is not enough

Coverage differs, benchmark families measure different behaviors, and cost or latency may matter more than a small score gap. Start with the task winner below, then check coverage, wall time, and first-party cost-equivalent.

industry board

Mean accuracy across available industry families

Top 15 by generated mean. Coverage is explicit because a 1/4-family mean is not comparable to a 4/4-family mean.

mean accuracy x/4 completed families
  1. 1 grok-4.6-direct91.3% 2/4
  2. 2 grok-4.20-reasoning-direct90.3% 2/4
  3. 3 grok-4.3-latest-direct90.3% 2/4
  4. 4 codex-gpt-5.6-luna89.7% 2/4
  5. 5 grok-4-fast-direct89.3% 2/4
  6. 6 codex-gpt-5.6-sol88.7% 4/4
  7. 7 grok-4-1-fast-direct87.8% 2/4
  8. 8 codex-gpt-5.586.7% 1/4
  9. 9 codex-gpt-5.6-terra86.7% 2/4
  10. 10 grok-4.5-direct86.7% 4/4
  11. 11 grok-3-mini-direct86.1% 2/4
  12. 12 grok-4.3-direct86.1% 2/4
  13. 13 grok-build-0.186.1% 2/4
  14. 14 claude-opus-5-high85.2% 2/4
  15. 15 muse-glimmer-30b-q4_k_xl-local85.1% 3/4
VIEW ALL BENCHMARK FAMILIES →
decision layer · generated from use_case_pairing.json

What wins where

The strongest observed custom-task pairings. Ties stay visible; scores are rubric means out of 5, not industry accuracy percentages.

CUSTOM SUITE →
constraint check · generated leaderboard

Quality, latency, and cost-equivalent together

The first eight rows from the generated rank-eligible board. First-party equivalent is normalized per recorded trial here so different run totals do not masquerade as a per-call price.

RANKMODELMEAN / 5PASSMED WALL1P EQUIV / RECORDED TRIALN
1claude-fable4.4487.7%18.13s$0.0147660
2claude-opus-5-high4.4487.4%23.05snot mapped190
3claude-opus-54.3988.9%15.92s$0.015490
4ollama-cl/gemma44.3884.6%3.78snot mapped52
5grok-4.5-direct4.3381.4%12.24s$0.0286140
6mistral/magistral-small4.3282.4%1.67s$0.000634
7grok-4.6-direct4.3180.0%27.34s$0.002845
8ollama-cl/deepseek-v4-flash4.2982.4%6.45s$0.000334

The per-trial equivalent reflects this benchmark's observed prompt and output mix. It is useful for comparing this workload, not a universal token-price or subscription-billing claim.

coverage drill-down

The industry-tested fleet

81 model IDs have at least one industry-family result. Cards retain each family score rather than flattening the data to the headline mean.

ALL MODELS →
xai-direct reasoning

grok-4.6-direct

ACC 91.3% WALL 27.3s COVER 2/4
GPQA Diamond 92.0%
IFEval 90.7%
xai-direct reasoning

grok-4.20-reasoning-direct

ACC 90.3% WALL 20.8s COVER 2/4
GPQA Diamond 89.1%
IFEval 91.5%
xai-direct general

grok-4.3-latest-direct

ACC 90.3% WALL 8.0s COVER 2/4
GPQA Diamond 90.9%
IFEval 89.7%
codex-cli fast

codex-gpt-5.6-luna

ACC 89.7% WALL 20.3s COVER 2/4
GPQA Diamond 83.3%
HumanEval+ 96.0%
xai-direct fast

grok-4-fast-direct

ACC 89.3% WALL 6.5s COVER 2/4
GPQA Diamond 92.0%
IFEval 86.7%
codex-cli reasoning

codex-gpt-5.6-sol

ACC 88.7% WALL 11.9s COVER 4/4
GPQA Diamond 93.3%
HumanEval+ 92.0%
IFEval 77.3%
MMLU-Pro 92.0%
xai-direct reasoning

grok-4-1-fast-direct

ACC 87.8% WALL 5.8s COVER 2/4
GPQA Diamond 90.0%
IFEval 85.6%
codex-cli reasoning

codex-gpt-5.5

ACC 86.7% WALL 13.2s COVER 1/4
GPQA Diamond 86.7%
codex-cli general

codex-gpt-5.6-terra

ACC 86.7% WALL 10.7s COVER 2/4
GPQA Diamond 93.3%
HumanEval+ 80.0%
xai-direct reasoning

grok-4.5-direct

ACC 86.7% WALL 12.2s COVER 4/4
GPQA Diamond 87.3%
HumanEval+ not scored
IFEval 86.1%
MMLU-Pro not scored
xai-direct fast

grok-3-mini-direct

ACC 86.1% WALL 8.8s COVER 2/4
GPQA Diamond 83.3%
IFEval 88.9%
xai-direct general

grok-4.3-direct

ACC 86.1% WALL 4.8s COVER 2/4
GPQA Diamond 83.3%
IFEval 88.9%
xai-direct general

grok-build-0.1

ACC 86.1% WALL 18.6s COVER 2/4
GPQA Diamond 80.0%
IFEval 92.2%
claude-cli reasoning

claude-opus-5-high

ACC 85.2% WALL 23.1s COVER 2/4
GPQA Diamond 88.0%
HumanEval+ 82.5%
llama.cpp-local reasoning

muse-glimmer-30b-q4_k_xl-local

ACC 85.1% WALL 34.2s COVER 3/4
GPQA Diamond 68.0%
IFEval 87.3%
HumanEval+ 100.0%
claude-cli reasoning

claude-opus-5

ACC 80.6% WALL 15.9s COVER 4/4
GPQA Diamond 84.8%
HumanEval+ 78.0%
IFEval 71.3%
MMLU-Pro 88.4%
xai-direct general

grok-4.20-non-reasoning-direct

ACC 80.5% WALL 2.1s COVER 2/4
GPQA Diamond 70.0%
IFEval 91.1%
codex-cli fast

codex-gpt-5.3-spark

ACC 75.7% WALL 7.7s COVER 2/4
GPQA Diamond 63.3%
HumanEval+ 88.0%
READ BEFORE ROUTING

Coverage, question count, and grader differ by family. GPQA Diamond: Saturated above ~93% for frontier models; score deltas at top are within noise.; MMLU-Pro: Saturated above ~80% for frontier models on most disciplines; per-discipline ranking is the real signal. Public benchmarks can also be exposed in training data, so close scores are evidence for a shortlist—not proof of a universal winner.