Accuracy
Percent passed by that benchmark's reference verifier. The homepage mean is an unweighted average across only the industry families a model has completed, so always read its coverage count.
9 custom tasks sit beside GPQA Diamond, MMLU-Pro, IFEval and HumanEval+. Industry suites use benchmark-specific deterministic graders; custom tasks use rubric scoring with a partial second-judge validation layer.
Corpus accounting: 23,680 source records = 23,566 published benchmark trials + 114 quarantined harness-control records. See methodology.
Percent passed by that benchmark's reference verifier. The homepage mean is an unweighted average across only the industry families a model has completed, so always read its coverage count.
Median end-to-end seconds observed by the harness, including any reasoning trace. Lower is faster; it is measured latency, not provider-advertised speed.
Estimated public-API cost for the recorded tokens using first_party_pricing.json. It normalizes the workload; it is not a subscription charge or a claim about the operator's bill.
2 of 4 industry families are flagged saturated in the generated metadata. Small gaps near the top can be noise, and public test exposure may inflate results.
Coverage differs, benchmark families measure different behaviors, and cost or latency may matter more than a small score gap. Start with the task winner below, then check coverage, wall time, and first-party cost-equivalent.
Top 15 by generated mean. Coverage is explicit because a 1/4-family mean is not comparable to a 4/4-family mean.
The strongest observed custom-task pairings. Ties stay visible; scores are rubric means out of 5, not industry accuracy percentages.
The first eight rows from the generated rank-eligible board. First-party equivalent is normalized per recorded trial here so different run totals do not masquerade as a per-call price.
| RANK | MODEL | MEAN / 5 | PASS | MED WALL | 1P EQUIV / RECORDED TRIAL | N |
|---|---|---|---|---|---|---|
| 1 | claude-fable | 4.44 | 87.7% | 18.13s | $0.0147 | 660 |
| 2 | claude-opus-5-high | 4.44 | 87.4% | 23.05s | not mapped | 190 |
| 3 | claude-opus-5 | 4.39 | 88.9% | 15.92s | $0.0154 | 90 |
| 4 | ollama-cl/gemma4 | 4.38 | 84.6% | 3.78s | not mapped | 52 |
| 5 | grok-4.5-direct | 4.33 | 81.4% | 12.24s | $0.0286 | 140 |
| 6 | mistral/magistral-small | 4.32 | 82.4% | 1.67s | $0.0006 | 34 |
| 7 | grok-4.6-direct | 4.31 | 80.0% | 27.34s | $0.0028 | 45 |
| 8 | ollama-cl/deepseek-v4-flash | 4.29 | 82.4% | 6.45s | $0.0003 | 34 |
The per-trial equivalent reflects this benchmark's observed prompt and output mix. It is useful for comparing this workload, not a universal token-price or subscription-billing claim.
81 model IDs have at least one industry-family result. Cards retain each family score rather than flattening the data to the headline mean.
Coverage, question count, and grader differ by family. GPQA Diamond: Saturated above ~93% for frontier models; score deltas at top are within noise.; MMLU-Pro: Saturated above ~80% for frontier models on most disciplines; per-discipline ranking is the real signal. Public benchmarks can also be exposed in training data, so close scores are evidence for a shortlist—not proof of a universal winner.