Methodology
80 blind pairwise bouts across 8 harder-than-grid prompts (bug hunting, ADR writing, complex JSON, ruthless code review, literary fiction, agent planning, math chains, SQL fan-trap fix). Each bout judged by an independent 3-judge panel:
- minimax-m3 (Skynet)
- Claude Haiku (Claude CLI)
- Grok 4.3 (xAI direct)
None of the contenders judges itself. A/B orientation is hash-randomized per bout to cancel position bias. Consensus = majority vote; split = tie.
Final Arena Elo (80 bouts)
| Rank | Model | Elo | W/L/T |
|---|---|---|---|
| 1 | Claude Fable | 1611 | 20/7/5 |
| 2 | Claude Opus | 1517 | 17/10/5 |
| 3 | Claude Sonnet | 1480 | 11/17/4 |
| 4 | Grok 4.5 | 1479 | 12/14/6 |
| 5 | GPT-5.6 Sol | 1415 | 6/18/8 |
Key finding
The 9-task aggregate grid had Grok 4.5 at #1 (4.53 mean). The blind arena tells a different story: Fable wins decisively (Elo 1611 vs Grok’s 1479). Grok 4.5 drops to #4 in head-to-head — it aces structured tasks but loses on the harder multi-axis prompts (bug hunting, architecture tradeoffs, creative fiction).
GPT-5.6 Sol, which looked strong on the grid (#3 at 4.16), lands at the bottom of the arena (Elo 1415). The grid rewards breadth-of-task; the arena rewards depth-of-reasoning.