Round 9 — Blind Arena: Top 5 Head-to-Head

Methodology

80 blind pairwise bouts across 8 harder-than-grid prompts (bug hunting, ADR writing, complex JSON, ruthless code review, literary fiction, agent planning, math chains, SQL fan-trap fix). Each bout judged by an independent 3-judge panel:

  • minimax-m3 (Skynet)
  • Claude Haiku (Claude CLI)
  • Grok 4.3 (xAI direct)

None of the contenders judges itself. A/B orientation is hash-randomized per bout to cancel position bias. Consensus = majority vote; split = tie.

Final Arena Elo (80 bouts)

RankModelEloW/L/T
1Claude Fable161120/7/5
2Claude Opus151717/10/5
3Claude Sonnet148011/17/4
4Grok 4.5147912/14/6
5GPT-5.6 Sol14156/18/8

Key finding

The 9-task aggregate grid had Grok 4.5 at #1 (4.53 mean). The blind arena tells a different story: Fable wins decisively (Elo 1611 vs Grok’s 1479). Grok 4.5 drops to #4 in head-to-head — it aces structured tasks but loses on the harder multi-axis prompts (bug hunting, architecture tradeoffs, creative fiction).

GPT-5.6 Sol, which looked strong on the grid (#3 at 4.16), lands at the bottom of the arena (Elo 1415). The grid rewards breadth-of-task; the arena rewards depth-of-reasoning.