Model evidence, kept out of the routing methodology.
HRB measures models on a frozen private task grid. This page owns those benchmark results. The Routing Lab separately explains how Dispatch Ledger chooses and verifies specialist lanes.
CONFIRMATORY · SINGLE-FOUNDATION-MODEL TRACK
Same private task grid. Exact model identities. One sealed run.
This is a direct foundation-model comparison on HRB’s 96-cell confirmatory matrix. It is not a routed-system result.
48
TASKS PER MODEL
24
GROUPED TEMPLATES
10000
BOOTSTRAP RESAMPLES
SEALED
INTEGRITY STATUS
EXACT MODEL ID
claude-opus-5
100%48/48 correct · pass@1
Median wall time
2.739 s
Wire cost
$0.654024
Cost coverage
48/48 trials
Quality CI
[1.00, 1.00]
EXACT MODEL ID
gpt-5.6-sol
100%48/48 correct · pass@1
Median wall time
13.892 s
Wire cost
UNREPORTED
Cost coverage
0/48 trials
Quality CI
[1.00, 1.00]
SEALED RECEIPTc141e8ef2bfa9d8e…
paired-grouped-percentile by template_id, seed 1729, 95% confidence.
LIMITATIONS
two models only
one pass@1 draw per task
24 bootstrap groups
three templates per family; family results are descriptive only