HERMAN ROUTING BENCHMARK · SEALED RESULTS

Model evidence, kept out of the routing methodology.

HRB measures models on a frozen private task grid. This page owns those benchmark results. The Routing Lab separately explains how Dispatch Ledger chooses and verifies specialist lanes.

CONFIRMATORY · SINGLE-FOUNDATION-MODEL TRACK

Same private task grid. Exact model identities. One sealed run.

This is a direct foundation-model comparison on HRB’s 96-cell confirmatory matrix. It is not a routed-system result.

48
TASKS PER MODEL
24
GROUPED TEMPLATES
10000
BOOTSTRAP RESAMPLES
SEALED
INTEGRITY STATUS
EXACT MODEL ID

claude-opus-5

100%48/48 correct · pass@1
Median wall time
2.739 s
Wire cost
$0.654024
Cost coverage
48/48 trials
Quality CI
[1.00, 1.00]
EXACT MODEL ID

gpt-5.6-sol

100%48/48 correct · pass@1
Median wall time
13.892 s
Wire cost
UNREPORTED
Cost coverage
0/48 trials
Quality CI
[1.00, 1.00]
SEALED RECEIPTc141e8ef2bfa9d8e…

paired-grouped-percentile by template_id, seed 1729, 95% confidence.

LIMITATIONS
  • two models only
  • one pass@1 draw per task
  • 24 bootstrap groups
  • three templates per family; family results are descriptive only