- task_id
- visual_research
- task_class
- research
- target_model
- grok-4.5
- dependencies
- none
- acceptance_rubric
- sources + patterns
- retry_budget
- 1
- fallback_model
- minimax-m3
One question. Many models. A measured choice.
Dispatch Ledger is Herman’s multi-model decision system: identify the work, compare specialists, respect runtime constraints, and preserve evidence for why one model received the query. It is not a global leaderboard—and it is not the upstream Tokscale package, which remains a separately attributed telemetry input.
- 12
- MODELS IN ATLAS
- 108
- OBSERVED CELLS
- 9
- TASK CLASSES
- 42
- FRONTIER RECORDS
- 2026-07-30
- EVIDENCE THROUGH
A complex request becomes a small, judged program.
The requestor model keeps the conversation. When the work crosses the complexity gate, the router emits typed task envelopes, dispatches specialists in parallel, judges every result, condenses accepted evidence, and hands one compact payload back.
- task_id
- data_contract
- task_class
- code_review
- target_model
- zai/glm-5.2
- dependencies
- benchmark JSON
- acceptance_rubric
- schema + tests
- retry_budget
- 2
- fallback_model
- kimi-k3
- task_id
- implementation
- task_class
- multi_file
- target_model
- claude-opus
- dependencies
- A + B findings
- acceptance_rubric
- build + visual QA
- retry_budget
- 1
- fallback_model
- parent integrator
Different work exposes different winners.
Each cell is an observed custom-task mean out of 5. Hover or focus to reveal the sample size. Missing evidence remains blank.
Quality, speed, and cost-equivalent occupy the same map.
Higher is better. Left is faster. Bubble size represents public-API first-party cost-equivalent per recorded trial—not a subscription bill.
A routing verdict should leave a receipt.
The public example combines observed task evidence with the router's runtime vocabulary. Availability is checked at dispatch time.
Diagnose a concurrency bug spanning several modules
Experiments climb only when the evidence survives.
Methodology ladder for Router School; statuses describe the research process, not unreported model wins.
- 01Freeze the question setCOMPLETE
Version tasks and source manifests before tuning.
- 02Block signature leakageCOMPLETE
Group equivalent prompt signatures into one split.
- 03Run deterministic baselinesCOMPLETE
Compare rules and simple classifiers before training a router.
- 04Train bounded candidatesACTIVE
Change one registered variable per experiment cell.
- 05Apply falsification gatesREQUIRED
Reject gains that fail paired uncertainty, coverage, or minority recall.
- 06Publish surviving evidencePENDING
Only gate-passing candidates can replace the deterministic router.
The gap is the research program.
The task-specialist line is hindsight over observed task means. The ceiling is the rubric maximum, not an attainable production oracle.
Read the instruments honestly
- Custom-task scores are rubric means out of 5 and must not be mixed with industry benchmark accuracy percentages.
- Cost is public-API first-party equivalent for the recorded token mix; it is not subscription billing.
- Task-specialist hindsight selects from observed task means and is not a production-router accuracy claim.
- Cell sample sizes vary; every heatmap cell exposes n and missing cells remain missing.
- The perfect-score ceiling is a rubric bound, not an attainable oracle measurement.
Data snapshot: 2026-07-30 · Full methodology · Sealed HRB results · Raw evidence