DISPATCH LEDGER · ROUTING LAB · OBSERVED DATA, EXPLAINED

One question. Many models. A measured choice.

Dispatch Ledger is Herman’s multi-model decision system: identify the work, compare specialists, respect runtime constraints, and preserve evidence for why one model received the query. It is not a global leaderboard—and it is not the upstream Tokscale package, which remains a separately attributed telemetry input.

12
MODELS IN ATLAS
108
OBSERVED CELLS
9
TASK CLASSES
42
FRONTIER RECORDS
2026-07-30
EVIDENCE THROUGH
FIGURE 01 · ORCHESTRATION LIFECYCLE

A complex request becomes a small, judged program.

The requestor model keeps the conversation. When the work crosses the complexity gate, the router emits typed task envelopes, dispatches specialists in parallel, judges every result, condenses accepted evidence, and hands one compact payload back.

DIRECT PATHRequestor answers without fan-outNo decomposition overhead when the task is bounded.
TASK Agrok-4.5
task_id
visual_research
task_class
research
target_model
grok-4.5
dependencies
none
acceptance_rubric
sources + patterns
retry_budget
1
fallback_model
minimax-m3
RUNNING
TASK Bzai/glm-5.2
task_id
data_contract
task_class
code_review
target_model
zai/glm-5.2
dependencies
benchmark JSON
acceptance_rubric
schema + tests
retry_budget
2
fallback_model
kimi-k3
ACCEPTED
TASK Cclaude-opus
task_id
implementation
task_class
multi_file
target_model
claude-opus
dependencies
A + B findings
acceptance_rubric
build + visual QA
retry_budget
1
fallback_model
parent integrator
FALLBACK
REJECTED RESULTretry same lane → fallback model → parent resolvesRetries are bounded; a failed specialist never becomes evidence just because it returned text.
INSPECT THE PIPELINESelect any stageEvery node exposes what enters, what leaves, and what must be true before the next edge fires.
FIGURE 02 · CAPABILITY ATLAS

Different work exposes different winners.

Each cell is an observed custom-task mean out of 5. Hover or focus to reveal the sample size. Missing evidence remains blank.

Select a cell to inspect its evidence.
FIGURE 03 · TRADE-OFF PLANE

Quality, speed, and cost-equivalent occupy the same map.

Higher is better. Left is faster. Bubble size represents public-API first-party cost-equivalent per recorded trial—not a subscription bill.

Choose a point or a key entry to inspect the model.
FIGURE 04 · INSPECTABLE CHOICE

A routing verdict should leave a receipt.

The public example combines observed task evidence with the router's runtime vocabulary. Availability is checked at dispatch time.

QUERY CLASScode_debug

Diagnose a concurrency bug spanning several modules

task fit100
observed task score5
sample size15
availabilitychecked at dispatch time
SELECTED LANEgrok-4.5-direct
FIGURE 05 · ROUTER SCHOOL

Experiments climb only when the evidence survives.

Methodology ladder for Router School; statuses describe the research process, not unreported model wins.

  1. 01
    Freeze the question set

    Version tasks and source manifests before tuning.

    COMPLETE
  2. 02
    Block signature leakage

    Group equivalent prompt signatures into one split.

    COMPLETE
  3. 03
    Run deterministic baselines

    Compare rules and simple classifiers before training a router.

    COMPLETE
  4. 04
    Train bounded candidates

    Change one registered variable per experiment cell.

    ACTIVE
  5. 05
    Apply falsification gates

    Reject gains that fail paired uncertainty, coverage, or minority recall.

    REQUIRED
  6. 06
    Publish surviving evidence

    Only gate-passing candidates can replace the deterministic router.

    PENDING
FIGURE 06 · THE HEADROOM

The gap is the research program.

The task-specialist line is hindsight over observed task means. The ceiling is the rubric maximum, not an attainable production oracle.

one best global model
4.44claude-fable
task-specialist hindsight
4.58best observed model mean per task, n >= 10
perfect-score ceiling
5.00rubric maximum; not an observed router

Read the instruments honestly

  • Custom-task scores are rubric means out of 5 and must not be mixed with industry benchmark accuracy percentages.
  • Cost is public-API first-party equivalent for the recorded token mix; it is not subscription billing.
  • Task-specialist hindsight selects from observed task means and is not a production-router accuracy claim.
  • Cell sample sizes vary; every heatmap cell exposes n and missing cells remain missing.
  • The perfect-score ceiling is a rubric bound, not an attainable oracle measurement.

Data snapshot: 2026-07-30 · Full methodology · Sealed HRB results · Raw evidence