HRB · SECONDARY DOCS · OWLFORGE LOCAL CALIBRATION

Owlforge local calibration

Privacy-safe projection of sealed v0.3 natural-format and v0.4 repaired uniform-JSON campaigns. The frozen v0.3 and repaired v0.4 claims remain separate because prompt/grader contracts and response-format policy changed together.

PRIMARY CLAIM

Natural-format exact success

v0.3 measures exact natural-format instruction following with a deterministic pilot grader. It is not a general intelligence score. Only two of twelve JSON-expecting tasks were schema-constrained.

REPAIRED CAMPAIGN · SEPARATE CLAIM

Repaired uniform-JSON exact success

v0.4 uses repaired prompt/grader contracts and a uniform native JSON response constraint. It is a separate sealed campaign, not a causal format-only comparison with v0.3. Source: sealed owlforge-calibration-v0.4 aggregate.

4
MODELS
12
TASKS / MODEL
960
MEASUREMENTS / CAMPAIGN
SEALED
INTEGRITY
MODEL COMPARISON

Frozen v0.3 beside repaired uniform-JSON v0.4

Each card shows two separately sealed campaign claims. Do not treat the change as a causal format uplift or average the columns.

GPT-OSS · 20B · think-low

GPT-OSS 20B

gpt-oss-20b
v0.3 natural-format exact 2/12 17% of 12 tasks
v0.4 repaired uniform-JSON 11/12 separate sealed campaign
v0.4 median latency
3.176 s
v0.4 P95 latency
5.552 s
v0.4 E2E throughput
6.4 tok/s
v0.4 native decode
172.4 tok/s
v0.4 energy / out tok
19.71 J
v0.4 peak VRAM
13414 MiB
v0.4 peak power
346 W
v0.4 peak temp
66 °C
v0.4 telemetry coverage
98.8%
v0.4 invalid grades
0
V0.4 FAILURE CLASSES
  • success 220
  • wrong_answer 20
Natural-format task matrix (2/12)
Per-task natural-format exact success for GPT-OSS 20B
TaskExactDominant class
failure-taxonomy-afailformat_failure
failure-taxonomy-bfailformat_failure
privacy-afailformat_failure
privacy-bfailformat_failure
provenance-apasssuccess
provenance-bfailformat_failure
receipt-afailformat_failure
receipt-bfailformat_failure
route-choice-afailwrong_answer
route-choice-bfailwrong_answer
structured-output-apasssuccess
structured-output-bfailwrong_answer
v0.4 repaired task matrix (11/12)
Per-task repaired uniform-JSON exact success for GPT-OSS 20B
TaskExactDominant class
failure-taxonomy-apasssuccess
failure-taxonomy-bfailwrong_answer
privacy-apasssuccess
privacy-bpasssuccess
provenance-apasssuccess
provenance-bpasssuccess
receipt-apasssuccess
receipt-bpasssuccess
route-choice-apasssuccess
route-choice-bpasssuccess
structured-output-apasssuccess
structured-output-bpasssuccess
QWEN3.5 · 9B · nonthink

Qwen3.5 9B

qwen35-9b-q4
v0.3 natural-format exact 5/12 42% of 12 tasks
v0.4 repaired uniform-JSON 8/12 separate sealed campaign
v0.4 median latency
2.307 s
v0.4 P95 latency
4.200 s
v0.4 E2E throughput
16.8 tok/s
v0.4 native decode
124.4 tok/s
v0.4 energy / out tok
7.35 J
v0.4 peak VRAM
8596 MiB
v0.4 peak power
274 W
v0.4 peak temp
64 °C
v0.4 telemetry coverage
98.8%
v0.4 invalid grades
0
V0.4 FAILURE CLASSES
  • format_failure 80
  • success 160
Natural-format task matrix (5/12)
Per-task natural-format exact success for Qwen3.5 9B
TaskExactDominant class
failure-taxonomy-afailwrong_answer
failure-taxonomy-bpasssuccess
privacy-apasssuccess
privacy-bfailpolicy_violation
provenance-apasssuccess
provenance-bpasssuccess
receipt-afailincomplete_artifact
receipt-bfailformat_failure
route-choice-afailwrong_answer
route-choice-bfailwrong_answer
structured-output-apasssuccess
structured-output-bfailwrong_answer
v0.4 repaired task matrix (8/12)
Per-task repaired uniform-JSON exact success for Qwen3.5 9B
TaskExactDominant class
failure-taxonomy-apasssuccess
failure-taxonomy-bpasssuccess
privacy-afailformat_failure
privacy-bfailformat_failure
provenance-apasssuccess
provenance-bpasssuccess
receipt-afailformat_failure
receipt-bfailformat_failure
route-choice-apasssuccess
route-choice-bpasssuccess
structured-output-apasssuccess
structured-output-bpasssuccess
QWEN3.5 · 4B · nonthink

Qwen3.5 4B

qwen35-4b-q4
v0.3 natural-format exact 3/12 25% of 12 tasks
v0.4 repaired uniform-JSON 4/12 separate sealed campaign
v0.4 median latency
2.442 s
v0.4 P95 latency
3.545 s
v0.4 E2E throughput
29.1 tok/s
v0.4 native decode
172.5 tok/s
v0.4 energy / out tok
4.35 J
v0.4 peak VRAM
5744 MiB
v0.4 peak power
353 W
v0.4 peak temp
64 °C
v0.4 telemetry coverage
98.3%
v0.4 invalid grades
20
V0.4 FAILURE CLASSES
  • format_failure 40
  • incomplete_artifact 20
  • output_budget_exhaustion 20
  • success 80
  • wrong_answer 80
Natural-format task matrix (3/12)
Per-task natural-format exact success for Qwen3.5 4B
TaskExactDominant class
failure-taxonomy-afailwrong_answer
failure-taxonomy-bfailwrong_answer
privacy-apasssuccess
privacy-bpasssuccess
provenance-afailwrong_answer
provenance-bpasssuccess
receipt-afailformat_failure
receipt-bfailformat_failure
route-choice-afailwrong_answer
route-choice-bfailwrong_answer
structured-output-afailwrong_answer
structured-output-bfailwrong_answer
v0.4 repaired task matrix (4/12)
Per-task repaired uniform-JSON exact success for Qwen3.5 4B
TaskExactDominant class
failure-taxonomy-afailoutput_budget_exhaustion
failure-taxonomy-bfailwrong_answer
privacy-afailformat_failure
privacy-bfailformat_failure
provenance-afailwrong_answer
provenance-bpasssuccess
receipt-apasssuccess
receipt-bfailincomplete_artifact
route-choice-afailwrong_answer
route-choice-bpasssuccess
structured-output-afailwrong_answer
structured-output-bpasssuccess
QWEN2.5 · 7B · nonthink

Qwen2.5 7B

qwen25-7b-1m-q4
v0.3 natural-format exact 1/12 8% of 12 tasks
v0.4 repaired uniform-JSON 9/12 separate sealed campaign
v0.4 median latency
1.500 s
v0.4 P95 latency
2.785 s
v0.4 E2E throughput
17.6 tok/s
v0.4 native decode
151.7 tok/s
v0.4 energy / out tok
7.09 J
v0.4 peak VRAM
6614 MiB
v0.4 peak power
165 W
v0.4 peak temp
65 °C
v0.4 telemetry coverage
96.2%
v0.4 invalid grades
0
V0.4 FAILURE CLASSES
  • incomplete_artifact 20
  • success 180
  • wrong_answer 40
Natural-format task matrix (1/12)
Per-task natural-format exact success for Qwen2.5 7B
TaskExactDominant class
failure-taxonomy-afailformat_failure
failure-taxonomy-bfailformat_failure
privacy-afailformat_failure
privacy-bfailformat_failure
provenance-afailformat_failure
provenance-bfailformat_failure
receipt-afailformat_failure
receipt-bfailformat_failure
route-choice-afailformat_failure
route-choice-bfailformat_failure
structured-output-apasssuccess
structured-output-bfailwrong_answer
v0.4 repaired task matrix (9/12)
Per-task repaired uniform-JSON exact success for Qwen2.5 7B
TaskExactDominant class
failure-taxonomy-afailwrong_answer
failure-taxonomy-bpasssuccess
privacy-apasssuccess
privacy-bpasssuccess
provenance-afailwrong_answer
provenance-bpasssuccess
receipt-apasssuccess
receipt-bfailincomplete_artifact
route-choice-apasssuccess
route-choice-bpasssuccess
structured-output-apasssuccess
structured-output-bpasssuccess
CONTEXT · COLD / WARM

v0.4: 8k vs 32k lanes, cold load vs warm residency

Cold cells unload every model and verify empty GPU residency before reload. Warm cells reuse a resident model. These rows come from the repaired v0.4 campaign.

ModelCtxCold med latWarm med latΔ load (c−w)Warm/cold tok/sCold VRAMWarm energy J/tok
GPT-OSS 20B32k5.019 s0.829 s4.012 s5.57×13414 MiB7.20
GPT-OSS 20B8k5.021 s0.832 s4.004 s5.56×12790 MiB7.43
Qwen2.5 7B32k2.724 s0.345 s2.321 s8.22×6614 MiB1.82
Qwen2.5 7B8k2.704 s0.345 s2.307 s8.19×5222 MiB1.82
Qwen3.5 4B32k3.223 s0.501 s2.672 s5.32×5744 MiB1.72
Qwen3.5 4B8k3.193 s0.496 s2.668 s5.28×4928 MiB1.72
Qwen3.5 9B32k3.832 s0.550 s3.255 s6.46×8596 MiB2.38
Qwen3.5 9B8k3.821 s0.552 s3.248 s6.44×7780 MiB2.39
INTEGRITY

Sealed hashes for the projected run

Status
SEALED
Run ID
owlforge-calibration-v0.3-20260817
Campaign
owlforge-calibration-v0.3
Aggregate canonical SHA-256
69e1480f25de2ebf1993b92d278f3a44aa61569affe7eaa7de952a67395e6b16
Aggregate file SHA-256
33d5bd7813188928bcb7d31d8dbab8d5d0c74548cd19118af705abed70fb7609
Contract SHA-256
517d6b66c34956635277320304f2c16b6ab6b4e8cfbf209c5395c5b528a4d160
Schedule SHA-256
dff46e6f9e76e997640081e12eba45fe9a935beaa4613afece21e230f8fb8991
Receipt head SHA-256
e39ed30eb4641df3b9461434cd94cb4e808f91e95cb9ee03b9550040ad36290b
Journal head SHA-256
90b9f04c295f9c78fc275b1947d2f31a51e1766cb5ff6f5398d60d3614f4bb5a
Raw SHA-256
ed6a95355f8a137d8bffc814fa927999c2c045ae3e08f080cdcec1f7b923002d
Source schema
hrb-local-calibration-aggregate-0.1
v0.4 status
SEALED
v0.4 run ID
owlforge-calibration-v0.4-20260817
v0.4 campaign
owlforge-calibration-v0.4
v0.4 aggregate canonical SHA-256
b1cda8850ce0ab56e1105ab3790eb7c5f8e0ada27537e4a0805b8ace2c170ff8
v0.4 contract SHA-256
96ea6ccca5b803a20cc429cb8d001428c52a01101168a058b57b76f33f97ed39
v0.4 receipt head SHA-256
04df4744a04a02fa93e0682a8a96b97d7995a99a10712009c90eb9a25533446a
v0.4 journal head SHA-256
5791ccc486a9d1785800bbb42105319f51a178d6e270c9b573215d591a2e3778
METRIC SEMANTICS

How to read the numbers

v0.3 natural-format campaign

energy
time-windowed nvidia-smi energy; null coverage is explicit and excluded from ratios
format_policy
mixed: native JSON constraint only where declared by campaign task response_formats
native_throughput
native Ollama token counts divided by native prompt/eval duration; reported separately
quality
exact deterministic pilot grader outcome; not a general intelligence score
repetition_reduction
mean_per_task_lane_cache_state_before_template_grouping
statistical_unit
template_id
thinking_mode_caveat
thinking mode is collinear with model identity in v0.3 and cannot estimate a causal mode effect
throughput
output_tokens_per_end_to_end_second

v0.4 repaired uniform-JSON campaign

energy
time-windowed nvidia-smi energy; null coverage is explicit and excluded from ratios
format_policy
uniform: native JSON constraint for all 12 campaign tasks
native_throughput
native Ollama token counts divided by native prompt/eval duration; reported separately
quality
exact deterministic pilot grader outcome; not a general intelligence score
repetition_reduction
mean_per_task_lane_cache_state_before_template_grouping
statistical_unit
template_id
thinking_mode_caveat
thinking mode is collinear with model identity and cannot estimate a cross-model causal mode effect
throughput
output_tokens_per_end_to_end_second
CAVEATS

Boundaries and non-claims

  • Natural-format exact success is the frozen v0.3 claim; repaired uniform-JSON exact success is the separate sealed v0.4 claim.
  • v0.4 changed prompt/grader contracts and response-format policy together, so the difference from v0.3 is not a causal JSON-format uplift.
  • The repaired v0.2 prompts expose every expected vocabulary while retaining exact keys, strict types, arithmetic, privacy, provenance, and receipt requirements.
  • Thinking mode is collinear with model identity in v0.3 and cannot estimate a causal mode effect.
  • Quality is exact deterministic pilot-grader outcome, not a general intelligence or chat preference score.
  • Cold measurements globally unload all models, verify empty residency, then reload the target; temporary GPU idle between cells is expected.
  • Energy is time-windowed nvidia-smi; null telemetry coverage is excluded from ratios.
  • This projection excludes prompts, receipts bodies, journal lines, and model artifacts.
  • Local calibration is not mixed into custom-suite 0-5 scores, industry accuracy, or Routing Lab atlas ranks.