Verdict: GLM-5.3-Flash is a credible high-volume fleet worker: fast on short calls, exact at 900,041 observed prompt tokens, capable of vision and native tool calls, and broad enough for coding/reasoning support. It is not a blanket premium-model replacement, and this campaign does not certify the custom paid route for private prompts.
Officially, Z.ai describes GLM-5.3-Flash as a 320B-total / 18B-active, natively multimodal MoE with a 1M-token context window. The measured route was zai-coding/glm-5.3-flash through paid Skynet subscription inference.
Integrity boundary
- Requested and served ID:
zai-coding/glm-5.3-flashon all successful measured calls. - Data policy for this run: synthetic/public data only; no private repository or operator material was sent.
- The public API terms say API inputs/outputs are not used for training absent explicit agreement, but the Coding Plan documents do not clearly bind this custom Skynet proxy path to those API Services terms. Private-data status is therefore not certified.
- Official sources identify the prerelease alias
ox-alphaas GLM-5.3-Flash and advertise a 1M-token context window.
Custom suite
Primary MiniMax-judge mean: 4.84/5 (95% bootstrap CI 4.69–4.96), 45/45 valid, all nine task cells covered. Claude cross-check: 4.78/5 across 45 matched trials.
Four mandatory-thinking cells (code_gen_long, creative_write, refactor_existing_code, agentic_tool_use) were rerun at 12,288 output tokens after the legacy 3,000–5,000-token budgets produced truncated answers. One creative row still exhausted 12,288 and required a 65,536-token completion window. The truncated rows remain in evidence but are excluded from this headline score.
| Task | Mean / 5 | n |
|---|---|---|
agentic_prompt | 5.00 | 5 |
agentic_tool_use | 5.00 | 5 |
code_debug | 5.00 | 5 |
code_gen_long | 4.80 | 5 |
creative_write | 5.00 | 5 |
json_strict | 5.00 | 5 |
reasoning_multistep | 5.00 | 5 |
refactor_existing_code | 4.20 | 5 |
summarize | 4.60 | 5 |
Deterministic public benchmark samples
These are fixed 25-question subsamples, not full-suite numbers.
| Benchmark | Correct | End-to-end pass | Wilson 95% CI | Median latency |
|---|---|---|---|---|
| GPQA Diamond | 14/25 | 56.0% | 37.1–73.3% | 41.19s |
| IFEval | 17/25 | 68.0% | 48.4–82.8% | 34.30s |
| MMLU-Pro | 21/25 | 84.0% | 65.3–93.6% | 11.85s |
Capability probes
- Exact synthetic needle retrieval at 512,041 and 900,041 prompt tokens; the 900k call completed in 57.61s.
- Native function-call object emitted with the expected function name and arguments.
- Vision probe correctly returned
orangefrom an inline synthetic image. - Structured JSON succeeded when the output budget was raised.
- Important failure mode: 2 low-budget HTTP-200 calls returned empty visible content because mandatory reasoning consumed the completion budget. Give this model at least ~512 output tokens even for terse answers.
Fleet fit
- High-volume coding and reasoning leaves: strong candidate for bounded implementation drafts, debugging, refactors, structured output, and research synthesis where a premium lane is unnecessary.
- Long-context public-corpus triage: excellent fit for large public docs, logs, standards, and code snapshots; the route carried and retrieved a synthetic needle at 900k.
- Tool-using workers: viable for native function calling, provided the wrapper preserves
tool_callsand does not starve the reasoning budget. - Vision triage: useful for lightweight screenshots/diagrams; this was a minimal color probe, not a full vision benchmark.
- Not the creative lead: keep Fable/Grok-class specialists for voice-sensitive or flagship prose unless a separate human-preference evaluation reverses that call.
- Do not send private code yet: paid subscription inference is materially different from the anonymous preview, but paid status alone is not contractual proof of no-training treatment for this proxy route.
Evidence
Raw synthetic/public trials, capability receipts, and SHA-256 manifest are published under /evidence/glm-5.3-flash-20260826/. Method limitations are carried in the machine-readable campaign artifact, not hidden in footnotes.