skynet · fast-reasoning
SKYNET

zai-coding/glm-5.3-flash

Z.ai GLM-5.3-Flash via paid Skynet subscription route; observed 900k context retrieval

⚖ COMPARE MODELS
VALS INDEX
69.3% +/- 11.5
Accuracy
VALS INDEX
45.5s
Latency (median)
$VALS INDEX
Cost / Test (1P-est)
SCORE STATUSscored
TEST SETS3
RECORDED TRIALS75
CUSTOM PASS
EXACT ROUTE · 2026-08-26

Paid-route capability campaign

FULL REPORT →

Run on synthetic/public data only. The primary 45-trial custom score uses MiniMax M3; a 45-trial independent cross-check uses Claude. Z.ai's API no-training language is documented, but coverage of this Coding Plan proxy route is not contractually clear, so private-data use is not certified.

PRIMARY MEAN4.84 / 5
CUSTOM COVERAGE45 / 45
CONTEXT PROVED900,041 tokens
PRIVATE DATANOT CERTIFIED

Industry benchmarks

Accuracyattempts · responded
GPQA_DIAMOND GPQA Diamond
56.0%
25 attempts · 25 responded
IFEVAL IFEval
68.0%
25 attempts · 25 responded
MMLU_PRO MMLU-Pro
84.0%
25 attempts · 25 responded
Industry benchmark Bars = accuracy %, black cap = that score

Custom 9-task suite

Mean /5n=trials
CUSTOM agentic_prompt
5.00
5 trials
CUSTOM agentic_tool_use
5.00
5 trials
CUSTOM code_debug
5.00
5 trials
CUSTOM creative_write
5.00
5 trials
CUSTOM json_strict
5.00
5 trials
CUSTOM reasoning_multistep
5.00
5 trials
CUSTOM code_gen_long
4.80
5 trials
CUSTOM summarize
4.60
5 trials
CUSTOM refactor_existing_code
4.20
5 trials
GLM campaign · displayed scores = MiniMax M3 primary judge ★ = task scored a 5 at least once · sorted by mean score

Per-task distribution

Score

The displayed custom scores are the MiniMax M3 primary-judge results for this 45-trial campaign. A blind claude-opus-5 re-judge covered the same 45 responses (4.78/5 vs 4.84/5 primary; 4.81/5 consensus). Industry suites use deterministic family-specific graders. See methodology and the campaign report.