zai-coding · coding
ZAI-CODING

zai-coding/glm-5.3

Z.ai GLM-5.3 via GLM Coding Plan subscription (1M context, max effort)

⚖ COMPARE MODELS
VALS INDEX
Accuracy
VALS INDEX
26.3s
Latency (median)
$VALS INDEX
Cost / Test (1P-est)
SCORE STATUSnot scored
TEST SETS9
RECORDED TRIALS45
CUSTOM PASS76%
DEEP VALIDATION · 2026-08-20

Dual-judge family campaign

FULL REPORT →
CONSENSUS MEAN3.833 / 5
BOTH-JUDGE PASS73.3%
CROSS-JUDGED45 / 45
MEDIAN WALL23.64s
GPQA_DIAMONDserialized verification sample
55.0%
60/60 responses
IFEVALserialized verification sample
58.3%
60/60 responses
MMLU_PROserialized verification sample
85.0%
60/60 responses

No full-corpus industry row for this model yet. The dedicated serialized validation sample is published above; it remains separate from the accumulated leaderboard.

Custom 9-task suite

Mean /5n=trials
CUSTOM agentic_prompt
5.00
5 trials
CUSTOM agentic_tool_use
5.00
5 trials
CUSTOM code_gen_long
5.00
5 trials
CUSTOM json_strict
5.00
5 trials
CUSTOM reasoning_multistep
5.00
5 trials
CUSTOM summarize
4.80
5 trials
CUSTOM code_debug
3.80
5 trials
CUSTOM creative_write
1.20
5 trials
CUSTOM refactor_existing_code -
1.20
5 trials
Custom suite · scored 0–5 by consensus (m3 + Fable-5) ★ = task scored a 5 at least once · sorted by mean score

Per-task distribution

Score

The accumulated custom suite below uses minimax-m3 + claude-fable-5 consensus. The dedicated 2026-08-20 campaign above is supplemental and uses minimax-m3 + a blind grok-4.6 re-judge on every successful response. Industry suites use deterministic family-specific graders. See methodology and the campaign report.