ZAI-CODING
zai-coding/glm-5.3
Z.ai GLM-5.3 via GLM Coding Plan subscription (1M context, max effort)
★VALS INDEX
—
Accuracy
⏱VALS INDEX
26.3s
Latency (median)
$VALS INDEX
—
Cost / Test (1P-est)
SCORE STATUSnot scored
TEST SETS9
RECORDED TRIALS45
CUSTOM PASS76%
DEEP VALIDATION · 2026-08-20
FULL REPORT →Dual-judge family campaign
CONSENSUS MEAN3.833 / 5
BOTH-JUDGE PASS73.3%
CROSS-JUDGED45 / 45
MEDIAN WALL23.64s
GPQA_DIAMONDserialized verification sample
55.0%
60/60 responses
IFEVALserialized verification sample
58.3%
60/60 responses
MMLU_PROserialized verification sample
85.0%
60/60 responses
No full-corpus industry row for this model yet. The dedicated serialized validation sample is published above; it remains separate from the accumulated leaderboard.
Custom 9-task suite
Mean /5n=trials
CUSTOM
agentic_prompt
★
5.00
5 trials
CUSTOM
agentic_tool_use
★
5.00
5 trials
CUSTOM
code_gen_long
★
5.00
5 trials
CUSTOM
json_strict
★
5.00
5 trials
CUSTOM
reasoning_multistep
★
5.00
5 trials
CUSTOM
summarize
★
4.80
5 trials
CUSTOM
code_debug
★
3.80
5 trials
CUSTOM
creative_write
★
1.20
5 trials
CUSTOM
refactor_existing_code
-
1.20
5 trials
Per-task distribution
Score
The accumulated custom suite below uses minimax-m3 + claude-fable-5 consensus. The dedicated 2026-08-20 campaign above is supplemental and uses minimax-m3 + a blind grok-4.6 re-judge on every successful response. Industry suites use deterministic family-specific graders. See methodology and the campaign report.