— · general

zai-coding/glm-4.6

No description available.

⚖ COMPARE MODELS
VALS INDEX
46.3% +/- 38.4
Accuracy
VALS INDEX
Latency (median)
$VALS INDEX
Cost / Test (1P-est)
SCORE STATUSscored
TEST SETS3
RECORDED TRIALS108
CUSTOM PASS

Industry benchmarks

Accuracyattempts · responded
GPQA_DIAMOND GPQA Diamond
3.3%
30 attempts · 2 responded
IFEVAL IFEval
77.2%
30 attempts · 30 responded
MMLU_PRO MMLU-Pro
58.3%
48 attempts · 46 responded
Industry benchmark Bars = accuracy %, black cap = that score

All trials on this page were run through the harness. Industry suites use deterministic family-specific reference graders (no LLM judge). The custom suite is graded by consensus of minimax-m3 + claude-fable-5. See /industry/ for the full benchmark suite.