Z.ai + MiniMax deep validation — 2026-08-20

A fresh, dedicated family campaign covering six live Z.ai and MiniMax routes. This is a supplemental validation layer beside the accumulated public corpus: it does not rewrite historical trials.

Protocol

  • Custom suite: all 9 tasks, 5 successful repetitions per model (45/model).
  • Judges: MiniMax M3 plus a blind independent Grok 4.6 re-judge on every successful custom response; the headline custom score is their mean.
  • Industry samples: GPQA Diamond, IFEval, and MMLU-Pro with deterministic benchmark-specific graders.
  • Failure handling: API failures are preserved as failed attempts, never scored as wrong answers. Z.ai 429 collisions from an initial concurrent wave were quarantined and rerun serially with backoff.
  • Confidence: custom mean uses a deterministic bootstrap 95% interval; industry end-to-end pass uses Wilson 95% intervals.
  • Judge-family caveat: MiniMax M3 shares a family with the MiniMax candidates. Grok 4.6 re-judged every custom output, but MiniMax custom consensus can retain self-family bias.

Custom-suite consensus

ModelConsensus /5 (95% CI)Both-judge passAPI successMedian wallJudge exact agree
zai-coding/glm-5-turbo4.167 (3.822–4.467)75.6%100.0%37.40s71.1%
minimax-m2.7-highspeed4.011 (3.578–4.422)71.1%100.0%25.61s73.3%
minimax-m33.922 (3.466–4.356)71.1%100.0%17.70s60.0%
zai-coding/glm-5.23.878 (3.344–4.400)75.6%100.0%18.87s66.7%
minimax-m2.73.867 (3.422–4.244)60.0%97.8%21.18s55.6%
zai-coding/glm-5.33.833 (3.278–4.367)73.3%76.3%23.64s77.8%

Deterministic industry samples

zai-coding/glm-5.3

BenchmarkEnd-to-end pass (95% CI)Successful / target
gpqa_diamond55.0% (42.5–66.9)60 / 60
ifeval58.3% (45.7–69.9)60 / 60
mmlu_pro85.0% (73.9–91.9)60 / 60

zai-coding/glm-5.2

BenchmarkEnd-to-end pass (95% CI)Successful / target
gpqa_diamond52.0% (33.5–70.0)23 / 25
ifeval76.0% (56.6–88.5)25 / 25
mmlu_pro84.0% (65.3–93.6)25 / 25

zai-coding/glm-5-turbo

BenchmarkEnd-to-end pass (95% CI)Successful / target
gpqa_diamond80.0% (60.9–91.1)23 / 25
ifeval80.0% (60.9–91.1)25 / 25
mmlu_pro48.0% (30.0–66.5)24 / 25

minimax-m3

BenchmarkEnd-to-end pass (95% CI)Successful / target
gpqa_diamond62.5% (47.0–75.8)40 / 40
ifeval62.5% (47.0–75.8)40 / 40
mmlu_pro67.5% (52.0–79.9)40 / 40

minimax-m2.7

BenchmarkEnd-to-end pass (95% CI)Successful / target
gpqa_diamond42.5% (28.5–57.8)40 / 40
ifeval67.5% (52.0–79.9)40 / 40
mmlu_pro47.5% (32.9–62.5)40 / 40

minimax-m2.7-highspeed

BenchmarkEnd-to-end pass (95% CI)Successful / target
gpqa_diamond48.0% (30.0–66.5)25 / 25
ifeval72.0% (52.4–85.7)25 / 25
mmlu_pro60.0% (40.7–76.6)25 / 25

Boundaries

  • The models were tested through the named subscription-backed routes; these results do not claim equivalence to every public API deployment or quantization.
  • Model identity is checked from returned model identifiers where the provider supplied one. Missing provider identity remains visible in the report data.
  • Creative-writing and refactor tasks remain judge-sensitive. Read the per-task values and judge-agreement fields rather than treating the family rank as universal.
  • ollama-cl/minimax-m2.5 was probed and excluded after a reproducible HTTP 403 route failure; it is not presented as a scored model.

Evidence manifest

FileBytesSHA-256
custom-5rep.jsonl1310313c989ffd8ff6521a76be91593bae52cde712cb5b6d06ffe88c89842939daeeea6
cross-judge-grok43.jsonl13571750172b9d05482dae5b0de980787cb8f4b0bd8f584fdc1b94bdd67be7b90aa3c8
industry-top.jsonl1023816f2ea45ac1f564b274243f3914f29d30a21698e135b9db9b1bdfc933c062f9b6a
industry-peers.jsonl8315764b6d9df131473b0cf443f254ee0af7af94879d0f427c121d7dae83ac9922554a
industry-glm53-clean.jsonl744138ef2ab74df4ef82e9418578f0c1973b8f3d73a66c96d62177a0f8c1fb26310275
industry-repair.jsonl5892fc00ebc9289f08baba8044d40a2da8c8a536e514ab16eba38a799d47b16f9f02

Machine-readable campaign data is published in data/zai_minimax_deep_20260820.json.