A fresh, dedicated family campaign covering six live Z.ai and MiniMax routes. This is a supplemental validation layer beside the accumulated public corpus: it does not rewrite historical trials.
Protocol
Custom suite: all 9 tasks, 5 successful repetitions per model (45/model).
Judges: MiniMax M3 plus a blind independent Grok 4.6 re-judge on every successful custom response; the headline custom score is their mean.
Industry samples: GPQA Diamond, IFEval, and MMLU-Pro with deterministic benchmark-specific graders.
Failure handling: API failures are preserved as failed attempts, never scored as wrong answers. Z.ai 429 collisions from an initial concurrent wave were quarantined and rerun serially with backoff.
Confidence: custom mean uses a deterministic bootstrap 95% interval; industry end-to-end pass uses Wilson 95% intervals.
Judge-family caveat: MiniMax M3 shares a family with the MiniMax candidates. Grok 4.6 re-judged every custom output, but MiniMax custom consensus can retain self-family bias.
The models were tested through the named subscription-backed routes; these results do not claim equivalence to every public API deployment or quantization.
Model identity is checked from returned model identifiers where the provider supplied one. Missing provider identity remains visible in the report data.
Creative-writing and refactor tasks remain judge-sensitive. Read the per-task values and judge-agreement fields rather than treating the family rank as universal.
ollama-cl/minimax-m2.5 was probed and excluded after a reproducible HTTP 403 route failure; it is not presented as a scored model.