Model
Model explorer
GLM-4.6
OPENZhipu AI / Z.ai (Tsinghua) · GLM-4.6 family · released Sep 30, 2025
Flagship coding/agentic update to GLM-4.5; expanded context to 200K, ~15% more token-efficient, near-parity with Claude Sonnet 4 on real-world coding (CC-Bench).
ReasoningCodingVisionFunction callingTool useAgentic
1899.3
Elo · rank #83
Parameters
357B
Active params
32B (MoE)
Context
200K tokens
Architecture
Mixture-of-Experts, 357B total / 32B active, 200K context (up from 128K in GLM-4.5), hybrid thinking/non-thinking response modes
License
MIT
Languages
—
API price (in/out)
$0.6 / $2.2
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
best: GPT-5.2 · 100.0%
best: Kimi K3 · 91.2%
best: GPT-6 Astra · 96.0%
best: Claude Opus 5 · 64.7%
best: MiniMax M3 · 83.0%
best: DeepSeek-V4-Pro (Think Max) · 93.5%
best: Claude Opus 5 · 96.0%
best: Claude Opus 4.6 · 99.3%
best: Claude Opus 4.5 (High) · 59.3%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
—
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
PermissiveQLoRA242.8 GB2× H200 141GB
LoRA760.4 GB4× B200 192GB
Full fine-tune5729.9 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $16.55 (2× H200 141GB)
API price $0.6/$2.2 · each benchmark row carries its own source badge (see methodology)