Model
Model explorer

Qwen3.8-Max-0902

CLOSED
Alibaba · Qwen3.8 family · released Sep 2, 2026

A September 2, 2026 post-training snapshot of Qwen3.8-Max (alias `qwen3.8-max-2026-09-02`), not a new model underneath — Alibaba's own changelog describes deeper coding on engineering-scale projects, stronger multi-tool orchestration and refined chart/document/multimodal perception, with the 1M context, thinking mode and tool ecosystem unchanged. Pricing is unchanged at $2/$6 per Mtok ($0.25 cached). The bare `qwen3.8-max` endpoint transitions to this snapshot on September 5, 2026 at the same rates. Only three rows are recorded: Alibaba published a release chart rather than a full comparison table, and its own framing is that the strongest evidence is the snapshot-over-snapshot improvement rather than a clean cross-lab ranking, since the comparison models are run under different harnesses. Recorded values come from BenchmarkList's transcription of that chart, which links the QwenCloud model page and the Alibaba_Qwen launch post as its sources. READ ITS RANK WITH THAT IN MIND: three rows across two categories is exactly the D20 coverage floor, and all three are vendor-chosen rows on which the snapshot does well, so its overall position rests on far thinner evidence than the models around it. It should move once a fuller comparison table or an independent evaluation lands.

ReasoningCodingVisionFunction callingTool useAgentic
3092.0
Elo · rank #4
Parameters
2400B
Active params
95B (MoE)
Context
1M tokens
Architecture
Post-trained refresh of the Qwen3.8-Max checkpoint — same 2.4T total / ~95B active sparse MoE underneath, 1M-token context, thinking mode and the full built-in tool ecosystem retained
License
Proprietary hosted snapshot (the August Qwen3.8-Max weights are published; this refresh is API-only)
Languages
API price (in/out)
$2 / $6
Modalities
text · vision · video
Benchmark results
Bar shows position within the tracked field; marker = field best
AutomationBenchAgents50.8%#1
best: this model · 50.8%
JobBenchAgents64.0%#3
best: Claude Opus 5 · 65.7%
MMMU-ProVision82.7%#5
best: Claude Opus 4.7 · 85.5%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$2
Output / M tok
$6
API price $2/$6 · each benchmark row carries its own source badge (see methodology)