Model
Model explorer

Qwen3.7-Max

CLOSED
Alibaba · Qwen3.7 family · released May 20, 2026

Agent-first flagship with 1M-token context (up from 256K on the prior generation); Alibaba reports 1,000+ consecutive tool calls and a 35-hour autonomous kernel-optimization run in internal testing (not independently verified). Not open-weight; no HF checkpoints as of this research.

ReasoningCodingVisionFunction callingTool useAgentic
2506.5
Elo · rank #29
Parameters
Undisclosed
Active params
Undisclosed
Context
1M tokens
Architecture
MoE, exact expert configuration undisclosed; positioned as an agent-first reasoning model
License
Proprietary
Languages
API price (in/out)
$2.5 / $7.5
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Agents' Last ExamAgents31.1%#17
best: GPT-6 Astra · 59.3%
GDPval-AAAgents1280#32
best: Claude Fable 5 · 1932
GPQA DiamondReasoning92.4%#18
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning41.4%#26
best: Claude Opus 5 · 64.7%
IFBenchReasoning79.1%#11
best: MiniMax M3 · 83.0%
IFEvalReasoning94.3%#6
best: Gemma 4 26B A4B · 98.5%
LiveCodeBenchCoding91.6%#4
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge89.6%#3
best: Claude Fable 5 · 91.5%
MMLU-ReduxKnowledge95.0%#1
best: this model · 95.0%
MMMLUKnowledge90.3%#6
best: Gemini 3.1 Pro · 92.6%
SuperGPQAReasoning73.6%#1
best: this model · 73.6%
SWE-bench ProCoding60.6%#16
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding80.4%#17
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding69.7%#28
best: Gemini 3.8 Flash · 89.4%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$2.5
Output / M tok
$7.5
API price $2.5/$7.5 · each benchmark row carries its own source badge (see methodology)