Model
Model explorer

Qwen3-235B-A22B (Thinking)

OPEN
Alibaba · Qwen3 family · released Apr 29, 2025

Gen-3 flagship MoE, thinking-mode setting; rivals DeepSeek-R1, o1, Gemini 2.5 Pro.

ReasoningCodingVisionFunction callingTool useAgentic
1516.1
Elo · rank #136
Parameters
235B
Active params
22B (MoE)
Context
128K tokens
Architecture
MoE Transformer (128 experts, 8 active, hybrid thinking mode)
License
Apache 2.0
Languages
119+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding49.8%#25
best: Claude Opus 4.5 · 89.4%
AIMEMath85.7%#67
best: GPT-5.2 · 100.0%
Arena-HardHuman preference95.6%#2
best: Qwen3-235B-A22B (Non-Thinking) · 96.1%
BFCL v3Agents70.8%#6
best: Hunyuan-A13B · 78.3%
C-EvalKnowledge89.6%#14
best: Qwen3.6-Plus · 93.3%
CodeforcesCoding2056#17
best: DeepSeek-V4-Pro (Think Max) · 3206
GPQA DiamondReasoning71.1%#129
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning12.0%#87
best: Claude Opus 5 · 64.7%
IFBenchReasoning39.0%#56
best: MiniMax M3 · 83.0%
IFEvalReasoning83.4%#78
best: Gemma 4 26B A4B · 98.5%
LiveCodeBenchCoding70.7%#70
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math98.0%#7
best: GPT-5 · 99.4%
MMLU-ProKnowledge68.2%#97
best: Claude Fable 5 · 91.5%
MMLU-ReduxKnowledge92.7%#17
best: Qwen3.7-Max · 95.0%
MMMLUKnowledge84.3%#20
best: Gemini 3.1 Pro · 92.6%
SWE-bench ProCoding21.4%#55
best: Claude Fable 5.1 · 81.2%
Run it locally
VRAM @ Q4
132 GB
VRAM @ FP16
470 GB
Fits on (Q4)
M3 Ultra 512GBB200 192GB
Exceeds single RTX 4090 24GB VRAM even at Q4 (~132GB); requires multi-GPU or CPU offload.
Quantizations
GGUF · MLX · FP8 · NVFP4
Fine-tune it
Permissive
QLoRA159.8 GB1× B200 192GB
LoRA500.6 GB4× H200 141GB
Full fine-tune3771.8 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $7.87 (1× B200 192GB)
API price weights · each benchmark row carries its own source badge (see methodology)