Model
Model explorer

Qwen3-32B (Thinking)

OPEN
Alibaba · Qwen3 family · released Apr 29, 2025

Largest Qwen3 dense model; thinking mode rivals OpenAI o3-mini (medium) on reasoning benchmarks.

ReasoningCodingVisionFunction callingTool useAgentic
1484.0
Elo · rank #140
Parameters
32.8B
Active params
32.8B (dense)
Context
128K tokens
Architecture
Dense transformer, GQA + SwiGLU + QK-Norm
License
Apache 2.0
Languages
119+
API price (in/out)
$0.287 / $1.147
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath81.4%#72
best: GPT-5.2 · 100.0%
Arena-HardHuman preference93.8%#3
best: Qwen3-235B-A22B (Non-Thinking) · 96.1%
Arena EloHuman preference1347#30
best: Claude Fable 5 · 1505
BFCL v3Agents70.3%#9
best: Hunyuan-A13B · 78.3%
C-EvalKnowledge87.3%#19
best: Qwen3.6-Plus · 93.3%
CodeforcesCoding1977#22
best: DeepSeek-V4-Pro (Think Max) · 3206
GPQA DiamondReasoning68.4%#136
best: GPT-6 Astra · 96.0%
IFEvalReasoning85.0%#60
best: Gemma 4 26B A4B · 98.5%
LiveCodeBenchCoding65.7%#84
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math97.2%#19
best: GPT-5 · 99.4%
MMLU-ReduxKnowledge90.9%#20
best: Qwen3.7-Max · 95.0%
MMMLUKnowledge80.6%#25
best: Gemini 3.1 Pro · 92.6%
Run it locally
VRAM @ Q4
19.9 GB
VRAM @ FP16
65.6 GB
Fits on (Q4)
RTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF · AWQ · GPTQ · MLX
Fine-tune it
Permissive
QLoRA22.7 GB1× RTX 3090 24GB
LoRA70.2 GB1× A100 80GB
Full fine-tune526.8 GB4× H200 141GB
QLoRA SFT on ~10k samples ≈ $16.89 (1× RTX 3090 24GB)
API price $0.287/$1.147 · each benchmark row carries its own source badge (see methodology)