Model
Model explorer

Qwen3-14B (Thinking)

OPEN
Alibaba · Qwen3 family · released Apr 29, 2025

Mid-size Gen-3 dense model, thinking-mode setting; Apache 2.0.

ReasoningCodingVisionFunction callingTool useAgentic
1382.9
Elo · rank #156
Parameters
14.8B
Active params
14.8B (dense)
Context
128K tokens
Architecture
Dense Transformer (GQA, QK-Norm, hybrid thinking mode)
License
Apache 2.0
Languages
119+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath79.3%#83
best: GPT-5.2 · 100.0%
Arena-HardHuman preference91.7%#9
best: Qwen3-235B-A22B (Non-Thinking) · 96.1%
BFCL v3Agents70.4%#8
best: Hunyuan-A13B · 78.3%
C-EvalKnowledge86.2%#22
best: Qwen3.6-Plus · 93.3%
CodeforcesCoding1766#28
best: DeepSeek-V4-Pro (Think Max) · 3206
GDPval-AAAgents500#49
best: Claude Fable 5 · 1932
GPQA DiamondReasoning64.0%#159
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning4.0%#139
best: Claude Opus 5 · 64.7%
IFBenchReasoning41.0%#51
best: MiniMax M3 · 83.0%
IFEvalReasoning85.4%#57
best: Gemma 4 26B A4B · 98.5%
LiveCodeBenchCoding63.5%#92
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math96.8%#22
best: GPT-5 · 99.4%
MMLU-ReduxKnowledge88.6%#27
best: Qwen3.7-Max · 95.0%
MMMLUKnowledge77.9%#32
best: Gemini 3.1 Pro · 92.6%
Terminal-Bench 2.0Coding5.0%#76
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
8.4 GB
VRAM @ FP16
29.6 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
~60 tok/s on RTX 4090 (Q4, llama.cpp)
Quantizations
GGUF Q4 · GGUF · FP8
Fine-tune it
Permissive
QLoRA11.3 GB1× RTX 3060 12GB
LoRA32.8 GB1× A100 80GB
Full fine-tune238.8 GB2× H200 141GB
QLoRA SFT on ~10k samples ≈ $8.66 (1× RTX 3060 12GB)
API price weights · each benchmark row carries its own source badge (see methodology)