Model
Model explorer

Gemma 4

OPEN
Google · Gemma 4 family · released Apr 2, 2026

Most capable open Gemma yet; 31B dense flagship ranks #3 among open models on the Arena AI text leaderboard.

ReasoningCodingVisionFunction callingTool useAgentic
1976.5
Elo · rank #78
Parameters
30.7B
Active params
30.7B (dense)
Context
256K tokens
Architecture
Dense Transformer, hybrid local/global attention (31B flagship variant)
License
Apache-2.0
Languages
140+
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath89.2%#50
best: GPT-5.2 · 100.0%
Arena EloHuman preference1452#7
best: Claude Fable 5 · 1505
CodeforcesCoding2150#11
best: DeepSeek-V4-Pro (Think Max) · 3206
GPQA DiamondReasoning84.3%#65
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning19.5%#70
best: Claude Opus 5 · 64.7%
LiveCodeBenchCoding80.0%#50
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MathVisionVision85.6%#9
best: Seed 2.1 Pro · 92.6%
MMLU-ProKnowledge85.2%#24
best: Claude Fable 5 · 91.5%
MMMLUKnowledge88.4%#13
best: Gemini 3.1 Pro · 92.6%
MMMU-ProVision76.9%#23
best: Claude Opus 4.7 · 85.5%
Run it locally
VRAM @ Q4
20 GB
VRAM @ FP16
64 GB
Fits on (Q4)
RTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Third-party RTX 4090 tok/s figures found in the wild referenced non-existent Gemma 4 sizes (7B/9B/27B) and were discarded as unreliable.
Quantizations
GGUF Q4 · GPTQ · MLX
Fine-tune it
Permissive
QLoRA21.3 GB1× RTX 3090 24GB
LoRA65.9 GB1× A100 80GB
Full fine-tune493.2 GB4× H200 141GB
QLoRA SFT on ~10k samples ≈ $15.81 (1× RTX 3090 24GB)
Gemma 4 family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)