Model
Model explorer

Nemotron 3 Ultra 550B A55B

OPEN
NVIDIA · Nemotron 3 family · released Jun 4, 2026

NVIDIA's frontier open MoE with 1M-token context; hybrid LatentMoE+Mamba-2 flagship of the Nemotron 3 family, tuned for long-running agents.

ReasoningCodingVisionFunction callingTool useAgentic
2155.8
Elo · rank #65
Parameters
550B
Active params
55B (MoE)
Context
1M tokens
Architecture
Hybrid LatentMoE + Mamba-2 (550B total / 55B active, MTP layers, NVFP4-native pretraining)
License
OpenMDW License Agreement v1.1 (Linux Foundation Model Openness Framework)
Languages
10+
API price (in/out)
$0.5 / $2.2
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
BrowseCompAgents44.4%#36
best: Kimi K3 · 91.2%
GDPvalAgents46.7%#6
best: Seed 2.1 Pro · 87.9%
GPQA DiamondReasoning87.0%#53
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning26.7%#57
best: Claude Opus 5 · 64.7%
IFBenchReasoning81.7%#4
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding89.0%#11
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge86.8%#16
best: Claude Fable 5 · 91.5%
PinchBenchAgents90.0%#2
best: Trinity-Large-Thinking · 91.9%
SWE-bench VerifiedCoding70.7%#62
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding56.4%#51
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
352 GB
VRAM @ FP16
1121 GB
Fits on (Q4)
M3 Ultra 512GB
Needs 8x GB200/B200/H200-class GPUs natively; community GGUF still needs ~300GB+ combined RAM/VRAM, far beyond a single RTX 4090.
Quantizations
NVFP4 · GGUF
Fine-tune it
Permissive
QLoRA374.0 GB2× B200 192GB
LoRA1171.5 GB8× B200 192GB
Full fine-tune8827.5 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $197 (2× B200 192GB)
API price $0.5/$2.2 · each benchmark row carries its own source badge (see methodology)