Model
Model explorer

Nemotron 3 Super

OPEN
NVIDIA · Nemotron 3 family · released Mar 11, 2026

Mid-size member of the Nemotron 3 family, released between Nano Omni (April) and the Ultra 550B already in the corpus (dated 2026-06-04). Fills the size gap between Nemotron-H 56B and Nemotron 3 Ultra 550B.

ReasoningCodingVisionFunction callingTool useAgentic
1805.3
Elo · rank #99
Parameters
120B
Active params
12B (MoE)
Context
1M tokens
Architecture
Hybrid Mamba-Transformer MoE with LatentMoE and multi-token prediction, 120B total / 12B active
License
NVIDIA Open Model License
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath90.2%#46
best: GPT-5.2 · 100.0%
Arena-Hard v2Human preference73.9%#4
best: Qwen3-Max · 86.1%
BIRD-SQLCoding41.8%#8
best: Command A · 59.5%
BrowseCompAgents31.3%#40
best: Kimi K3 · 91.2%
GDPval-AAAgents700#45
best: Claude Fable 5 · 1932
GPQA DiamondReasoning79.2%#91
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning18.3%#72
best: Claude Opus 5 · 64.7%
IFBenchReasoning72.6%#26
best: MiniMax M3 · 83.0%
LiveCodeBenchCoding81.2%#48
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge83.7%#34
best: Claude Fable 5 · 91.5%
SWE-bench VerifiedCoding60.5%#82
best: Claude Opus 5 · 96.0%
τ²-Bench TelecomAgents64.4%#31
best: Claude Opus 4.6 · 99.3%
Terminal-Bench 2.0Coding31.0%#70
best: Gemini 3.8 Flash · 89.4%
Terminal-BenchCoding25.8%#13
best: Claude Opus 4.5 (High) · 59.3%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Conditional / custom
QLoRA81.6 GB1× H200 141GB
LoRA255.6 GB2× H200 141GB
Full fine-tune1926.0 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $62.06 (1× H200 141GB)
API price weights · each benchmark row carries its own source badge (see methodology)