Model
Model explorer

gpt-oss-120b (High)

OPEN
OpenAI · gpt-oss family · released Aug 5, 2025

OpenAI's first open-weight model since GPT-2; 117B MoE (5.1B active), Apache 2.0.

ReasoningCodingVisionFunction callingTool useAgentic
1839.3
Elo · rank #93
Parameters
117B
Active params
5.1B (MoE)
Context
128K tokens
Architecture
Mixture-of-experts transformer, 128 experts/layer, top-4 active, 36 layers
License
Apache 2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding44.4%#29
best: Claude Opus 4.5 · 89.4%
AIMEMath95.8%#11
best: GPT-5.2 · 100.0%
CodeforcesCoding2463#8
best: DeepSeek-V4-Pro (Think Max) · 3206
GPQA DiamondReasoning80.1%#86
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning14.9%#81
best: Claude Opus 5 · 64.7%
MMLUKnowledge90.0%#11
best: OpenAI o3 · 92.9%
MMMLUKnowledge81.3%#23
best: Gemini 3.1 Pro · 92.6%
SWE-bench VerifiedCoding62.4%#79
best: Claude Opus 5 · 96.0%
tau-benchAgents67.8%#10
best: Claude Opus 4.1 · 82.4%
Run it locally
VRAM @ Q4
73 GB
VRAM @ FP16
234 GB
Fits on (Q4)
M3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Does not fit on a single RTX 4090 (24GB); OpenAI specifies a single 80GB GPU (H100/MI300X) for native MXFP4 inference.
Quantizations
MXFP4 (native) · GGUF Q4_K_M · GGUF Q8_0
Fine-tune it
Permissive
QLoRA79.6 GB1× A100 80GB
LoRA249.2 GB2× H200 141GB
Full fine-tune1877.8 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $3.59 (1× A100 80GB)
API price weights · each benchmark row carries its own source badge (see methodology)