Model
Model explorer

GPT-5.4

CLOSED
OpenAI · GPT-5.4 family · released Mar 5, 2026

Flagship/'Thinking' tier of the GPT-5.4 family, confirmed real via TechCrunch and The Decoder launch coverage. GPT-5.4 Pro ($30/$180 per M, ARC-AGI-2 83.3%, GDPval 82%) and cheaper GPT-5.4 mini/nano also shipped the same month but are out of scope for this pass (not requested).

ReasoningCodingVisionFunction callingTool useAgentic
2445.3
Elo · rank #37
Parameters
Undisclosed
Active params
Context
1.05M tokens
Architecture
Undisclosed architecture (presumed dense transformer, reasoning-tuned); first general-purpose OpenAI model with native computer-use; ships in Thinking/Pro/mini/nano tiers
License
Proprietary (OpenAI API Terms)
Languages
API price (in/out)
$2.5 / $15
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
Agents' Last ExamAgents37.3%#14
best: GPT-6 Astra · 59.3%
AIMEMath95.3%#14
best: GPT-5.2 · 100.0%
ARC-AGI-1Reasoning93.7%#7
best: GPT-6 Astra · 98.5%
ARC-AGI-2Reasoning73.3%#9
best: GPT-6 Astra · 95.0%
BrowseCompAgents82.7%#17
best: Kimi K3 · 91.2%
best: GPT-5.6 Sol · 89.0%
best: GPT-5.6 Sol · 83.0%
GDPvalAgents83.0%#3
best: Seed 2.1 Pro · 87.9%
GPQA DiamondReasoning92.8%#15
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning39.8%#31
best: Claude Opus 5 · 64.7%
LiveCodeBenchCoding84.1%#29
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMMUVision87.5%#2
best: Claude Fable 5 · 89.3%
OSWorld-VerifiedAgents75.0%#11
best: Claude Fable 5 · 85.0%
SWE-bench ProCoding57.7%#25
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding76.9%#35
best: Claude Opus 5 · 96.0%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$2.5
Output / M tok
$15
API price $2.5/$15 · each benchmark row carries its own source badge (see methodology)