Model
Model explorer

OpenAI o3

CLOSED
OpenAI · o3 family · released Apr 16, 2025

Flagship reasoning model with native tool use; major gains on math, coding, and ARC-AGI. Price reflects OpenAI's June 2025 80% cut from the original $10/$40 launch price.

ReasoningCodingVisionFunction callingTool useAgentic
1896.3
Elo · rank #84
Parameters
Undisclosed
Active params
Context
200K tokens
Architecture
Dense transformer reasoning model (exact size undisclosed)
License
Proprietary (OpenAI API Terms)
Languages
API price (in/out)
$2 / $8
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
Aider PolyglotCoding79.6%#6
best: Claude Opus 4.5 · 89.4%
AIMEMath91.6%#41
best: GPT-5.2 · 100.0%
ARC-AGI-1Reasoning53.0%#15
best: GPT-6 Astra · 98.5%
ARC-AGI-2Reasoning2.9%#28
best: GPT-6 Astra · 95.0%
BrowseCompAgents49.7%#33
best: Kimi K3 · 91.2%
CharXivVision78.6%#19
best: Qwen3.8-Flash-Next · 90.6%
CodeforcesCoding2706#7
best: DeepSeek-V4-Pro (Think Max) · 3206
best: GPT-5.6 Sol · 89.0%
GAIAAgents28.5%#12
best: MiniMax-M2 · 75.7%
GPQA DiamondReasoning83.3%#72
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning24.9%#62
best: Claude Opus 5 · 64.7%
HumanEvalCoding87.4%#33
best: Claude Opus 4.5 · 99.4%
LiveCodeBenchCoding75.8%#58
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MATH-500Math97.8%#12
best: GPT-5 · 99.4%
MathVistaVision86.8%#6
best: Seed 2.1 Pro · 90.7%
MMLUKnowledge92.9%#1
best: this model · 92.9%
MMMUVision82.9%#15
best: Claude Fable 5 · 89.3%
SWE-bench VerifiedCoding69.1%#66
best: Claude Opus 5 · 96.0%
Run it locally
Closed weights — available via API only. No local deployment.
Input / M tok
$2
Output / M tok
$8
API price $2/$8 · each benchmark row carries its own source badge (see methodology)