Model
Model explorer

Phi-3-medium (14B)

OPEN
Microsoft · Phi-3 family · released May 21, 2024

Largest dense Phi-3 checkpoint (~78% MMLU); top of the Phi-3 quality-cost curve, 4K/128K context, MIT-licensed.

ReasoningCodingVisionFunction callingTool useAgentic
714.8
Elo · rank #276
Parameters
14B
Active params
14B (dense)
Context
128K tokens
Architecture
Dense decoder-only transformer, 4K/128K context variants
License
MIT
Languages
API price (in/out)
$0.17 / $0.68
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AGIEvalReasoning50.2%#18
best: OLMo 3-Think 32B · 88.2%
ANLIReasoning55.8%#2
best: Phi-3-small (7B) · 58.1%
ARC-ChallengeReasoning91.6%#21
best: Llama 3.1 405B · 96.9%
ARC-EasyReasoning97.7%#1
best: this model · 97.7%
BIG-Bench HardReasoning81.4%#30
best: ERNIE 4.5 300B-A47B · 94.3%
GSM8KMath91.0%#36
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning82.4%#66
best: Claude 3 Opus · 95.4%
HumanEvalCoding62.2%#99
best: Claude Opus 4.5 · 99.4%
LogicKorHuman preference4.64#28
best: GPT-4o · 9.33
MBPPCoding75.2%#24
best: Llama-3.3-Nemotron-Super-49B v1 (Reasoning On) · 91.3%
MMLU (EU-21 languages)Knowledge57.4%#6
best: Llama 3.1 70B · 77.1%
MMLUKnowledge78.0%#107
best: OpenAI o3 · 92.9%
OpenBookQAReasoning87.4%#4
best: Claude 1 · 90.8%
PIQAReasoning87.9%#5
best: GPT-4o mini · 93.1%
Social IQaReasoning80.2%#2
best: Apple DCLM-Baseline 7B · 82.9%
TriviaQAKnowledge73.9%#17
best: Sarvam-1 (2B) · 90.6%
TruthfulQAKnowledge75.1%#2
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
WinoGrandeReasoning81.5%#28
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
9 GB
VRAM @ FP16
28 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF · GPTQ · AWQ · MLX
Fine-tune it
Permissive
QLoRA10.8 GB1× RTX 3060 12GB
LoRA31.1 GB1× RTX 5090 32GB
Full fine-tune226.0 GB2× H200 141GB
QLoRA SFT on ~10k samples ≈ $8.19 (1× RTX 3060 12GB)
API price $0.17/$0.68 · each benchmark row carries its own source badge (see methodology)