Model
Model explorer

OLMo 2 13B

OPEN
Allen Institute for AI (Ai2) · OLMo 2 family · released Nov 26, 2024

Flagship of Ai2's fully-open Nov 2024 OLMo 2 launch; Instruct variant beats Qwen 2.5 14B Instruct on average score

ReasoningCodingVisionFunction callingTool useAgentic
575.7
Elo · rank #299
Parameters
13B
Active params
13B (dense)
Context
4K tokens
Architecture
Dense Transformer (Olmo2 arch: 40 layers, QK-norm, reordered post-norm)
License
Apache-2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AGIEvalReasoning54.2%#11
best: OLMo 3-Think 32B · 88.2%
AlpacaEvalHuman preference39.5%#27
best: Tülu 2+DPO 70B · 95.1%
ARC-ChallengeReasoning83.5%#36
best: Llama 3.1 405B · 96.9%
BIG-Bench HardReasoning58.8%#81
best: ERNIE 4.5 300B-A47B · 94.3%
DROPReasoning71.5#43
best: Hunyuan-T1 · 93.1
GSM8KMath87.4%#54
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning86.4%#27
best: Claude 3 Opus · 95.4%
IFEvalReasoning82.6%#85
best: Gemma 4 26B A4B · 98.5%
MATH-500Math39.2%#148
best: GPT-5 · 99.4%
MMLU-ProKnowledge35.1%#159
best: Claude Fable 5 · 91.5%
MMLUKnowledge68.5%#156
best: OpenAI o3 · 92.9%
TriviaQAKnowledge81.9%#11
best: Sarvam-1 (2B) · 90.6%
TruthfulQAKnowledge64.3%#17
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
WinoGrandeReasoning81.5%#27
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
8.4 GB
VRAM @ FP16
27.4 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF
Fine-tune it
Permissive
QLoRA10.2 GB1× RTX 3060 12GB
LoRA29.0 GB1× RTX 5090 32GB
Full fine-tune210.0 GB2× H200 141GB
QLoRA SFT on ~10k samples ≈ $7.61 (1× RTX 3060 12GB)
API price weights · each benchmark row carries its own source badge (see methodology)