Model
Model explorer

Mistral Large 4

OPEN
Mistral AI · Mistral Large family · released Sep 9, 2026

Mistral's September 9, 2026 flagship — a modest 685B/43B scale-up over Large 3 (675B/41B) that stays Apache 2.0 and doubles context to 512K. The honest read is unchanged from the Large line's recent history: competent open generalist, visibly behind the reasoning-specialized crowd — HLE 9.4 (independent), Terminal-Bench 2.1 at 21.6, SWE-bench 52.7. As with Large 3, the filings here are the independent/third-party measurements Mistral's own docs now link out to rather than vendor table claims; capabilities.reasoning stays false (no thinking mode).

ReasoningCodingVisionFunction callingTool useAgentic
1412.6
Elo · rank #161
Parameters
685B
Active params
43B (MoE)
Context
512K tokens
Architecture
685B-total/43B-active sparse MoE, Apache 2.0 open weights; 512K-token context
License
Apache 2.0
Languages
—
API price (in/out)
$0.6 / $1.8
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath52.6%#119
best: GPT-5.2 · 100.0%
GPQA DiamondReasoning72.8%#133
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning9.4%#110
best: Claude Opus 5 · 64.7%
HumanEvalCoding91.3%#16
best: Claude Opus 4.5 · 99.4%
IFBenchReasoning44.2%#52
best: Inkling-Medium · 85.3%
LiveCodeBenchCoding62.9%#100
best: DeepSeek-V4.1 (Think Max) · 94.2%
MATH-500Math94.2%#38
best: GPT-5 · 99.4%
MMLU-ProKnowledge78.4%#68
best: Claude Fable 5 · 91.5%
MMMU-ProVision62.7%#51
best: Claude Opus 4.7 · 85.5%
MMMUVision71.8%#50
best: Claude Fable 5 · 89.3%
SWE-bench VerifiedCoding52.7%#100
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding21.6%#85
best: Gemini 3.8 Flash · 89.4%
Video-MME ProVision71.4%#4
best: Seed 2.2 Pro · 90.4%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
—
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
Permissive
QLoRA465.8 GB4× H200 141GB
LoRA1459.0 GB8× B200 192GB
Full fine-tune10994.3 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $22.24 (4× H200 141GB)
API price $0.6/$1.8 · each benchmark row carries its own source badge (see methodology)