Model
Model explorer

DeepSeek-V4-Flash (Non-Think)

OPEN
DeepSeek · DeepSeek-V4 family · released Apr 24, 2026

Smaller, cheaper sibling MoE released alongside V4-Pro on the same day; fast default mode. (fast/low-latency mode; the default tier a user gets without selecting anything).

ReasoningCodingVisionFunction callingTool useAgentic
1596.6
Elo · rank #120
Parameters
284B
Active params
13B (MoE)
Context
1M tokens
Architecture
Mixture-of-Experts (284B total / 13B active), hybrid Compressed Sparse Attention + Heavily Compressed Attention
License
MIT
Languages
API price (in/out)
$0.14 / $0.28
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
GPQA DiamondReasoning71.2%#126
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning8.1%#104
best: Claude Opus 5 · 64.7%
LiveCodeBenchCoding55.2%#114
best: DeepSeek-V4-Pro (Think Max) · 93.5%
MMLU-ProKnowledge83.0%#38
best: Claude Fable 5 · 91.5%
SWE-bench ProCoding49.1%#51
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding73.7%#49
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding49.1%#63
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA193.1 GB2× H200 141GB
LoRA604.9 GB4× B200 192GB
Full fine-tune4558.2 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $6.72 (2× H200 141GB)
API price $0.14/$0.28 · each benchmark row carries its own source badge (see methodology)