Model
Model explorer

DeepSeek-V4-Flash-0731

OPEN
DeepSeek · DeepSeek-V4 family · released Jul 31, 2026

Retrained checkpoint of DeepSeek-V4-Flash published July 31, 2026, same 284B/13B architecture and API pricing as the April siblings. DeepSeek's own model card only reruns the launch-day agentic/coding table (Terminal-Bench 2.1, NL2Repo, CyberGym, DeepSWE) with an unreleased internal harness ('DeepSeek Harness'); general reasoning/knowledge benchmarks were not rerun and still carry April Preview values, so this row is coding/agentic-only evidence, not a full refresh.

ReasoningCodingVisionFunction callingTool useAgentic
2777.6
Elo · unrated
Parameters
284B
Active params
13B (MoE)
Context
1M tokens
Architecture
Mixture-of-Experts (284B total / 13B active), hybrid Compressed Sparse Attention + Heavily Compressed Attention — retrained checkpoint of DeepSeek-V4-Flash on the same architecture
License
MIT
Languages
API price (in/out)
$0.14 / $0.28
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
CyberGymCoding76.7%#6
best: GPT-5.6 Sol · 84.5%
DeepSWECoding54.4%#5
best: GPT-5.6 Sol · 72.7%
Terminal-Bench 2.0Coding82.7%#9
best: GPT-5.6 Sol · 88.8%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA193.1 GB2× H200 141GB
LoRA604.9 GB4× B200 192GB
Full fine-tune4558.2 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $6.72 (2× H200 141GB)
API price $0.14/$0.28 · each benchmark row carries its own source badge (see methodology)