Model
Model explorer
DeepSeek-Prover-V2 671B
OPENDeepSeek · DeepSeek-Prover family · released Apr 30, 2025
Formal theorem-proving model for Lean 4 built on V3-Base via recursive proof search and subgoal decomposition; state-of-the-art neural theorem prover.
ReasoningCodingVisionFunction callingTool useAgentic
2140.6
Elo · unrated
Parameters
671B
Active params
37B (MoE)
Context
128K tokens
Architecture
DeepSeekMoE (same architecture as DeepSeek-V3): 256 routed + 1 shared expert, top-8 routing, MLA attention, initialized from DeepSeek-V3-Base
License
MIT
Languages
—
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
best: this model · 88.9%
best: this model · 90.6%
best: this model · 7.1%
Run it locally
VRAM @ Q4
404 GB
VRAM @ FP16
1342 GB
Fits on (Q4)
Multi-node cluster required
671B MoE far exceeds a single 24GB 4090; only Unsloth Dynamic GGUF quants for CPU/multi-GPU offload are available.
Quantizations
GGUF Q4_K_M (Unsloth Dynamic)
Fine-tune it
PermissiveQLoRA456.3 GB4× H200 141GB
LoRA1429.2 GB8× B200 192GB
Full fine-tune10769.5 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $19.14 (4× H200 141GB)
API price weights · each benchmark row carries its own source badge (see methodology)