Model
Model explorer

DeepSeek-VL 7B

OPEN
DeepSeek · DeepSeek-VL family · released Mar 8, 2024

First DeepSeek VLM (1.3B/7B); hybrid SigLIP+SAM vision encoder trained jointly with the LLM to preserve language ability.

ReasoningCodingVisionFunction callingTool useAgentic
-29.9
Elo · rank #419
Parameters
7B
Active params
7B (dense)
Context
4K tokens
Architecture
Dense 7B transformer (DeepSeek-LLM-7B base) + SigLIP-L/SAM-B hybrid vision encoder (up to 1024x1024)
License
DeepSeek Model License (code: MIT)
Languages
2+
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AGIEvalReasoning27.8%#39
best: OLMo 3-Think 32B · 88.2%
GSM8KMath55.0%#125
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning68.4%#131
best: Claude 3 Opus · 95.4%
MathVistaVision36.1%#59
best: Seed 2.1 Pro · 90.7%
MBPPCoding35.2%#91
best: Llama-3.3-Nemotron-Super-49B v1 (Reasoning On) · 91.3%
MMBench (Chinese)Vision72.8%#9
best: ERNIE 4.5 VL 424B-A47B · 90.9%
MMBenchVision73.2%#5
best: Phi-3.5-vision (4.2B) · 81.9%
MMLUKnowledge52.4%#230
best: OpenAI o3 · 92.9%
MMMUVision36.6%#105
best: Claude Fable 5 · 89.3%
OCRBenchVision456#13
best: InternVL3-78B · 906
SEED-BenchVision70.4%#5
best: NAVER HyperCLOVA X SEED Think 32B · 77.9%
Run it locally
VRAM @ Q4
5 GB
VRAM @ FP16
14 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
No public RTX 4090 llama.cpp benchmark found for the full vision-language model
Quantizations
MLX
Fine-tune it
Conditional / custom
QLoRA6.4 GB1× RTX 3060 12GB
LoRA16.6 GB1× RTX 3090 24GB
Full fine-tune114.0 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.10 (1× RTX 3060 12GB)
DeepSeek-VL family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)