Model
Model explorer

CogVLM-17B

OPEN
Zhipu AI / THUDM (Tsinghua) · CogVLM family · released Oct 1, 2023

Open VLM (10B vision expert + 7B Vicuna-based LM) with SOTA-level results on 10 cross-modal benchmarks at 490x490 resolution.

ReasoningCodingVisionFunction callingTool useAgentic
465.4
Elo · unrated
Parameters
17B
Active params
17B (dense)
Context
4K tokens
Architecture
Dense dual-tower vision-language Transformer (trainable visual expert modules)
License
Apache-2.0 (code); CogVLM Model License for weights (non-commercial use / commercial by agreement)
Languages
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
COCO CaptionsVision148.7#1
best: this model · 148.7
Flickr30kVision94.9#1
best: this model · 94.9
MathVistaVision34.5%#60
best: Seed 2.1 Pro · 90.7%
MMBenchVision77.6%#3
best: Phi-3.5-vision (4.2B) · 81.9%
MMMUVision41.1%#104
best: Claude Fable 5 · 89.3%
OK-VQAVision64.8%#1
best: this model · 64.8%
SEED-BenchVision72.5%#4
best: NAVER HyperCLOVA X SEED Think 32B · 77.9%
TextVQAVision70.4%#22
best: Molmo 2 8B · 85.7%
VQAv2Vision82.3%#4
best: Reka Edge (2026) · 88.4%
Run it locally
VRAM @ Q4
11 GB
VRAM @ FP16
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
INT4 · INT8
Fine-tune it
Research only
QLoRA12.7 GB1× RTX 4070 Ti 16GB
LoRA37.4 GB1× A100 80GB
Full fine-tune274.0 GB2× H200 141GB
QLoRA SFT on ~10k samples ≈ $6.22 (1× RTX 4070 Ti 16GB)
CogVLM family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)