Model
Model explorer
MiniCPM-o 2.6
OPENOpenBMB · MiniCPM-o family · released Jan 13, 2025
8B any-to-any omni model (vision+speech+language) with real-time streaming, matching GPT-4o-202405 on OpenCompass (70.2 avg).
ReasoningCodingVisionFunction callingTool useAgentic
988.9
Elo · unrated
Parameters
8B
Active params
8B (dense)
Context
32K tokens
Architecture
SigLip-400M + Whisper-medium-300M + ChatTTS-200M + Qwen2.5-7B, end-to-end omni model
License
MiniCPM Model License (custom; free for research and, after registration, commercial use) + Apache-2.0 code
Languages
2+
API price (in/out)
No hosted API
Modalities
text · vision · video · audio
Benchmark results
Bar shows position within the tracked field; marker = field best
best: Molmo 72B · 96.3%
best: MiniMax-VL-01 · 91.7%
best: Qwen2-VL-72B · 96.5%
best: Seed 2.1 Pro · 92.6%
best: Seed 2.1 Pro · 90.7%
best: Qwen3.5-397B-A17B · 93.7%
best: Claude Fable 5 · 89.3%
best: InternVL3-78B · 906
best: Gemini 1.5 Pro · 70.3%
best: Molmo 2 8B · 85.7%
best: Seed 2.1 Pro · 89.2%
Run it locally
VRAM @ Q4
7 GB
VRAM @ FP16
18 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF Q4_K_M · BNB INT4
Fine-tune it
Conditional / customQLoRA7.0 GB1× RTX 3060 12GB
LoRA18.6 GB1× RTX 3090 24GB
Full fine-tune130.0 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.68 (1× RTX 3060 12GB)
API price weights · each benchmark row carries its own source badge (see methodology)