Model
Model explorer
Command R
OPENCohere · Command R family · released Mar 11, 2024
35B RAG/tool-use model with 128K context; first open-weights (CC-BY-NC) Command release, later superseded by the 08-2024 refresh.
ReasoningCodingVisionFunction callingTool useAgentic
385.9
Elo · rank #343
Parameters
35B
Active params
35B (dense)
Context
128K tokens
Architecture
Dense transformer (35B), optimized for RAG and multi-step tool use
License
CC-BY-NC-4.0 (+ Acceptable Use Policy)
Languages
10+
API price (in/out)
$0.5 / $1.5
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
best: Llama 3.1 405B · 96.9%
best: ERNIE 4.5 300B-A47B · 94.3%
best: GPT-6 Astra · 96.0%
best: Llama 3.1 405B · 96.8%
best: Claude 3 Opus · 95.4%
best: Gemma 4 26B A4B · 98.5%
best: Llama 3.1 70B · 77.1%
best: Claude Fable 5 · 91.5%
best: OpenAI o3 · 92.9%
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
21.6 GB
VRAM @ FP16
70 GB
Fits on (Q4)
RTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF Q4
Fine-tune it
Research onlyQLoRA24.1 GB1× RTX 5090 32GB
LoRA74.8 GB1× A100 80GB
Full fine-tune562.0 GB4× H200 141GB
QLoRA SFT on ~10k samples ≈ $17.07 (1× RTX 5090 32GB)
API price $0.5/$1.5 · each benchmark row carries its own source badge (see methodology)