Model
Model explorer

Falcon Mamba 7B

OPEN
TII · Falcon Mamba family · released Aug 12, 2024

First competitive attention-free State-Space (Mamba) open model; constant memory over sequence length, beats Llama 3.1 8B and Mistral 7B on Open LLM Leaderboard v1.

ReasoningCodingVisionFunction callingTool useAgentic
148.0
Elo · rank #381
Parameters
7B
Active params
7B (dense)
Context
8K tokens
Architecture
Pure Mamba State-Space Model, 7B, attention-free with extra RMS norm layers
License
TII Falcon-Mamba License 2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
ARC-ChallengeReasoning62.0%#80
best: Llama 3.1 405B · 96.9%
BIG-Bench HardReasoning19.9%#138
best: ERNIE 4.5 300B-A47B · 94.3%
GSM8KMath52.5%#129
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning80.8%#73
best: Claude 3 Opus · 95.4%
IFEvalReasoning33.4%#141
best: Gemma 4 26B A4B · 98.5%
MMLU-ProKnowledge14.5%#176
best: Claude Fable 5 · 91.5%
MMLUKnowledge62.1%#196
best: OpenAI o3 · 92.9%
TruthfulQAKnowledge53.4%#43
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
WinoGrandeReasoning73.6%#76
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
5 GB
VRAM @ FP16
15 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF · bitsandbytes 4-bit
Fine-tune it
Conditional / custom
QLoRA6.4 GB1× RTX 3060 12GB
LoRA16.6 GB1× RTX 3090 24GB
Full fine-tune114.0 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.10 (1× RTX 3060 12GB)
Falcon Mamba family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)