Model
Model explorer

Falcon-7B

OPEN
TII · Falcon family · released May 25, 2023

Apache-2.0 causal decoder trained on 1.5T RefinedWeb tokens; the small, widely-adopted entry point of the original Falcon line.

ReasoningCodingVisionFunction callingTool useAgentic
-307.6
Elo · rank #448
Parameters
7B
Active params
7B (dense)
Context
2K tokens
Architecture
Causal decoder-only Transformer, 7B (32 layers, multiquery attention, RoPE, FlashAttention)
License
Apache 2.0
Languages
11+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
ARC-ChallengeReasoning44.5%#137
best: Llama 3.1 405B · 96.9%
ARC-EasyReasoning73.6%#38
best: Phi-3-medium (14B) · 97.7%
GSM8KMath4.6%#186
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning76.3%#106
best: Claude 3 Opus · 95.4%
LAMBADAReasoning74.9%#10
best: GPT-3 175B · 86.4%
MMLUKnowledge28.0%#275
best: OpenAI o3 · 92.9%
OpenBookQAReasoning44.6%#38
best: Claude 1 · 90.8%
PIQAReasoning80.3%#40
best: GPT-4o mini · 93.1%
TruthfulQAKnowledge34.3%#98
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
WinoGrandeReasoning67.2%#110
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
6 GB
VRAM @ FP16
16.2 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF · GPTQ · bitsandbytes 4-bit
Fine-tune it
Permissive
QLoRA6.4 GB1× RTX 3060 12GB
LoRA16.6 GB1× RTX 3090 24GB
Full fine-tune114.0 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.10 (1× RTX 3060 12GB)
API price weights · each benchmark row carries its own source badge (see methodology)