Model
Model explorer

Falcon3-10B

OPEN
TII · Falcon 3 family · released Dec 17, 2024

Depth-upscaled flagship of Falcon 3; topped Hugging Face's sub-13B leaderboard at release.

ReasoningCodingVisionFunction callingTool useAgentic
415.4
Elo · rank #335
Parameters
10B
Active params
10B (dense)
Context
32K tokens
Architecture
Dense decoder-only transformer (depth-upscaled from Falcon3-7B, GQA, SwiGLU)
License
TII Falcon-LLM License 2.0
Languages
4+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
AlpacaEvalHuman preference24.7%#34
best: Tülu 2+DPO 70B · 95.1%
ARC-ChallengeReasoning66.2%#61
best: Llama 3.1 405B · 96.9%
BIG-Bench HardReasoning44.8%#105
best: ERNIE 4.5 300B-A47B · 94.3%
BFCLAgents90.5%#2
best: SmolLM3 3B (No Thinking) · 92.3%
GPQA DiamondReasoning10.5%#292
best: GPT-6 Astra · 96.0%
GSM8KMath84.9%#68
best: Llama 3.1 405B · 96.8%
HumanEval+Coding74.7%#18
best: Mistral Small 3.2 (24B) · 92.9%
IFEvalReasoning78.2%#96
best: Gemma 4 26B A4B · 98.5%
MATH-500Math25.9%#170
best: GPT-5 · 99.4%
MMLU-ProKnowledge38.1%#155
best: Claude Fable 5 · 91.5%
MMLUKnowledge73.9%#127
best: OpenAI o3 · 92.9%
MT-BenchHuman preference8.2#40
best: Hunyuan-Large (A52B) · 9.4
OpenBookQAReasoning48.2%#30
best: Claude 1 · 90.8%
PIQAReasoning78.4%#52
best: GPT-4o mini · 93.1%
WinoGrandeReasoning71.0%#96
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
6 GB
VRAM @ FP16
22.94 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
GGUF · AWQ · GPTQ · MLX
Fine-tune it
Conditional / custom
QLoRA8.3 GB1× RTX 3060 12GB
LoRA22.8 GB1× RTX 3090 24GB
Full fine-tune162.0 GB1× B200 192GB
QLoRA SFT on ~10k samples ≈ $5.85 (1× RTX 3060 12GB)
Falcon 3 family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)