Model
Model explorer

BLOOM-7B1

OPEN
BigScience · BLOOM family · released Jul 11, 2022

7.07B multilingual BLOOM checkpoint, the practical size for downstream fine-tuning and local deployment.

ReasoningCodingVisionFunction callingTool useAgentic
-490.3
Elo · rank #463
Parameters
7.069B
Active params
7.069B (dense)
Context
2K tokens
Architecture
Dense decoder-only transformer (ALiBi positional encoding, GPT-2-style, GeLU)
License
BigScience RAIL License v1.0
Languages
57+
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
ARC-ChallengeReasoning41.1%#146
best: Llama 3.1 405B · 96.9%
BIG-Bench HardReasoning30.9%#130
best: ERNIE 4.5 300B-A47B · 94.3%
GPQA DiamondReasoning26.4%#280
best: GPT-6 Astra · 96.0%
GSM8KMath1.4%#196
best: Llama 3.1 405B · 96.8%
HellaSwagReasoning62.0%#139
best: Claude 3 Opus · 95.4%
HumanEvalCoding8.1%#186
best: Claude Opus 4.5 · 99.4%
IFEvalReasoning9.1%#157
best: Gemma 4 26B A4B · 98.5%
MATH-500Math0.5%#209
best: GPT-5 · 99.4%
MMLU (EU-21 languages)Knowledge25.6%#17
best: Llama 3.1 70B · 77.1%
MMLU-ProKnowledge11.1%#183
best: Claude Fable 5 · 91.5%
MMLUKnowledge26.3%#282
best: OpenAI o3 · 92.9%
TruthfulQAKnowledge38.9%#85
best: Phi-3.5-MoE (16x3.8B, 6.6B active) · 77.5%
WinoGrandeReasoning65.4%#119
best: PaLM 2 · 90.9%
Run it locally
VRAM @ Q4
5 GB
VRAM @ FP16
14 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
INT8 (bitsandbytes) · GGUF
Fine-tune it
Conditional / custom
QLoRA6.5 GB1× RTX 3060 12GB
LoRA16.7 GB1× RTX 3090 24GB
Full fine-tune115.1 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.14 (1× RTX 3060 12GB)
BLOOM family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)