Model
Model explorer

Muse Glimmer 8B

OPEN
Meta · Muse Glimmer family · released Sep 15, 2026

Meta's September 15, 2026 open-weights release — the 8.2B consumer-GPU tier of the Muse Glimmer line that debuted at 30B in August. The point of the release is footprint: 5.8GB at Q4_K_M (16.5GB at BF16) with ~95 tok/s on an RTX 4090, i.e. a genuinely local multimodal model. The capability story is proportionate — most scores land at 70-80% of the 30B sibling, with agentic rows (MCP Atlas, OSWorld-Verified, Terminal-Bench) falling off hardest, as expected at this size. Apache 2.0, weights on HuggingFace, no first-party API.

ReasoningCodingVisionFunction callingTool useAgentic
1566.8
Elo · rank #133
Parameters
8.2B
Active params
8.2B (dense)
Context
131K tokens
Architecture
8.2B dense causal transformer with a distilled ViT-S/16 perception encoder (~0.4B); 131K-token context, the Glimmer line's consumer-GPU tier
License
Apache 2.0
Languages
—
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath82.1%#73
best: GPT-5.2 · 100.0%
CharXivVision63.5%#38
best: Qwen3.8-Flash-Next · 90.6%
GDPval-AAAgents812#53
best: Claude Fable 5 · 1932
GPQA DiamondReasoning68.7%#149
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning12.4%#97
best: Claude Opus 5 · 64.7%
IFBenchReasoning64.8%#38
best: Inkling-Medium · 85.3%
MCP AtlasAgents61.2%#10
best: Kimi K3.1 · 85.6%
MMMU-ProVision58.3%#54
best: Claude Opus 4.7 · 85.5%
OSWorld-VerifiedAgents52.7%#30
best: Claude Fable 5 · 85.0%
ScreenSpot-ProVision61.8%#7
best: GPT-6 Astra · 92.7%
SWE-bench ProCoding38.4%#67
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding61.2%#87
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding36.9%#81
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
5.8 GB
VRAM @ FP16
16.5 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
~95 tok/s on RTX 4090 (Q4, llama.cpp)
Quantizations
GGUF Q4 · Q4_K_M · BF16
Fine-tune it
Permissive
QLoRA7.2 GB1× RTX 3060 12GB
LoRA19.1 GB1× RTX 3090 24GB
Full fine-tune133.2 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.80 (1× RTX 3060 12GB)
Muse Glimmer family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)