Model
Model explorer

Muse Glimmer 30B

OPEN
Meta · Muse Glimmer family · released Aug 10, 2026

Meta's return to open weights, August 10, 2026: a 30B always-on local agent model under a real Apache 2.0 licence, distilled from Muse Spark and explicitly sized for consumer hardware. Quantized to ~4 bits the language model drops under 20 GB, leaving room for the KV cache, the perception encoder and the speculative-decoding drafter inside a 24 or 32 GB envelope — Meta says with minimal-to-no degradation on agentic tasks. BF16 weights, two 4-bit quantizations and the perception encoder all ship; llama.cpp, MLX, ExecuTorch, vLLM and SGLang integrations followed. It is the strongest model in its size class on the agentic rows Meta cares about (MCP Atlas 75.5 vs Gemma4-31B's 54.2 and Qwen3.6-27B's 62.5, DeepSearchQA 74.6, GAIA2 43.3, tau3-Banking 23.5) and on AIME 2026 (94.7), while Qwen3.6-27B still leads it on OSWorld-Verified, SWE-bench Verified, Terminal-Bench 2.1 and GDPval-AA. VRAM figures below are Meta's own stated envelopes: ~55 GB at full precision, under 20 GB for the language model at ~4-bit (Meta also ships a 17 GB K-Quant build; the ~20 GB figure is the one that includes headroom for the encoder and drafter). No first-party API rate — it is a download.

ReasoningCodingVisionFunction callingTool useAgentic
2077.8
Elo · rank #70
Parameters
29.6B
Active params
29.6B (dense)
Context
131K tokens
Architecture
29.6B dense causal transformer with a frozen ViT-G/14 perception encoder (~1.8B), 52 layers, hidden 6656, [Local, Local, Local, Global] attention with a 2048 sliding window, SwiGLU FFN, 202,048-token vocabulary; distilled from Muse Spark by logit distillation and shipped with a DFlash speculative-decoding drafter
License
Apache 2.0
Languages
API price (in/out)
No hosted API
Modalities
text · vision
Benchmark results
Bar shows position within the tracked field; marker = field best
AIMEMath94.7%#20
best: GPT-5.2 · 100.0%
CharXivVision78.8%#18
best: Qwen3.8-Flash-Next · 90.6%
DeepSearchQAAgents74.6%#5
best: GPT-5.6 Sol · 93.0%
GAIA2Agents43.3%#1
best: this model · 43.3%
GDPval-AAAgents953#43
best: Claude Fable 5 · 1932
GPQA DiamondReasoning83.5%#70
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning22.0%#67
best: Claude Opus 5 · 64.7%
IFBenchReasoning77.0%#13
best: MiniMax M3 · 83.0%
MCP AtlasAgents75.5%#4
best: Kimi K3 · 84.2%
MMMU-ProVision74.0%#31
best: Claude Opus 4.7 · 85.5%
OSWorld-VerifiedAgents65.9%#20
best: Claude Fable 5 · 85.0%
SciCodeCoding43.6%#1
best: this model · 43.6%
ScreenSpot-ProVision75.4%#5
best: GPT-6 Astra · 92.7%
SWE-bench ProCoding51.2%#48
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding76.0%#42
best: Claude Opus 5 · 96.0%
τ³-BankingAgents23.5%#5
best: GPT-5.6 Sol · 33.0%
Terminal-Bench 2.0Coding51.7%#58
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
20 GB
VRAM @ FP16
55 GB
Fits on (Q4)
RTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Meta benchmarks the 17 GB K-Quant build with a quantized DFlash drafter on MacBook M4-Max, M5-Max and RTX 5090 rather than a 4090, and reports 'fluid conversation and real-time agent interaction' without publishing a tokens/sec figure.
Quantizations
Q4_K_M · BF16
Fine-tune it
Permissive
QLoRA20.6 GB1× RTX 3090 24GB
LoRA63.6 GB1× A100 80GB
Full fine-tune475.6 GB4× H200 141GB
QLoRA SFT on ~10k samples ≈ $15.25 (1× RTX 3090 24GB)
Muse Glimmer family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)