Muse Glimmer 30B
OPENMeta's return to open weights, August 10, 2026: a 30B always-on local agent model under a real Apache 2.0 licence, distilled from Muse Spark and explicitly sized for consumer hardware. Quantized to ~4 bits the language model drops under 20 GB, leaving room for the KV cache, the perception encoder and the speculative-decoding drafter inside a 24 or 32 GB envelope — Meta says with minimal-to-no degradation on agentic tasks. BF16 weights, two 4-bit quantizations and the perception encoder all ship; llama.cpp, MLX, ExecuTorch, vLLM and SGLang integrations followed. It is the strongest model in its size class on the agentic rows Meta cares about (MCP Atlas 75.5 vs Gemma4-31B's 54.2 and Qwen3.6-27B's 62.5, DeepSearchQA 74.6, GAIA2 43.3, tau3-Banking 23.5) and on AIME 2026 (94.7), while Qwen3.6-27B still leads it on OSWorld-Verified, SWE-bench Verified, Terminal-Bench 2.1 and GDPval-AA. VRAM figures below are Meta's own stated envelopes: ~55 GB at full precision, under 20 GB for the language model at ~4-bit (Meta also ships a 17 GB K-Quant build; the ~20 GB figure is the one that includes headroom for the encoder and drafter). No first-party API rate — it is a download.