Meta open-sources Muse Glimmer 30B Apache 2.0 for local always-on agents
Meta Superintelligence Labs released Muse Glimmer, a ~30B dense multimodal model with open weights under Apache 2.0 on Hugging Face, distilled from Muse Spark for always-on local agents on a Mac or PC with a single consumer GPU (coding, tool use, document/screenshot understanding, LLM-as-judge). Training used logit distillation from Muse Spark, mid-training on longer agent traces, then SFT plus on-policy distillation and RL; ~4-bit K-quant packs the LM under ~20 GB with a perception encoder and DFlash speculative-decoding drafter for responsive on-device generation. Day-0 paths include transformers, llama.cpp, vLLM, SGLang, and partners (Ollama, LM Studio, Unsloth, Together, Fireworks, OpenRouter), with AMD, Arm, Dell, Intel, and NVIDIA optimizing device runtimes—Meta’s first major open-weight return since Llama 4, distinct from API-only Muse Spark 1.2 / Muse Code.







