Garp Independent AI & technology journalism
Saturday, September 26, 2026 Sign In · Join Subscribe
Latest Ando wants to take on Slack with a team messaging app that lets humans and agents work together

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  Meta is back with Muse Glimmer: local, agentic, multimodal, and open source

AI News

Meta is back with Muse Glimmer: local, agentic, multimodal, and open source

Meta is back with Muse Glimmer: local, agentic, multimodal, and open source

Hugging Face architecture Text Decoder Perception Encoder Transformers Llama.cpp Speculative Decoding Speculative Decoding with llama.cpp Support for Muse Glimmer vLLM with transformers backend Fine-tuning with TRL Demos Connect OpenClaw to Muse Glimmer Hey Muse Glimmer, quantize yourself Hey Muse Glimmer, deploy yourself Hey Muse Glimmer, optimize yourself Wrapping Up Hey Muse Glimmer, research the Hub Great news from the OGs of open source LLMs! Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for local agentic use cases.

Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it’s ideal deploying locally for privacy, reducing costs, or just hacking around. It’s intended for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups. To celebrate, we are shipping with Meta day-0 support in transformers, llama.cpp, vLLM, Inference Endpoints, and other libraries. We built a few cool things and explain our findings in this blog. Check out the demos below for inspiration. You can find all Muse Glimmer models in this collection. Muse Glimmer is a dense 30B parameter model consisting of: In addition to the main VLM, there’s also a speculative decoding drafter implemented on DFlash. Usage of this module is optional, and it can provide much faster generation in exchange for some memory cost. We found this drafter to be particularly well suited to structured content generation such as coding. The language model uses the following architecture components: Muse Glimmer uses one image encoder to handle both images and videos. Unlike the relatively small vision encoders used in other VLMs, this is a sizable 2B ViT-like model designed after the Perception Encoder architecture. Perception Encoder was previously introduced by Meta as a backbone for various downstream spatial and multimodal tasks.The encoder patchifies images to a shape of 2 frames x 3 channels x 14 x 14, and passes them through a linear layer for projection. An interpolated absolute position embedding from a learned position table is then added to these embeddings. These are then sent to the vision tower which consist of 50 layers and GELU MLPs. Similar to the language model, the attention pattern consists of three window attention layers followed by one full attention layer.