Meta has released Apache-licensed weights for Muse Glimmer, a 30-billion-parameter agentic model intended for local deployment. The hardware and benchmark claims make the proposition concrete, but they are Meta's measurements, and the harder test is whether developers can safely operate the tools around it.
Meta, the social-media company behind a previous open-weight Llama model line, has released Muse Glimmer: a 30-billion-parameter AI model whose weights developers can download under the Apache 2.0 license. Meta says the model is designed for agentic work—multi-step tasks that call software tools—and can run on a Mac or a PC with one consumer graphics card.
The immediate bet is on where the agent runs. Unlike a hosted model accessed through an API, Meta says Glimmer can run without cloud infrastructure or a network connection. That gives developers more control over customization and the operating environment. It does not demonstrate that a downloaded model will be cheaper, safer or easier to operate once an application adds tools, permissions and long-running tasks.

Meta’s introductory illustration for Muse Glimmer, its open-weight agentic model. Source: Meta AI Research.

Meta’s company-reported quantization comparison for Muse Glimmer’s local-hardware configurations. Source: Meta AI Research.

Meta’s company-reported DFlash decoding results across three hardware platforms. Source: Meta AI Research.
Muse Glimmer is a dense causal transformer with a perception encoder, or about 29.6 billion parameters including its vision component, according to Meta's model card. It accepts text and images, generates text and supports a context length of at least 131,072 tokens. Its stated uses include coding agents, schema-based tool calls, interpreting documents and screenshots, synthetic-data generation and evaluating other models' outputs.
Meta says Glimmer was distilled from Muse Spark, its larger model. The company describes a sequence of pre-training on Spark outputs through logit distillation, mid-training on longer-context agent-focused data and post-training using supervised fine-tuning, on-policy distillation and reinforcement learning. The model card says the broader training data came from publicly available material, third parties, Meta products and services, and vendor and Meta curation.
That pedigree is important, but it is not a claim that the smaller model matches the teacher. Meta's safety documentation says Glimmer is generally less capable than Spark. An available preview of a separate account describes it as nearly identical to Spark; the official model card instead supplies the more limited comparison and the detailed 30B specification.
Meta's announcement says full-precision weights would require more than 55 GB of memory. Its 4-bit releases compress the language model to under 20 GB, leaving room for working memory, the perception encoder and a speculative-decoding drafter. The model card reports these two target configurations:
| Release configuration | Target hardware | Meta-reported average degradation |
|---|---|---|
| K-Quant-Dynamic | 32 GB VRAM | 0.2% |
| K-Quant-17GB | 24 GB VRAM | 1.0% |
The degradation is an average across accuracy metrics on 15 common benchmarks, not a universal measure of agent quality. It also does not include the resource demands of every tool, document or application built around the model.
Meta also reports that, with its DFlash speculative-decoding drafter, the 17 GB configuration reached 233.4 tokens a second on an Nvidia RTX 5090, versus 74.9 without speculation. On an Apple M4 Max it reported 37.8 versus 23.7 tokens a second, and on an M5 Max 50.2 versus 26.6. Those results use batch size one and greedy decoding; they describe a controlled inference test, not end-to-end agent latency.

Meta’s company-reported benchmark comparison of Muse Glimmer, Gemma4-31B and Qwen3.6-27B. Source: Meta AI Research.
Meta compared Glimmer's high-reasoning mode with Gemma4-31B and Qwen3.6-27B in their thinking modes. It reported leads on MCP Atlas, DeepSearch QA and SWE-Bench Pro. But Qwen led on GDPVal-AA v2, OSWorld-Verified and TerminalBench 2.1, while Gemma led on GPQA Diamond and Humanity's Last Exam text.
The safety comparison points the same way. On Siren AgentDojo, Meta reports a 28.4% attack-success rate for Glimmer, against 25.6% for Gemma and 40.3% for Qwen; lower is better on that measure. On CI Memories, Glimmer's 26.4 violation score is above Gemma's 12.1, while its 64.8 coverage score is below Qwen's 66.9. These are useful signals about the tested models, but not evidence of an across-the-board lead or a deployment guarantee.
Mark Zuckerberg, Meta's chief executive, used the launch to argue that U.S. policy should reduce what he called extra restrictions on domestic labs' training data if American open-weight models are to lead. In a report on his statement, Zuckerberg also said restricting access to foreign open-source models would not be an effective answer. The report identifies Moonshot's Kimi K3, Alibaba's Qwen3.8-Max and DeepSeek's V4-Flash as Chinese open-weight competitors, while leading models from OpenAI, Anthropic and Google are closed.
Meta has a record in this argument: its Llama open-weight series began in 2023, though an account of the launch says not every version impressed developers. That context makes Glimmer less a first experiment in openness than a new attempt to make the case with a locally deployable agent model.
The release is not, however, a permanent split between an open small model and a closed flagship. The same account says Muse Spark was released earlier this year as Meta's closed, more powerful model and that Meta plans to release an open-weight version of Spark. Alexandr Wang, the former Scale AI chief executive who now leads Meta's Superintelligence division, is overseeing the group that produced Glimmer. For now, the relevant comparison is between a downloadable 30B model and competing local models—not a final statement of Meta's access policy for its strongest system.
Meta assessed Glimmer as moderate or lower risk in chemical and biological, cyber and loss-of-control categories under its Advanced AI Scaling Framework. The cyber and loss-of-control designations were inferred from Spark because Glimmer is broadly weaker, the model card says.
The same document sets a clear boundary on that assessment. Meta says its testing cannot cover every scenario, recommends deployers conduct application-specific safety testing and advises additional guardrails. For agentic systems that can take real-world actions, it specifically recommends human confirmation before irreversible steps. Its evaluation categories also include prompt-injection resistance, data minimization and respecting the limits imposed by an agent scaffold.
Zuckerberg has said in the same report that Meta would create a governance structure giving independent directors authority to approve model-release safety criteria. That is a company proposal, not a substitute for the decisions made by a developer who chooses what an installed agent may read, send, change or purchase.
The runtime ecosystem is still forming. Meta says integrations for llama.cpp, MLX and ExecuTorch are due in coming days; a launch-day report says Ollama 0.32.7 already added Muse Glimmer support. These integrations may make the model easier to run, but they do not establish that a particular agent setup is safe or reliable.
The next evidence should come from reproducible, application-level tests: independent benchmarks on the released quantizations, measurements that include tools and long context rather than generation alone, and evaluations of the permissions and prompt-injection defenses deployed around the model.
Glimmer's release makes a real local-agent option available under a permissive license. Whether that option becomes a credible alternative to hosted agents depends less on the download than on whether developers can make those surrounding controls dependable without sacrificing the local control that gives the model its appeal.
Get concise AI news and useful context from the Magica team.
Read the newsletterKeel Infrastructure has completed its U.S. bitcoin-mining shutdown and is developing former power sites for high-performance computing. But its filing says none of the three priority U.S. sites had begun HPC operations or recorded related revenue as of Aug. 7, leaving contracts, funding and execution—not the shutdown itself—as the next test.
Bitdeer’s conditional lease for its Tydal, Norway data-centre conversion is material because the company plans to move 225 megawatts of the former mining site to colocation. But public accounts describe different contract values, terms and counterparties, leaving the customer, funding and revenue implications unresolved.
Nvidia is reportedly discussing an AI-infrastructure funding package of up to $500 billion with large private-capital firms. The talks underline the need to finance chips, data centres and power, but no commitments or risk-sharing terms have been disclosed.
Micron has signed 16 multiyear strategic customer agreements that commit buyers to specified memory volumes and, in many cases, pricing ranges. The contracts support its expansion plans and reduce risk for a portion of the business, but they do not yet make most of Micron’s revenue contractual or settle whether demand will outlast the supply shortage.
Anthropic, Macquarie Asset Management and GIC have formed Theseus Infrastructure to develop, operate and lease dedicated data centers to Anthropic. The investors will own the platform and provide most project equity, but the companies have not disclosed sites, capacity, costs or a delivery timetable—and Anthropic already has a separate $50 billion infrastructure plan with Fluidstack.
Sen. Bernie Sanders has asked OpenAI, Anthropic and Meta to pause AI development, invoking their conditional safety commitments. The request is not a pause order, but it extends his earlier push to halt new AI data-center construction pending safeguards.
Intel has proposed a $15 billion underwritten common-stock offering, with a further $2.25 billion underwriter option. The gross base deal equals about half of Intel’s June cash and short-term investments, yet the preliminary prospectus leaves the price, share count, net proceeds and allocation of capital unspecified.
Archer has agreed to acquire Boeing subsidiaries Wisk Aero, Insitu and SkyGrid in a stock-and-warrant transaction that keeps Boeing economically and technologically involved. Insitu gives Archer an operating defense business; the deal leaves unresolved how Archer will fund and organize the autonomous Wisk program alongside its piloted Midnight aircraft.
The police watchdog is investigating an unnamed senior Derbyshire detective over alleged AI use in decision logs—records that may be disclosed in court—as ministers press ahead with a £75 million national police-AI programme.
Industries that buy memory chips say tighter supply and higher prices are threatening products from network equipment to medical devices. Their competing requests to Washington—add capacity, reserve supply for domestic industry, or keep Chinese suppliers out—leave Commerce with a policy conflict, not a settled response.
Moore Threads has approved a plan to issue H shares and seek a Hong Kong main-board listing after first-half revenue rose 147%. The proposal has no disclosed size, price, timetable or use of proceeds, leaving investors to weigh a fast-growing cloud-products business against continuing losses, cash use and GPU competition.
A media tally says at least 37 people have been arrested in U.S. data-center protests in 2026, though it is not an official database. Two local cases show a narrower, documented conflict: who can speak at hearings over proposed infrastructure, under what rules, and with what recourse when those rules are enforced.
Mark Zuckerberg says Meta will pursue personal superintelligence for billions of people, with free versions, paid compute and private agents. The plan turns on capacity, release rules and governance that the company has not yet specified.
TSMC reported NT$467.58 billion in July revenue, up 44.7% from a year earlier and described by contemporaneous reports as a monthly record. The total supports its bullish AI-and-HPC outlook, but the disclosure does not break out the customers, products or AI contribution behind the increase.
South Australia plans to begin a royal commission on artificial intelligence on 1 October and report by 1 July 2027. Its immediate test is whether the government can define a focused remit on AI use, work, public services and infrastructure consequences before federal standards move ahead.
Meta has seeded a $1 billion fund for U.S. communities where it owns and operates data centers. The commitment places community benefits at the center of Mark Zuckerberg’s AI expansion case, but the company has not disclosed the timing, allocation rules or reporting that would let a host community judge its bargain.
California’s SB 903 would stop companion chatbots from being presented as psychotherapy and set consent, clinician-review and accountability rules for specified AI uses in mental-health care. The bill is not a blanket ban on AI support; its unresolved challenge is where clinical assistance becomes clinical judgment.
Microsoft’s latest in-house text-to-image model is second in Arena’s August 10 ranking, behind OpenAI’s GPT-Image-2. The result is an early competitive signal from a live head-to-head-vote leaderboard, while Microsoft’s promised reference, grounding and control features—and their wider rollout—remain to be detailed.
ChatGPT search can surface available times from OpenTable, Resy and Yelp and complete some reservations. The rollout makes ChatGPT a new place to start a booking, but coverage, confirmation and every post-booking change remain with the reservation providers.
Yang Qi, Game Science’s co-founder and Black Myth art director, says the team will keep AI-generated content out of Black Myth: Zhong Kui’s design and asset-production stages. The Weibo response is a limited production choice for an early project, not a disclosed company-wide AI ban or a release-date update.