Google is reportedly designing Frozen v2, a Gemini-specific inference chip that could hardwire part of a model architecture while retaining updateable weights. The unconfirmed project’s six-to-10-times tokens-per-watt estimate is not comparable to Google’s published TPU claims, leaving its practical advantage unproved.
Google is reportedly considering a server chip tailored to Gemini, an approach that would make the model’s underlying design part of the hardware. The project, informally called Frozen v2, is reported to be a new line alongside—not a replacement for—Google’s tensor processing units.
The consequential point is not that Google is entering custom silicon; it already has. It is that Frozen v2 would reportedly commit a future chip to parts of one model architecture. The available reporting does not show whether that commitment will make a production service cheaper or faster than Google’s existing TPU stack.
People familiar with the work told the report that engineers are still settling the design and the amount of model information to hardwire. Deployment is planned as soon as 2028, according to the account, which traces the claims to reporting by The Information.
Google did not confirm Frozen v2. In a statement published in a separate account, the company said its teams research hardware-and-software co-design and cautioned that not every project reaches production. That is evidence of a general research program, not confirmation of the reported architecture or timetable.
The report, relaying The Information’s account, says the project is meant to help with an AI-computing capacity squeeze that has created internal tension and led Google Cloud to decline some outside-customer deals. That account identifies a possible reason to pursue efficiency, but it does not establish how much capacity Frozen v2 would add, who would receive it, or whether it would change cloud pricing.

Google Cloud’s official product image of an Ironwood TPU. Source: Google Cloud Blog.
The reported projection is six to 10 times as many AI tokens served per unit of power as Google’s latest homegrown AI chips. It is a forecast by project workers, with no disclosed model version, workload, latency target, chip configuration, measurement method or deployment cost. It therefore cannot support a claim of six-to-10-times cheaper service.
Google’s own published TPU figures illustrate the mismatch. In a November 2025 product announcement, it said its seventh-generation Ironwood TPU offered 10 times peak performance over TPU v5p and more than four times per-chip performance for training and inference over TPU v6e, or Trillium. Those are Google performance claims against named prior generations, not tokens-per-watt figures for Frozen v2; they cannot be converted into a head-to-head result.
The existing TPU system is also not a generic baseline with no inference work behind it. Google describes Ironwood as serving high-volume, low-latency inference as well as training and reinforcement learning. It says its GKE Inference Gateway can reduce time-to-first-token latency by up to 96% and serving costs by up to 30% through load balancing across TPU servers. Those are separate company claims, with no stated equivalence to Frozen v2’s estimate, but they show the relevant alternative is an evolving hardware-and-software stack rather than an idle legacy chip.
The reported design would not put one fixed Gemini release on a chip. It would encode parts of the neural-network architecture while permitting new weights to be loaded. An earlier concept reportedly led by Google DeepMind Chief Scientist Jeff Dean would have embedded the weights themselves and was abandoned because it would have tied the chip to a single model version, the account says.
That distinction limits the usual “frozen model” description, but it does not resolve the durability question. Google’s own Ironwood announcement says that model architectures are constantly shifting. Separately, a Bloomberg report cited in the reporting said Google delayed its latest Gemini model after it fell short of internal goals. Neither account says Frozen v2’s fixed elements will become obsolete; together, they show why a 2028 hardware choice cannot be treated as a settled model-roadmap decision.
The concept also is not a new category of chip. Taalas has demonstrated a chip with an 8-billion-parameter Llama model, including its architecture and weights, embedded in silicon, according to one account. The same reporting describes d-Matrix and SambaNova as pursuing other inference-oriented designs. Their products and claims are not a like-for-like benchmark for a prospective Gemini chip, but they undercut any suggestion that model-specific hardware begins with Frozen v2.
Other AI companies are pursuing custom chips without evidence here that their technical designs match Google’s reported approach. The reporting says OpenAI announced an inference processor called Jalapeño in June and that Anthropic had been reported to be discussing a chip partnership with Samsung. Those developments establish competitive interest in serving efficiency, not a common price, performance or deployment comparison.
Frozen v2’s next decision is whether to turn a research effort into a product. Before its projected efficiency is meaningful, Google would need to identify what architectural elements are fixed and publish comparable tests that specify the Gemini model, tokens-per-watt denominator, latency, hardware configuration and workload.
It would also need to show how those results change serving cost, capacity and reliability relative to contemporary TPUs. Until then, the strongest supported conclusion is narrower than a claim of a decisive new chip: Google is reportedly testing a more rigid form of co-design while its flexible inference platform continues to advance.
Get concise AI news and useful context from the Magica team.
Read the newsletterEnigma has emerged from stealth with a $71 million seed round and a public test involving at least 100 real robots. The startup says the interactions could improve robot interfaces and models, but it has not disclosed a specific commercial use case, customers or performance evidence.
Starbucks has retired Automated Counting, a computer-vision tool for counting milk and beverage components in North American coffeehouses. The company says it is standardizing inventory counts; reported miscounts leave the economics and operational value of the replacement unproven as Starbucks pursues more frequent replenishment.
HSBC plans to open a Global AI Centre of Excellence in Singapore in the second half of 2026 and hire more than 100 specialists. The bank is adding a deployment hub to an existing Google Cloud programme; the unresolved question is whether it can substantiate the promised gains while retaining human accountability in wealth, payments and treasury work.
Microsoft says MAI-Cyber-1-Flash lets its MDASH vulnerability system route routine work to a smaller in-house model and reserve GPT-5.4 for difficult cases. The resulting benchmark and cost claims apply to the combined system, leaving customer deployment, pricing and remediation outcomes to test in Project Perception’s preview.
A federal judge treated Anthropic’s purchased-book scanning and its model training as distinct fair uses, while leaving its earlier pirate-library copying exposed. Newly unsealed Project Panama records show how the company tried to replace that source of books at industrial scale.
Meta has raised its planned investment in the Hyperion campus in Richland Parish from $10 billion to more than $50 billion. Its promises of jobs, local investment and lower power bills now depend on tax terms and a 20-year utility arrangement whose long-run protections are still being tested.
AT&T has signed an agreement to explore D-Wave's annealing technology in more network operations after an early optimization workload dropped from about an hour to under 15 seconds. The commercial question is whether that company-reported result can deliver better decisions at the cost, scale and reliability of a live carrier network.
Yuyuantantian, an account linked to China Central Television, has proposed matching access to AI-model capabilities and risks rather than treating models as simply open or closed. The commentary is not a regulation, and separate reported discussions of limits on overseas access remain preliminary.
NVIDIA has made an undisclosed investment in Safe Superintelligence and promised access to its Vera Rubin platform that the companies say can increase the lab’s compute tenfold. The partnership gives Ilya Sutskever’s closed AI lab another hardware route, but leaves its research, capacity allocation and commercial terms unexamined in public.
NVIDIA has formed the Open Secure AI Alliance with more than 30 technology and security partners, using Hugging Face’s response to an AI-enabled intrusion as its case for deployable open-weight defensive tools. Members have identified work on identity, scanning and patching, but the coalition has not yet set out a common process for evaluating, disclosing or remediating failures.
CXMT's 466% Shanghai debut made the DRAM maker China's most valuable listed company, but the price was set with only 6.73% of its enlarged share capital freely tradable and before the company has proved how far its new capital can take it against tooling restrictions and established memory rivals.
Investigations and a TikTok search study found synthetic or impersonated clinicians reaching large audiences with dubious health claims. The evidence is not a platform-wide measure, but it shows why an AI label alone may not tell viewers whether a medical recommendation is credible.
Alphabet, Amazon and Meta have disclosed 2026 capital-spending plans that add to $520B to $550B at their stated ranges. The total conveys the scale of their infrastructure push, but it is not a comparable measure of AI-only spending or of the obligations each company is taking on.
Anthropic has made Claude Fable 5 generally available with classifiers that redirect certain risky requests to Claude Opus 4.8, while the same underlying model, Claude Mythos 5, remains available with cyber safeguards lifted only to approved organizations. The models share a listed token price, making vetting, monitoring and capacity—not a purchasable premium tier—the immediate constraint on less-restricted access.
Nvidia is reportedly discussing a roughly $250 billion financing backstop that could help OpenAI lease a proposed 10GW Ohio data-center campus. But no deal has been announced, other AI companies are pursuing the federally controlled power, and the reported 800MW first phase would be only a fraction of the planned buildout.
xAI has added a built-in /deep-research command to Grok Build, turning a coding agent's existing delegation controls into a bounded, source-backed reporting process. The documentation is unusually specific about workflow limits and failure reporting, but it does not establish that the process improves accuracy, speed or cost.
An independently reported account says President Donald Trump posted AI-generated images of an Iranian tanker seizure, a burning tanker and a strike on Kharg Island. The posts did not announce a new operation, but they followed reported U.S. strikes on the island’s military assets and renewed attention to its oil terminal.
Alphabet’s proposed $80 billion equity package is partly tied to AI infrastructure, but its $40 billion at-the-market program is primarily intended to cover employee-equity tax obligations. Together with Microsoft’s $190 billion 2026 capital-expenditure forecast, the disclosures show a more complicated shift in how hyperscalers fund constrained compute capacity.
Intel is putting €5 billion into equipment and connections at its operating Leixlip campus to increase output of Intel 3-based Xeon processors. The investment strengthens an existing European manufacturing base, but Intel has not disclosed its added wafer capacity, the split between its own products and foundry work, or an outside customer.
Unitree’s As2-W pairs powered wheels with a quadruped chassis and advertises fast travel over rough terrain. Its official page, however, leaves price unpublished and gives materially different range figures, while launch coverage conflicts on core specifications—making a working deployment the meaningful next test.