AI Pulse by Inblix

Google's 'Frozen v2' chip hardcodes Gemini's structure for 10x faster AI

The Decoder · Jul 20, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Google's 'Frozen v2' chip hardcodes Gemini's structure for 10x faster AI

Google is building a server chip that doesn’t just run its Gemini AI model — it has the model’s architecture physically etched into the silicon. Codenamed ‘Frozen v2,’ the chip could deliver 6 to 10 times the efficiency of Google’s current TPUs when serving AI responses, according to sources who spoke with The Information. The company plans to deploy it starting in 2028, though as a limited-run test vehicle rather than a mass-market replacement for its TPU line.

The name borrows from the machine learning concept of ‘freezing’ parameters — locking in values so they stop updating during training. In this case, portions of Gemini’s model structure get permanently baked into the hardware, eliminating compute steps and accelerating inference. Jeff Dean, Google DeepMind’s chief scientist, reportedly pitched the original idea with a more radical approach: embedding the model’s actual weights directly into silicon. That version died on the vine. A chip with frozen weights would only work with a single Gemini release and would be obsolete the moment Google trained a new version.

Frozen v2 sidesteps that problem. Instead of hardcoding the tuned parameters, it embeds the model architecture — the blueprint, not the specific settings. New weights can still be loaded onto the chip. How much of the architecture gets locked in hasn’t been finalized yet, and that decision will determine just how flexible — or brittle — the final silicon turns out to be.

Don’t expect to see this in Google Cloud’s catalog. The chip only works as long as Gemini’s fundamental architecture remains stable, which makes it useless for external customers who want general-purpose accelerators. Google already leases TPUs to Meta and pitches them as an Nvidia alternative through its ‘TPU@Premises’ program, with internal targets to capture 10% of Nvidia’s annual revenue. Frozen v2 serves a different master: Google’s own crushing demand for inference capacity. If the efficiency gains hold, Google could run powerful models at a fraction of the cost and undercut competitors like OpenAI and Anthropic on price. Inference economics are quietly becoming the real battlefield in AI.

💡 Key Takeaways

  1. Frozen v2 embeds Gemini's model architecture — not weights — into silicon, targeting a 6-10x efficiency leap over Google's current TPUs for inference workloads.
  2. Google scrapped an earlier design that froze model weights directly because it would have locked the chip to a single, quickly obsolete Gemini version.
  3. The chip won't be sold to cloud customers; it's strictly for Google's internal inference crunch and could reshape the cost structure against rivals like OpenAI and Anthropic.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles