TLDRocket
Sign in

Google is building a chip with Gemini baked into the silicon

TNW Covered by 19 sources

Google is reportedly baking Gemini's architecture directly into a future chip, nicknamed Frozen v2. Trade flexibility for speed and power savings — a huge bet if it works.

Based on reporting by TNW — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Chips usually run models. Google wants a chip that is the model. According to reporting from The Information, picked up by Reuters and Bloomberg Law, Google is developing a project internally called Frozen v2 that etches Gemini's neural-network architecture into the silicon itself, rather than loading it as software the way every GPU or TPU does today. Alphabet's stock ticked up as much as 3.7% on the news, even though Google hasn't confirmed anything and the chip, if real, is still years from shipping.

The mechanics are the interesting part. Normal chips shuffle a model's weights in and out of memory constantly, and that movement burns power and time. Frozen v2 would lock the shape of Gemini's architecture into the hardware permanently, while still letting engineers swap in new weights as the model improves. Google hasn't decided how much of the model gets hardwired this way, per the report, but the payoff being floated is striking: 6 to 10 times better efficiency than Google's current custom AI chips, measured in tokens served per watt. This would sit alongside Google's TPU line, not replace it, with a target deployment as early as 2028.

The timing lines up with a real problem. The Information frames Frozen v2 partly as Google's answer to an internal AI capacity crunch so tight that Google Cloud has reportedly turned away outside customers. At data-center scale, every watt saved translates directly into money, and a chip built for exactly one model can shed a lot of the overhead a general-purpose chip has to carry. It would also cut latency, which matters for anything real-time, like voice assistants, where waiting even a beat feels broken. There's a strategic layer too: Google already builds TPUs to reduce its Nvidia dependence, and a Gemini-specific chip pushes that self-reliance even further.

Google isn't the only one chasing this. A startup called Taalas is already selling silicon with a model's weights and architecture printed directly onto it, in a chip it calls Hardcore. The company claims up to 17,000 tokens a second, compared to roughly 150 per user on a leading Nvidia GPU, and says it skips the expensive high-bandwidth memory that's currently in short supply industry-wide. The trade being made across the board is the same: give up flexibility to gain speed, cost savings and lower power draw. For a company serving one dominant model to billions of users, that trade starts to look less crazy.

The catch is obvious. AI architectures evolve fast, and a chip frozen around today's Gemini design could feel stale by 2028, even with updatable weights. And Google still hasn't confirmed the project exists; a spokesperson would only say teams experiment with efficiency ideas and that not everything ships. So this is a signal more than a product announcement. But it's a signal worth watching, because it marks a shift from building chips that can run any model to building chips that are fused to one.

My take — AI-written commentary, not fact-checked reporting

This is the logical endpoint of building your own model and your own silicon under one roof, and it's exactly why vertical integration scares competitors who rely on general-purpose hardware. I'd bet Google ships something like this before 2028 if the capacity crunch stays this bad, because desperation is a great forcing function. The real tell will be whether OpenAI or Anthropic, who don't control fabs, start scrambling for similar deals — that's when you'll know the silicon war has actually started.

Read more about this at: TNW

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.