Google just bet its inference future on a chip built for one model
The New Stack Amanda Caswell ● Covered by 2 sources
Google is developing a specialized chip called Frozen v2 designed specifically for its Gemini AI model, which would hardwire parts of Gemini's architecture while keeping weights updatable. The chip is projected to deliver six to ten times more tokens per watt compared to Google's current AI chips. If successful, this approach could significantly reduce inference costs for developers using Gemini while establishing a trend toward model-specific silicon rather than general-purpose accelerators.
Why it matters
The race to make AI inference cheaper is pushing chip design beyond general-purpose accelerators. We’re now moving toward silicon that The post Google just bet its inference future on a chip built for one model appeared first on The New Stack.