Google just bet its inference future on a chip built for one model
The New Stack 1 month ago 41 ● 19 sources
Google is developing a specialized chip called Frozen v2 designed specifically for its Gemini AI model, which would hardwire parts of Gemini's architecture while keeping weights updatable. The chip is projected to deliver six to ten times more tokens per watt compared to Google's current AI chips. If successful, this approach could significantly reduce inference costs for developers using Gemini while establishing a trend toward model-specific silicon rather than general-purpose accelerators.