Google is developing a new server chip called Frozen v2, which integrates parts of the Gemini AI model's architecture directly into silicon. This design aims to improve efficiency in serving AI responses. The chip is expected to be deployed starting in 2028 and is seen as a test run for specialized chips with smaller production volumes than the TPU line. According to sources, the chip could be 6 to 10 times more efficient than Google's current TPU chips. This efficiency gain is achieved by hardcoding portions of the model's architecture into the hardware, reducing compute steps and speeding up responses.

The concept for Frozen v2 reportedly originated from Jeff Dean, Google Deepmind's chief scientist. His initial design aimed to embed model weights directly into the chip, but this approach was abandoned due to its limited flexibility. Instead, Frozen v2 focuses on embedding the model's architecture, allowing for updates to weights while maintaining the underlying blueprint. The extent to which the architecture will be hardcoded remains undetermined, according to The Information.

Google's TPU chips are used internally and leased to external customers, with the company aiming to capture 10% of Nvidia's annual revenue. Frozen v2, however, is primarily intended to address internal challenges with AI compute capacity. If successful, the chip could provide a competitive edge by enabling Google to run powerful models at lower costs, potentially challenging OpenAI and Anthropic in the market.

Source: thedecoder