Google is reportedly developing a new server chip designed to make its Gemini artificial intelligence models respond faster while consuming less energy. The project, internally known as “Frozen v2,” could place some parts of the AI model directly into the chip itself.

According to The Information, citing people familiar with the project, Google hopes to begin using the new semiconductor as early as 2028. The design is still under development, and engineers are reportedly deciding how much of Gemini’s model architecture or stored information should be permanently embedded in silicon.

Bringing Gemini closer to the hardware

Large AI models depend on constant data transfers between processors and high-bandwidth memory. Moving model parameters back and forth consumes energy and can slow the generation of responses.

Frozen v2 reduces that traffic by placing selected fixed elements of Gemini directly into the chip architecture. The project’s name reportedly comes from the idea of “freezing” part of the model into the hardware instead of repeatedly loading it from memory.

The chip could deliver between six and 10 times more AI tokens per unit of electricity than Google’s current custom processors, according to the report. That figure, however, represents an internal development target rather than a performance result independently confirmed by Google.

Embedding model components directly into silicon could improve speed and efficiency, but it would also reduce flexibility. Software teams can update AI models frequently, while hardware teams may need years to design and manufacture a new chip.

Google would therefore need to balance performance gains against the risk that the hardware could become too closely tied to one generation of Gemini.

Frozen v2 would not replace Google’s TPUs

Reports do not describe the chip as a new version of Google’s Tensor Processing Units, or TPUs. Instead, Google is developing Frozen v2 as a separate class of specialized hardware that could operate alongside the company’s existing AI accelerators.

Google originally created TPUs for internal machine-learning workloads, including systems supporting its search business. The company later expanded their use to Google Cloud and to the training and operation of Gemini models.

Its eighth-generation TPU family includes hardware designed for both AI training and inference. Training chips handle the computationally intensive process of building models, while inference chips generate responses after a user submits a prompt.

Frozen v2 would reportedly perform a narrower role. By optimizing the chip more closely for a specific Gemini architecture, engineers could sacrifice general-purpose flexibility in exchange for faster and more energy-efficient inference.

AI companies are accelerating the custom chip race

The project reflects growing pressure across the technology industry to reduce the cost of running large AI systems.

As model sizes and user demand increase, companies require larger data centers, more electricity and greater access to advanced processors. Custom chips can help reduce dependence on external suppliers while allowing hardware to be designed around a company’s own AI software.

Google is already one of the most experienced technology companies in this field. Its TPUs have become an important part of the company’s internal infrastructure and cloud computing strategy.

Other AI companies are moving in the same direction. OpenAI has worked with Broadcom on its own AI hardware, while Anthropic has reportedly explored semiconductor partnerships with Samsung Electronics.

These efforts could eventually create stronger competition for Nvidia, whose graphics processors remain widely used for training and operating advanced AI systems.

A chip built around a specific model

Frozen v2 remains an unconfirmed project, and Google has not publicly announced the chip or its planned specifications.

The launch date reported for 2028 could also change as engineers refine the design. Semiconductor teams frequently revise projects before production, particularly when they build them around rapidly evolving AI models.

Still, the idea behind Frozen v2 highlights a broader shift in artificial intelligence development. Competition is no longer focused only on creating larger or more capable models.

The companies building those systems are also trying to determine how they can deliver billions of AI responses faster, at lower cost and with less energy.