Google develops Frozen v2 chip to hardwire Gemini architecture into silicon; targets 6–10x inference efficiency by 2028
Google is developing "Frozen v2," an inference-specific chip that embeds Gemini's neural-network architecture directly into silicon, according to sources cited by The Information on July 20. The chip is projected to deliver 6 to 10 times more tokens per watt than Google's current generation Tensor Processing Units (TPU 8i), with deployment targeted as early as 2028. Unlike general-purpose accelerators from NVIDIA or Google's own TPU line, Frozen v2 will hardwire Gemini's structural logic into the chip's circuitry, eliminating runtime overhead for architectural decisions.
The design shifts from an earlier "fully hardwired" concept that would have frozen model weights (rendering the chip obsolete with each model update) to a "flexible hardwired" approach that locks only the stable architectural foundation. This preserves architectural efficiency gains while maintaining a multi-year chip lifecycle. Google's internal compute crunch drove the project; the company is facing severe bottlenecks in serving Gemini at the scale required for current and planned workloads.
Frozen v2 is positioned as a parallel track alongside Google's general-purpose TPU line, not a replacement. A $2 billion capex bet with no confirmed timeline, the project trades versatility for extreme specialization—a bet that Gemini's transformer structure remains stable through 2028 and beyond. If Gemini undergoes major architectural revision (e.g., shifting to state-space models), the hardware could become partially obsolete.
For infrastructure planners: Google's move signals a broader shift toward model-specific silicon over general-purpose accelerators. Competitors without equivalent vertical integration—including cloud providers leaning on NVIDIA—face margin pressure if Frozen v2 reaches production and lowers Gemini's serving cost. Watch Alphabet's Q2 earnings (July 22) and any public disclosures about the project's status.
Sources
- Primary source
- techtimes.com
“Frozen v2 hardwires Gemini architecture, targets 6–10x tokens per watt over TPU 8i”
- techcrunch.com
“Google didn't confirm or deny the project; stock climbed 3% on The Information report”
- fourweekmba.com
“Frozen v2 bakes Gemini blueprint into silicon; 2028 deployment target”