Google’s Gemini 2026: The Future of AI with Frozen v2 Chip

Table of Contents


Google’s Frozen v2 chip aims for AI efficiency

Frozen v2 Efficiency Gains

  • Reported codename: “Frozen v2” (internal)
  • Reported goal: improve Gemini inference efficiency
  • Reported metric: tokens generated per unit of power (tokens-per-watt)
  • Reported magnitude: ~6x to 10x vs Google’s existing AI chips
  • Reported timing: sometime in 2028
  • Confirmation status: not confirmed by Google; described publicly as ongoing research/experimentation

What is Google’s new AI chip Frozen v2 and how does it improve efficiency?

Alphabet is reportedly designing a new server chip—internally dubbed “Frozen v2”—with a clear goal: make Google’s in-house Gemini models cheaper and more efficient to run at scale. The chip is positioned as part of Google’s “full stack” approach, where the company co-designs hardware and software to optimize real-world workloads.

What makes Frozen v2 notable, according to The Information’s reporting as relayed by TechCrunch, is the idea that it’s built specifically to serve Gemini more efficiently than Google’s existing AI chips. Efficiency here is framed in operational terms. That metric matters because inference—serving model outputs to users—can become a massive ongoing cost as AI features spread across products.

Google didn’t validate the specific Frozen v2 details. But it did underline the strategy behind such efforts: constant experimentation, selective projects moving to production, and deep integration “from the ground up” so systems are “highly optimized for real-world workloads.”

Tokens per Watt Explained
Tokens-per-watt (the metric in the reporting) can be thought of as:

  • Output: how many tokens the model generates
  • Cost: how much electrical power the hardware consumes to generate them

So a “6x tokens-per-watt” claim means: for the same power budget, the system could generate ~6x as many tokens (or generate the same tokens using ~1/6 the power), assuming comparable conditions.
“Hardwiring model elements” (as described in reporting elsewhere about the project) generally implies moving some repeated inference work from software into fixed-function silicon. The upside is efficiency; the tradeoff is that changing those hardwired parts later can be slower than updating software.

When is the Frozen v2 chip expected to be released?

The reported timeline is long by consumer-tech standards but typical for advanced silicon: Frozen v2 is slated for release sometime in 2028, according to The Information, which cited anonymous sources. That puts the chip firmly in the “next cycle” of AI infrastructure planning—something that matters to investors and enterprise customers trying to forecast capacity, costs, and competitive positioning.

A 2028 target also signals that Frozen v2 is not a quick patch for today’s compute crunch, but a strategic bet on where AI economics are headed. If the chip is designed around Gemini’s needs, it suggests Google expects its model family to remain central to its product roadmap for years—and that serving those models efficiently will be a differentiator.

Google’s public statement to TechCrunch adds an important caveat: not every research project becomes a production chip. The company described ongoing research and experimentation, emphasizing that “rigorous exploration” is core to its approach. In other words, 2028 is a reported plan, not a guaranteed launch date.

Still, the existence of a reported schedule—paired with Google’s broader hardware-software co-design messaging—helps explain why the market treated the story as more than idle speculation.

Gemini “Frozen v2” Timeline
Timeline clarity (based on public reporting):

  • Now: reports describe an internal project (“Frozen v2”) aimed at Gemini efficiency
  • 2028 (reported): potential release window cited by The Information via anonymous sources
  • Unknowns: final product name, whether it ships at all, and what “6–10x” looks like under real production workloads

How much more efficient is the Frozen v2 chip compared to existing AI chips?

The headline claim is striking: Frozen v2 could be between six and 10 times more efficient than Google’s existing AI chips, based on the number of tokens generated per unit of power, according to The Information’s reporting as relayed by TechCrunch. That’s not a generic “faster” claim; it’s explicitly about energy efficiency per unit of model output, which is increasingly the metric that matters when AI moves from demos to daily usage.

If that range proved accurate in production, it would translate into a meaningful reduction in the power required to serve the same volume of AI responses—potentially lowering operating costs and easing pressure on data center power budgets. It could also allow Google to serve more Gemini capacity with the same energy envelope, a critical advantage amid broader concerns about AI computing constraints.

It’s also notable that the comparison is to Google’s existing AI chips, not to Nvidia’s GPUs or other competitors’ accelerators. That framing suggests Frozen v2 is an internal step-change—an attempt to improve Google’s own baseline economics for Gemini inference.

Google itself did not confirm the 6–10x figure in its response, sticking instead to general language about performance and efficiency research.

Item What’s being claimed (as reported) What it’s compared against How it’s measured What’s not specified in the report
Efficiency gain 6x to 10x Google’s existing AI chips Tokens generated per unit of power Exact workload mix, model version, batch sizes, latency targets, and whether gains hold across products

Has Google confirmed the development of the Frozen v2 chip?

Google has not directly confirmed the specific report about Frozen v2. In a response to TechCrunch, the company didn’t confirm the chip’s existence or the reported details—such as the name, the 2028 timeline, or the 6–10x efficiency claim. But it also didn’t deny the report.

Instead, Google offered a carefully worded statement that reinforces the plausibility of such a project without validating it. The company said its teams are “constantly researching and experimenting with new innovations” to deliver “maximum performance and efficiency” for users and customers. It added that not every project moves into production, but that this exploration is central to its “full stack approach.”

The key line is strategic: Google highlighted “co-designing our hardware and software from the ground up” to ensure systems are “integrated and highly optimized for real-world workloads.” That’s consistent with the broader industry trend toward tighter coupling between models and the infrastructure that runs them.

For readers, the practical takeaway is that Frozen v2 remains reported, not officially launched or productized. But Google’s statement makes clear that custom silicon and deep optimization are not side projects—they’re part of how the company intends to compete in AI.

Interpreting Reported Certainty Levels
How to read the certainty level in this story:

  • What Google said (to TechCrunch): it is “constantly researching and experimenting,” not every project ships, and it co-designs hardware/software for “real-world workloads.”
  • What was reported (via The Information, per TechCrunch): an internal chip project called “Frozen v2,” a possible 2028 release window, and a 6–10x tokens-per-power efficiency claim.

Practical checkpoint: treat the codename, timeline, and 6–10x figure as reported plans/estimates until Google (or a product launch) confirms them.

Why are AI companies focusing on producing their own chips?

The push toward custom chips is driven by a mix of economics, supply constraints, and strategic control. As TechCrunch summarized, AI companies increasingly want their own silicon to make in-house models run more efficiently and to address global shortages in AI computing capacity. When demand for AI accelerators outstrips supply, relying entirely on third parties can become a bottleneck—both for scaling products and for controlling costs.

Efficiency has also become a selling point in a market that’s more skeptical than it was during earlier waves of AI hype. Concerns about AI spending have “dampened the market euphoria,” making it harder for companies to justify massive infrastructure bills without a clear path to improved unit economics. A chip that generates more tokens per watt is, in effect, a narrative about discipline: better performance, lower power, and potentially lower cost per response.

There’s also the Nvidia factor. The industry has long depended on Nvidia’s AI hardware dominance, and major AI makers are trying to wean themselves off that dependency. Custom chips don’t eliminate Nvidia overnight, but they can reduce exposure to pricing, supply timing, and platform constraints.

Recent examples underscore the trend: OpenAI announced its first custom chip, an inference processor dubbed Jalapeño, and Anthropic was reported to be discussing a chipmaking partnership with Samsung.

Drivers of Custom AI Chips
Why big AI labs build custom chips (three drivers):
1) Unit economics: better tokens-per-watt can lower the long-run cost per response
2) Capacity & supply: shortages/lead times make “buying everything off-the-shelf” a scaling risk
3) Strategic control: less dependence on a single dominant vendor’s pricing, roadmap, and platform constraints

What is Google’s investment plan for its AI strategy?

Alphabet’s AI ambitions come with a price tag large enough to shape investor sentiment on their own. Earlier this year, Google said it plans to spend between $180 billion and $190 billion as part of building out its AI strategy. That level of expenditure has been a source of investor concern—less about whether AI matters, and more about whether the returns will justify the scale and timing of the outlay.

This is where a project like Frozen v2 fits into the story even before it exists as a shipping product. If Google can credibly argue that it’s improving the efficiency of running Gemini—especially in a way that reduces power consumption per token—then the company can frame its spending as investment in long-term cost advantages, not just capacity accumulation.

In practical terms, custom silicon can be a lever for controlling the ongoing costs of AI services. Training is expensive, but inference at scale can become the recurring bill that defines margins. A chip designed to run Gemini more efficiently would support Google’s broader “full stack” positioning: models, infrastructure, and deployment optimized together.

Google’s public comments to TechCrunch emphasized exactly that philosophy—hardware and software co-designed—even as it avoided confirming any specific Frozen v2 roadmap.

Google AI Capex Snapshot
Capex snapshot (as stated by Google earlier this year):

  • Planned spend: $180B–$190B to build out its AI strategy
  • Why efficiency stories matter to investors: if inference gets cheaper per token, the same infrastructure budget can support more usage (or better margins) over time

How did the news of the Frozen v2 chip affect Google’s stock?

The market reaction was immediate and positive. After The Information’s report was published, Alphabet’s stock climbed about 3% on Monday morning, according to TechCrunch. The timing mattered: the move came ahead of Google’s earnings report later that week, when investors tend to be especially sensitive to signals about costs, margins, and capital intensity.

The logic behind the bump is straightforward. Investors have been wary of Alphabet’s massive planned AI expenditures, and a credible report suggesting a major efficiency leap—6x to 10x tokens per unit of power—reads like a partial answer to the “show me the payoff” question. Even if the chip is years away, it implies Google is working on structural improvements that could reduce the long-run cost of serving Gemini across products.

It’s also a reminder that, in the AI era, hardware strategy is no longer a back-office detail. For hyperscalers, chips can influence everything from product pricing to data center expansion plans. A report about a more efficient inference chip can therefore move the stock not because it changes next quarter’s revenue, but because it changes the perceived trajectory of future AI economics.

Alphabet Shares Rise After Report
Market reaction (as reported by TechCrunch): Alphabet shares rose ~3% Monday morning after the report, ahead of the company’s earnings later that week.

The Future of Google’s AI Chip Development

Google’s reported Frozen v2 effort sits at the intersection of three pressures: the rising cost of serving AI, the scarcity of compute, and the strategic risk of over-dependence on a single dominant chip supplier. Whether Frozen v2 ships in 2028 as reported—or evolves into something else—the direction is clear: AI leaders increasingly see silicon as a competitive moat, not a commodity.

Understanding the Significance of the Frozen v2 Chip

Frozen v2 matters less as a codename and more as a signal of intent. The report suggests Google is pursuing a step-change in tokens-per-watt efficiency for Gemini, and Google’s own statement reinforces the underlying method: hardware-software co-design aimed at “highly optimized” real-world workloads. In an environment where AI features are proliferating, efficiency becomes product strategy—because it determines what can be offered, at what price, and at what scale.

Implications for AI Efficiency and Market Dynamics

If the industry is moving from “who has the best model” to “who can serve it sustainably,” then efficiency gains become a market-moving story. Frozen v2, OpenAI’s Jalapeño, and Anthropic’s reported chip discussions all point to the same conclusion: the next phase of AI competition will be fought not only in model architectures, but in the infrastructure economics underneath them.

This lens is shaped by Martin Weidemann’s work building and scaling technology businesses where unit economics, infrastructure constraints, and hardware-software tradeoffs directly impact what can be shipped sustainably.

Efficiency Gains, Flexibility Tradeoffs
What a “hardwired-for-efficiency” chip can trade off (even if the upside is real):

  • Efficiency vs flexibility: the more logic you bake into silicon, the harder it can be to adapt to fast-changing model architectures
  • Update cadence: software can ship weekly; silicon typically moves on multi-year cycles
  • Workload specificity: big gains may depend on serving a narrower set of Gemini inference patterns
  • Engineering complexity: co-designing model + compiler + silicon can raise execution risk and lengthen timelines

This article reflects publicly available information and Google’s public statements as of the time of writing. Some details—such as timelines, naming, and performance figures—come from anonymous-source reporting and may be incomplete or change as plans evolve. If Google later confirms or releases related hardware, official product disclosures will be the most reliable source for final specifics.

Scroll to Top