Google Cloud opened Next 2026 at Mandalay Bay in Las Vegas on 22 April with its eighth generation of TPUs, split into two chips: TPU 8t for training and TPU 8i for inference. Thomas Kurian also launched the Gemini Enterprise Agent Platform and the Agentic Data Cloud, and backed the pitch with usage numbers, including 330 customers who each processed more than 1 trillion tokens in the past 12 months.
Two chips for two different jobs
Training a model and serving it are different problems. Training wants enormous clusters that behave like one machine for weeks. Inference wants to answer millions of requests cheaply and quickly. Google has decided those needs are now different enough to justify separate designs.
TPU 8t is the training chip. It scales to 9,600 TPUs in a single superpod with 2 PB of shared high bandwidth memory. Google says it delivers “3x processing power” compared with Ironwood, the previous generation, and “2x better performance/Watt”.
TPU 8i is the inference chip, with 1,152 TPUs per pod. Here the headline claim is about money rather than raw speed: Google says it offers “80% better performance per dollar for inference” than the prior generation.
What caught my attention is that the inference chip is sold on cost, not speed. That tells you where Google thinks the pressure is. Training a frontier model is a huge but occasional bill. Serving it to every user, every day, is the bill that never stops.
Agents and data
The software announcements follow the same logic. As the names suggest, Gemini Enterprise Agent Platform is aimed at businesses building and running AI agents, and the Agentic Data Cloud at connecting those agents to company data. Agents are exactly the kind of workload that consumes tokens around the clock, which is why the chips and the software arrived on the same stage.
The usage figures are there to prove demand already exists. Beyond the 330 customers above 1 trillion tokens, 35 customers passed 10 trillion tokens in the same 12 months. Traffic through Google’s direct API has reached more than 16 billion tokens per minute, up from 10 billion the previous quarter. And paid monthly active users of Gemini Enterprise grew 40% quarter on quarter in the first quarter of 2026.
By the numbers
| Item | Figure | Source |
|---|---|---|
| TPU 8t chips per superpod | 9,600 | Google Cloud Blog |
| Shared HBM per TPU 8t superpod | 2 PB | Google Cloud Blog |
| TPU 8t processing power vs Ironwood | 3x | Google Cloud Blog |
| TPU 8i chips per pod | 1,152 | Google Cloud Blog |
| TPU 8i inference performance per dollar vs prior generation | 80% better | Google Cloud Blog |
| Customers above 1 trillion tokens in 12 months | 330 (35 above 10 trillion) | Google Cloud Blog |
| Direct API throughput | 16+ billion tokens per minute | Google Cloud Blog |
Why it matters
The split chip strategy is the most interesting part of the day for me. For years the conversation about AI hardware has been about training, because training is where the dramatic cluster sizes and headlines are. Google designing a dedicated inference chip, and leading with performance per dollar, is a signal that the industry’s centre of gravity is moving towards running models, not just building them.
That shift matters to developers and startups in the Arab world more than the training race does. Very few teams in the region will ever train a frontier model. Many are already building products on top of hosted models, including Arabic language assistants, customer service tools and document processing for banks and government. For them, the price of inference decides whether a product is viable. If cheaper inference hardware leads to lower prices on the platforms they use, that is a direct benefit, although nothing announced at the keynote tells us how much of the saving will reach API prices.
What to watch
- Next runs until 24 April, and the remaining sessions should add detail on the agent platform and pricing.
- When TPU 8t and TPU 8i become available to customers, and in which cloud regions.
- Whether the token growth continues at the same pace when Google next reports quarterly figures.
Sources
- Google Cloud Blog, Thomas Kurian, “Welcome to Google Cloud Next ’26”, 22 April 2026, https://cloud.google.com/blog/topics/google-cloud-next/welcome-to-google-cloud-next26
- Google, “Google Cloud Next ’26 recap”, April 2026, https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/google-cloud-next-26-recap/