Google Cloud Next Unveils Eighth-Generation TPUs for Agentic AI

Google split its custom AI chips into training and inference models, signaling infrastructure costs now drive deployment architecture

Illustration: Google Cloud Next Unveils Eighth-Generation TPUs for Agentic AI
AI-generated illustration · Sylvaris

Two Chips, One Strategy

Google announced its eighth-generation Tensor Processing Units at Cloud Next in Las Vegas on April 22, breaking from tradition by introducing two distinct chips. TPU 8t is optimized for training large models, while TPU 8i targets cost-effective inference with near-zero latency. The split reflects a broader industry shift: as AI agents handle multi-step reasoning tasks, inference costs now rival training expenses.

The announcement positions Google's custom silicon as a direct alternative to NVIDIA's GPU dominance. Anthropic, Google Cloud's largest TPU customer, secured 3.5 gigawatts of next-generation capacity starting in 2027 through a deal announced alongside the TPU roadmap.

Platform for Agent Deployment

Beyond hardware, Google introduced the Gemini Enterprise Agent Platform, designed to build, govern, and scale autonomous AI systems. The platform includes an Agent Designer interface, persistent agent memory, and orchestration tools for multi-agent workflows.

The conference drew 32,000 attendees and focused on production deployment rather than research previews. Google also announced BigQuery fluid scaling, which reduces autoscaling costs by up to 34 percent, and new C4N and M4N compute families targeting agent workloads.

What It Means for Cloud Customers

The specialized chip architecture matters because it affects how organizations budget for AI. Training a frontier model once is expensive; running inference for millions of users continuously is often more expensive. Google's dual-chip approach allows customers to optimize costs by matching workload to silicon.

Anthropic's reported $30 billion annualized revenue in April, surpassing OpenAI for the first time, is tied directly to its use of Google Cloud TPUs. Every performance improvement and cost reduction announced at Cloud Next directly affects Anthropic's per-token pricing and competitive position ahead of its anticipated October 2026 IPO.

sources
more in Cloud & Infrastructure
Google Cloud becomes Alphabet's fastest-growing segment at over 20% of revenue Cloud infrastructure now represents more than one-fifth of Alphabet's total revenue and operating profit, marking a strategic shift beyond advertising dominance. US Army exhausts annual AI token allocation, implements usage limits Rapid AI adoption in military operations faces practical constraints when organizations underestimate infrastructure requirements. AMD and Anthropic reach $5 billion AI infrastructure deal AMD's largest AI infrastructure commitment to date signals intensifying competition in the data center GPU market dominated by NVIDIA.