AI Briefing
KO

Two Chips for the Agentic Era: Google's 8th-Gen TPU

·2026.04.23 10:00

Key point

Google unveiled TPU 8t for training and TPU 8i for inference.

Details

Google announced its 8th-generation TPU, splitting it into TPU 8t for training and TPU 8i for inference. Judging that the demands of agentic workloads have diverged, the company optimized learning and serving with separate architectures.

TPU 8t focuses on large-scale training.

  • A single superpod scales up to 9,600 chips and 121 ExaFlops.
  • Per-Pod compute performance improved roughly 3x over the previous generation.
  • TPUDirect, storage access 10x faster, and the combination of Virgo Network and JAX/Pathways target large-scale expansion and near-linear scaling.
  • The goodput target is 97% or higher, and RAS features include telemetry, automatic bypass routing, and OCS.

TPU 8i is tailored for latency-sensitive inference and agent execution.

  • 288GB HBM and 384MB on-chip SRAM reduce memory bottlenecks.
  • It applies an Axion ARM-based CPU host with NUMA isolation.
  • For MoE models, ICI bandwidth was expanded to 19.2 Tb/s, and the Boardfly topology reduces network diameter.
  • The on-chip CAE lowers global compute latency by up to 5x, and performance per dollar improved by 80%.

Both chips are combined with Google's own Axion hosts, 4th-generation liquid cooling, and integrated power management to boost system-level efficiency. Google has signaled general availability in the second half of this year, and it will be offered as part of AI Hypercomputer. It natively supports JAX, MaxText, PyTorch, SGLang, vLLM, and also provides bare-metal access.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.