AI Briefing
KO

Two chips for the agent era, our 8th-generation TPU

·2026.04.22 21:15

Key point

Google unveiled the training-focused TPU 8t and the inference-focused TPU 8i.

1 / 2

Details

At Google Cloud Next, Google unveiled its 8th-generation TPU. This generation is split into TPU 8t for training and TPU 8i for inference, targeting model training, deployment, and agent workloads respectively.

TPU 8t is designed for large-scale training.

  • About 3x compute performance per pod compared to the previous generation
  • 9,600 chips, 2PB shared HBM, and 2x interchip bandwidth in a single superpod
  • 121 ExaFlops of compute
  • Improved utilization through TPUDirect and 10x faster storage access
  • Combined with Virgo Network, JAX, and Pathways, it aims for near-linear scaling up to a logical cluster of up to 1 million chips
  • Targets over 97% goodput through automatic link bypass, OCS, and real-time telemetry

TPU 8i is built for latency-sensitive inference and agent execution.

  • 288GB HBM and 384MB on-chip SRAM keep the working set on-chip
  • Uses Axion ARM-based CPU hosts to optimize system efficiency
  • ICI bandwidth of 19.2 Tb/s and the Boardfly architecture reduce network latency
  • CAE offloads global operations, cutting on-chip latency by up to 5x
  • Claims an 80% improvement in performance-per-dollar over the previous generation

Both chips support JAX, MaxText, PyTorch, SGLang, and vLLM, and also offer bare metal access. Google is bundling these as part of AI Hypercomputer, positioning it as agent-era infrastructure that integrates compute, storage, networking, and software. Both chips are scheduled for general availability later this year.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.