AI Briefing
KO

SemiAnalysis TPUv7 Ironwood Analysis: Costs Approximately 19% Lower Than B200 at 100 Tokens Per Second Per User

·2026.09.07 09:00

Key point

SemiAnalysis analyzed that under external buyer TCO and 100 tokens per second per user conditions for Qwen3.5 397B FP8 inference, Ironwood's cost per token is approximately 19% lower than B200.

1 / 15

Details

SemiAnalysis compared the inference performance of TPUv7 Ironwood and B200/B300 on the Qwen3.5 397B FP8 model with an 8k1k workload. Under the condition of 20 tokens per second per user, Ironwood recorded a total throughput of 9,364 tokens per second per chip, higher than B200 (8,903) and B300 (8,925). Separately, under the condition of 100 tokens per second per user with external buyer TCO applied, Ironwood's total cost per 1 million tokens was approximately $0.181, about 19% lower than B200 ($0.222). When applying Google's internal TCO ($1.03/chip-hour), the throughput per cost at a concurrency of 256 was analyzed to be 76.7% higher than B200 and 130.2% higher than B300. However, at the same concurrency of 256, Ironwood's average Time to First Token (TTFT) was 5.41 seconds, slower than B200 (3.75 seconds) and B300 (2.40 seconds). These results are comparisons under these specific conditions and do not apply to all latency targets. The software stack is transitioning from the existing TorchAX to a new native TorchTPU backend.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.