SemiAnalysis TPUv7 Ironwood Analysis: Costs Approximately 19% Lower Than B200 at 100 Tokens Per Second Per User
Key point
SemiAnalysis analyzed that under external buyer TCO and 100 tokens per second per user conditions for Qwen3.5 397B FP8 inference, Ironwood's cost per token is approximately 19% lower than B200.
Details
SemiAnalysis compared the inference performance of TPUv7 Ironwood and B200/B300 on the Qwen3.5 397B FP8 model with an 8k1k workload. Under the condition of 20 tokens per second per user, Ironwood recorded a total throughput of 9,364 tokens per second per chip, higher than B200 (8,903) and B300 (8,925). Separately, under the condition of 100 tokens per second per user with external buyer TCO applied, Ironwood's total cost per 1 million tokens was approximately $0.181, about 19% lower than B200 ($0.222). When applying Google's internal TCO ($1.03/chip-hour), the throughput per cost at a concurrency of 256 was analyzed to be 76.7% higher than B200 and 130.2% higher than B300. However, at the same concurrency of 256, Ironwood's average Time to First Token (TTFT) was 5.41 seconds, slower than B200 (3.75 seconds) and B300 (2.40 seconds). These results are comparisons under these specific conditions and do not apply to all latency targets. The software stack is transitioning from the existing TorchAX to a new native TorchTPU backend.