AI Briefing
KO

Google Cloud C4, GPT OSS TCO Improved by 70%

·2025.10.16 09:00

Key point

Google Cloud C4 VMs based on Intel Xeon 6 improved TCO by 1.7x for GPT OSS inference.

Details

Through collaboration between Intel and Hugging Face, benchmarking of Google Cloud's latest C4 VM (equipped with Intel Xeon 6 processors) performance showed a 1.7x improvement in TCO (Total Cost of Ownership) compared to the previous generation C3 VM.

Key results are as follows:

  • 1.4-1.7x improvement in TPOT (Time Per Output Token) throughput per vCPU
  • Reduced cost per hour compared to C3 VM

This performance improvement is thanks to the MoE (Mixture of Experts) expert execution optimization merged into the Hugging Face transformers library. This optimization moves away from the existing approach where all experts processed all tokens, instead guiding each expert to process only its assigned tokens, reducing wasted computation (FLOPs) and increasing utilization.

The benchmark used OpenAI's open-source MoE model GPT OSS, demonstrating CPU inference efficiency for large-scale models through the Intel Xeon 6 (Granite Rapids) architecture.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.