AI Briefing
KO

OpenAI Reveals Performance of Its Proprietary Inference Chip Jalapeño: Up to 1.9x Throughput per Watt and Up to 3.6x Latency Improvement

·2026.08.26 09:00

Key point

OpenAI's proprietary inference chip Jalapeño has demonstrated superiority over existing hardware in terms of throughput per watt and latency.

Details

OpenAI has released the initial performance test results for Jalapeño, its first self-developed inference chip. Jalapeño overcomes the trade-off between throughput and latency faced by existing hardware within a single architecture, simultaneously improving AI workload throughput per watt and response speed.

Performance Benchmark Results

OpenAI compared Jalapeño with major commercial AI systems using the public benchmark InferenceX. The models tested were GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T.

  • Throughput per watt: Improved by 1.5~1.9x compared to comparison systems based on peak throughput
  • End-to-end latency: Reduced by 1.7~3.6x compared to comparison systems
  • Interactive workloads: Performance measured to be 2.1~4.1x higher

Notably, for the largest model, Kimi K2.5 1T, peak performance per watt improved by approximately 1.5x, and latency improved by 3.4x. Jalapeño's rated power is 700W, but sustained power remained below 550W during test workloads.

Architecture and Strategy

Jalapeño features a co-designed integration of chips, memory, network, software, and rack-scale systems tailored to actual language model workloads. This is part of OpenAI's 'Full-stack' strategy, where the company directly controls the entire stack from models to hardware.

OpenAI plans to scale up Jalapeño production within the coming months to provide faster and more efficient AI services. This is expected to serve as a foundation for delivering the benefits of AGI (Artificial General Intelligence) to more users affordably and reliably.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.