Jalapeño's First Results Demonstrate Industry-Leading Speed and Efficiency in AI Inference
Key point
OpenAI's proprietary inference chip Jalapeño has been confirmed to significantly outperform existing systems in performance per watt and latency.
Details
OpenAI has released the initial performance test results for Jalapeño, its first self-developed inference chip. Jalapeño demonstrated the ability to simultaneously improve throughput and latency, which previously required trade-offs in existing hardware systems, within a single architecture.
In tests targeting various models such as GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, Jalapeño achieved 1.5–1.9x higher AI workloads per watt at peak throughput and 1.7–3.6x lower end-to-end latency. Notably, it recorded 2.1–4.1x higher performance in interactive workloads.
The performance evaluation was conducted using SemiAnalysis's public benchmark InferenceX, ensuring matched user experience levels. While Jalapeño was rated at 700W, it operated at sustained power of 550W or less during actual test workloads. OpenAI stated that Jalapeño is the result of a full-stack integrated design spanning models, chips, memory, and networks, and will expand into faster and more efficient products in the future.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.