OpenAI announced that its Jalapeño chip, the company's first custom inference chip, has achieved significant performance improvements in AI inference. The results show that Jalapeño can serve more AI work per unit of power while also returning responses more quickly. This marks a major advancement in the field, with the chip delivering both higher throughput and lower latency using a single architecture, unlike existing hardware systems that often have to make a tradeoff between the two. For customers, this can mean faster responses, more responsive agents, and more reliable access as demand grows. OpenAI's mission is to ensure that artificial general intelligence benefits all of humanity. These gains will help make increasingly capable AI more affordable and more broadly available. Source: openai

OpenAI models also accelerated Jalapeño’s development. Earlier generations helped the team design and bring up the chip, while the latest models are accelerating how the team optimizes and programs it. Jalapeño’s performance extends across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T, showing that the architecture works across models developed both inside and outside OpenAI. Across all three, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance. Source: openai

How we measured Jalapeño’s performance: OpenAI evaluated performance at a matched user experience, measuring how much useful AI work each system can complete per unit of power while meeting the latency customers and interactive agents require. This matters especially for agents, which need to complete many steps in sequence, so delays can compound across an entire task. To understand how Jalapeño performs in practice, the team tested it on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. They compared Jalapeño with leading commercially available AI systems across the tested operating range, from high-throughput serving to highly interactive, low-latency use. Jalapeño delivered a better combination of throughput, power efficiency, and latency. Although performance is sometimes reported per chip, the team believes the more useful standard is performance per unit of power. Source: openai