OpenAI unveiled benchmark results for its first custom chip, Jalapeño, at the Hot Chips conference, showing it outperforms Nvidia's Blackwell and Rubin in throughput per watt and token latency. The chip is designed for inference tasks, running AI models without training them. It is a general-purpose LLM inference accelerator, not specifically tuned for OpenAI's models. OpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems. For interactive workloads, the company says performance is 2.1x to 4.1x higher.

The results are based on tests using SemiAnalysis's public InferenceX benchmark. OpenAI provided the numbers, and SemiAnalysis verified some runs on-site in the lab. The models tested were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T. On GPT-OSS, Jalapeño achieved about 1,400 tokens per second per user. On Deepseek R1, it topped 700 tokens per second on a single concurrent request. At matched decoding speed, Jalapeño achieves 54x to 104x the token throughput per kilowatt compared to the best available accelerator, depending on the model. These numbers were posted without using techniques like multi-token prediction or speculative decoding, while some of the comparison systems did rely on those optimizations, so there's still room for improvement.

SemiAnalysis notes that the fairer comparison isn't Blackwell but Nvidia's newer Vera Rubin platform, since both use HBM4 memory. Even here, Jalapeño outperforms Vera Rubin, even though Nvidia's accelerator uses the multi-token prediction optimization that Jalapeño hasn't adopted yet. On total cost of ownership per token, the two come out roughly even. However, Nvidia and AMD have already published results with larger models like Deepseek V4 Pro and Kimi K3 that haven't been tested on Jalapeño yet. Also, while Rubin systems are already shipping to customers, Jalapeño reportedly hasn't moved beyond engineering samples.

Source: thedecoder