OpenAI's Jalapeño Chip Delivers Breakthrough Inference Performance, Benchmarks Show

At the Hot Chips conference on Tuesday, OpenAI unveiled detailed specifications and the first batch of benchmark results for its custom inference chip, codenamed Jalapeño. In tests using SemiAnalysis's InferenceX benchmark, Jalapeño outperformed currently available state-of-the-art inference processors, achieving higher token throughput per user and greater efficiency per kilowatt.

“The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI's head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”

Notably, the comparison was against an Nvidia Blackwell system—but by the time Jalapeño reaches full deployment, the competitive landscape may have evolved considerably. Ho estimated that Jalapeño would begin deployment in very small volumes by the end of 2026, with more significant scaling expected in 2027.

First announced in October 2025, Jalapeño was developed by OpenAI in close collaboration with Broadcom, with OpenAI's own models assisting in the development process. The company plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory to be developed in concert.

This full-stack approach enabled OpenAI to target specific bottlenecks in the inference pipeline. In particular, Jalapeño is designed to minimize delays during the prefill and communication phases, which often act as performance limiters. By optimizing data movement and keeping critical state local, the chip reduces latency and improves overall efficiency.

“We designed Jalapeño to minimize data movement and communication delays,” the company stated in a blog post. “This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.”

As AI workloads continue to scale, the ability to deliver high throughput with low latency and energy efficiency becomes increasingly critical. Jalapeño's architecture reflects a shift toward purpose-built silicon that can adapt to the demands of real-time AI services, positioning OpenAI for the next wave of inference innovation.

via TechCrunch AI

Related