OpenAI published the first measured benchmark results for Jalapeño, the custom inference chip it developed with Broadcom, on August 25, 2026. Across three open-weight models, the company reported 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the Nvidia systems it tested against. OpenAI plans to begin deploying the chip inside its own compute infrastructure by the end of this year.
The numbers OpenAI published
Against an Nvidia GB200 system running GPT-OSS 120B, OpenAI reported roughly 1.9 times higher peak mixed throughput per kilowatt, at 85,448 versus 44,960, and end-to-end latency of 1.03 seconds versus 1.80 seconds. Against a GB300 running DeepSeek R1 670B, it reported about 1.7 times higher peak throughput per kilowatt and latency of 1.65 seconds versus 5.99 seconds. On Kimi K2.5 1T, the largest public model tested, the figures were approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency.
Power ratings matter to how those comparisons were built. Jalapeño is rated at 700 watts, though OpenAI says measured sustained draw stayed at or below 550 watts on the tested workloads. The Nvidia systems it was measured against carry package ratings of 1,200 watts for the GB200 and 1,400 watts for the GB300. For the most interactive workloads, where a user or an agent is waiting on each response, OpenAI reported 2.1 to 4.1 times higher performance.
How the test was run
The results were presented at the Hot Chips conference and measured on InferenceX, a public benchmark from SemiAnalysis that covers the full process of serving an AI request rather than raw chip throughput. OpenAI normalized results using each accelerator’s published chip power rating, arguing that performance per unit of power is a more useful standard than performance per chip.
Richard Ho, OpenAI’s head of hardware, called the results “a very, very significant performance advance over state of the art” on a press call. The same reporting flags the obvious caveat: the comparison is against Nvidia Blackwell hardware available today, and Ho estimated Jalapeño would deploy at the end of 2026 in very small volumes, with meaningful deployment arriving in 2027. Nvidia’s lineup will not stand still in between. OpenAI also disclosed that its internal testing on frontier OpenAI models showed a wider advantage, a claim no outside party has verified.
Does cheaper inference reach smaller users?
Inference cost sits underneath the price of every AI tool a freelancer or small business pays for each month. OpenAI’s own framing is candid about who benefits first: producing more useful work from the same power improves operating leverage, letting useful work and revenue grow faster than the cost to serve. Broader adoption is listed as a downstream effect, not the primary one.
So the realistic expectation is that efficiency gains show up as capacity before they show up as price. When serving gets cheaper, what usually changes first is how much work a plan allows, how tight the rate limits are, and how many multi-step agent runs fit inside a subscription, rather than the monthly figure itself. That is worth watching for anyone leaning on AI agents to run parts of a client workflow, because latency compounds across an agent task. A chain of twenty steps inherits every delay in every step, which is precisely the bottleneck OpenAI says Jalapeño was designed around.
One limit is worth stating plainly: Jalapeño is an inference chip, built to serve models after they are trained. It does not change the economics of training the next generation of frontier models, and OpenAI has said it will keep deploying Nvidia and other partner accelerators widely for both training and inference.
The claim still has to survive 2027
Jalapeño went from initial design to tapeout in nine months, with OpenAI’s own models assisting the design and optimization work. For selected GPT-OSS attention and mixture-of-experts blocks, the company said AI-generated implementations ran 1.5 to 1.8 times faster than versions written by human experts, a figure that applies to those blocks rather than a full model.
Gen 2 is described as deep in development and Gen 3 as taking shape. Whether the first generation holds its lead depends on what Nvidia has fielded by the time Jalapeño ships in volume, and on whether production qualification and software maturity land on schedule. Until independent testing arrives on deployed hardware, these remain first-party numbers from a company with a strong interest in reducing what it pays for compute.





