OpenAI’s Jalapeño chip now has a number attached to it. Back in June, when OpenAI and Broadcom first announced it, I wrote that the performance claims were the softest part of the rollout: no FLOPS, no bandwidth figures, just a line about “substantially better performance per watt than current state of the art” with nothing to check it against. At Hot Chips this week, OpenAI finally put an actual benchmark on the table, and it is not flattering to Nvidia.
SemiAnalysis ran Jalapeño through its InferenceX suite inside OpenAI’s own lab, working alongside OpenAI’s engineers, and came back with more tokens per user and more throughput per kilowatt than a currently shipping Nvidia Blackwell system. Richard Ho, OpenAI’s head of hardware, put it plainly on the press call: Jalapeño can serve more AI work per unit of power while also returning responses faster. Those are the two numbers that actually matter for a company burning through inference compute at OpenAI’s rate, not peak FLOPs on a spec sheet nobody’s workload ever touches.
I want to be careful about how independent this actually is. SemiAnalysis was invited into OpenAI’s lab, benchmarked a chip OpenAI controlled access to, and published results the same day as OpenAI’s announcement. That is a step up from a marketing deck with nobody outside the building in the room, and SemiAnalysis has embarrassed vendors before when their numbers did not hold up under its testing, so I am not waving this off. But “independent” is doing a lot of work in every headline running with this story this week, and the honest read is that OpenAI picked the referee and the referee showed up ready to be impressed.
The architecture explanation is more convincing than the marketing copy. Jalapeño is built to minimize data movement during the prefill and communication phases of inference, the stretches where a chip sits mostly idle waiting for the KV cache to move instead of actually computing anything. OpenAI can make that tradeoff because it controls the model, the serving stack, and the chip all at once, which is the same full stack argument I made about this chip back in June, when it was still a press release with no silicon to check it against. A merchant GPU has to stay general enough to run whatever ships next quarter. Jalapeño does not, because OpenAI already knows exactly what it is running on it.
None of this ships in volume yet. Ho says small-volume deployment lands by the end of 2026, with anything that matters showing up in 2027, which is the same window I flagged as the real test when this was still an announcement with no chip behind it. Sixteen months from a hiring push to a working tapeout is unusually fast for an ASIC, and OpenAI apparently proved the thing actually runs by porting Doom to it using nothing but Codex prompts, which is a better demo than anything in the benchmark deck.
Nvidia is not losing to Jalapeño specifically. It is losing the same argument for roughly the fourth time this year, after AMD made a similar CUDA moat claim when Helios went into production, and after TSMC quietly started routing Nvidia’s own CoWoS overflow to outside packaging houses because Nvidia’s supply chain cannot keep up with itself. Nvidia’s answer to all of this has mostly been financial instead of technical, locking customers into hundreds of billions in vendor financing so the demand curve stays fixed no matter how crowded the chip race gets. That is a rational move if you are worried about the field catching up. It is not the move of a company confident its performance lead holds on its own.
I am not ready to call Blackwell beaten. A benchmark run in the vendor’s own lab against a shipping competitor from two product cycles back is not the same as a chip surviving a real fleet at production scale, and Nvidia’s next architecture lands before Jalapeño does in volume. I am also leaving the Broadcom side of this alone today, because who actually captures the margin on a chip partnership like this one is its own post. What I did not expect was OpenAI showing a real number this soon. The direction has been obvious since June. It is just a better-defined direction now, and I will believe the 2027 deployment claim once the racks are running and someone other than OpenAI can time them.
Sources
- TechCrunch, OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show, August 25, 2026
- SemiAnalysis, OpenAI Jalapeño: Better Than Nvidia Blackwell, August 25, 2026
- OpenAI, Jalapeño: First Results, August 25, 2026
- Yahoo Finance, OpenAI claims chips outperform Nvidia, August 25, 2026