OpenAI published the first performance numbers for Jalapeño on Tuesday, the custom inference chip it co-developed with Broadcom and unveiled in June. The company’s summary of the results was four words long: “We made a chip and it is fast.”
The claims are specific. Across three tested models — GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T — OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput, and 1.7 to 3.6 times lower end-to-end latency, than the comparison systems. On highly interactive workloads, the kind that decide whether a chatbot feels instant or sluggish, it reports 2.1 to 4.1 times higher performance. Tom’s Hardware put the physical comparison plainly: a roughly 700-watt part measured against a 1,400-watt Nvidia flagship.
The more interesting technical claim is architectural. Inference hardware usually forces a trade: tune for throughput and latency suffers, tune for latency and you leave throughput on the table. OpenAI says Jalapeño delivers both from a single architecture. If that holds up under independent testing, it is a design result, not a benchmark result.
The timing was the message
Nvidia reports fiscal Q2 after the close today, with the Street looking for roughly $92 billion in revenue. Publishing custom-silicon benchmarks into that news cycle is not an accident of the calendar. Broadcom shares moved on the release; the read-through is that the biggest buyer of AI compute in the world has a credible second source, and it is telling everyone about it.
Our take: This is a margin story for Nvidia before it is ever a volume story. OpenAI has said it will keep buying Nvidia accelerators — demand is growing faster than any one supplier can fill. But the moment a large customer can point at a working alternative, the negotiation changes, and Nvidia’s pricing power is the whole reason its gross margin sits where it does. Watch the margin line tonight, not the revenue line.
Three caveats worth holding onto
- OpenAI graded its own homework. These are vendor-published benchmarks with no independent reproduction. Every chip company publishes numbers that flatter its chip; the methodology, not the multiple, is what matters.
- The comparison is a generation old. Jalapeño was measured against Blackwell-class systems. Nvidia’s Rubin, with newer HBM4 memory, is already shipping to customers. A 1.9x advantage over last year’s part is a different claim from a 1.9x advantage over what Nvidia sells today.
- It is not deployed yet. OpenAI plans to start putting Jalapeño into its own infrastructure by year end. Broadcom chief executive Hock Tan has described a meaningful ramp through 2027 and full-scale operation in the first half of 2028. Silicon that works on a test bench and silicon that runs a global service at 99.9% uptime are separated by about eighteen months of unglamorous engineering.
What to watch
- Nvidia’s gross margin guidance tonight, and whether management is asked directly about custom ASICs.
- Broadcom’s next quarter — its AI revenue line is now the cleanest public proxy for how fast custom inference silicon is actually shipping.
- Any independent benchmark of Jalapeño against Rubin rather than Blackwell.
- OpenAI’s own API pricing. If the chip does what the deck says, the savings should eventually show up as a price cut — and that is the only proof that ends the argument.
