OpenAI is benchmarking its custom silicon against Nvidia’s best, and according to early numbers, it’s winning. At the Hot Chips conference this week, OpenAI shared the first public benchmark results for Jalapeño, its in-house inference chip built with Broadcom. As TechCrunch reported, those results show Jalapeño outperforming current state-of-the-art inference processors on both tokens per user and throughput per kilowatt, tested against Semianalysis’s InferenceX benchmark.
That’s a meaningful combination to beat. Tokens per user reflects how well the chip handles many simultaneous requests, while throughput per kilowatt speaks directly to operating cost at scale. Richard Ho, OpenAI’s head of hardware, described it plainly on a press call: “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”
The catch is the comparison. Jalapeño’s benchmarks are measured against an Nvidia Blackwell system, which is the current generation. By the time Jalapeño reaches meaningful deployment, Nvidia will almost certainly have moved forward. Ho placed full-scale rollout in 2027, with only “very small volumes” shipping by end of 2026. That’s a long window for a competitor like Nvidia, or Google with its TPUs, or Amazon with Trainium, to close any gap.
Still, the underlying architecture tells a more interesting story than the headline numbers. Jalapeño was designed specifically to reduce friction at the prefill and communication phases of inference, two spots that regularly create bottlenecks in production systems. OpenAI says the chip keeps model state, including the KV cache, local and explicitly placed, so the system can activate the right mix of compute, memory, and networking for each inference phase without unnecessary data movement. That’s not a general-purpose design choice. It’s targeted at the specific workloads OpenAI runs at scale.
That full-stack thinking is what separates this from a typical chip announcement. OpenAI built Jalapeño as a multigenerational platform, with models, chips, memory, and AI products all developed together. This mirrors what Google has done with TPUs for years, and what Meta is attempting with MTIA. The difference is OpenAI is doing this while also being one of the largest external customers of Nvidia. Jalapeño isn’t a replacement yet. It’s a hedge, and a signal of where OpenAI wants to be in three to five years.
For developers and infrastructure teams, the near-term impact is limited. But the direction matters. If Jalapeño delivers on efficiency at scale, it gives OpenAI more control over its cost structure and latency profile than any API pricing negotiation ever could. That’s the real play here.




