logo-darklogo-darklogo-darklogo-dark
  • Tool Categories
    • 🎨Art & Creative Design505
    • 🏢Business Management644
    • 💻Coding & Development514
    • 👮Detection83
    • 🧠General Use728
    • 🏥Health & Wellness55
    • 📷Image & Photo Analysis100
    • 🖼️Image Generation & Editing618
    • 📐Interior & Architectural Design37
    • 🎓Learning & Education483
    • ⚖️Legal & Finance90
    • 🎭Lifestyle & Entertainment236
    • 📢Marketing & Advertising627
    • 🎧Music & Audio138
    • 👔Office & Workplace1,014
    • 🔬Research & Data Analysis373
    • 👥Social Media245
    • 🎥Video Generation & Editing426
    • 👧🏻Virtual Companion135
    • 🎤Voice Generation & Editing381
    • ✍️Writing & Editing808
    • All Categories
    • AI Use Cases
  • News
  • Events
    • Academic Conferences
    • Developer Conferences
    • Expos / Trade Shows
    • Industry Summits
    • Workshops / Training
    • All Events
    • Past Events
  • Saved Tools
  • Suggest a Tool
✕
Home › News › OpenAI’s Jalapeño chip beats Nvidia Blackwell on inference benchmarks, but the real test comes later

OpenAI’s Jalapeño chip beats Nvidia Blackwell on inference benchmarks, but the real test comes later

August 25, 2026
OpenAI’s Jalapeño chip beats Nvidia Blackwell on inference benchmarks, but the real test comes later

OpenAI is benchmarking its custom silicon against Nvidia’s best, and according to early numbers, it’s winning. At the Hot Chips conference this week, OpenAI shared the first public benchmark results for Jalapeño, its in-house inference chip built with Broadcom. As TechCrunch reported, those results show Jalapeño outperforming current state-of-the-art inference processors on both tokens per user and throughput per kilowatt, tested against Semianalysis’s InferenceX benchmark.

That’s a meaningful combination to beat. Tokens per user reflects how well the chip handles many simultaneous requests, while throughput per kilowatt speaks directly to operating cost at scale. Richard Ho, OpenAI’s head of hardware, described it plainly on a press call: “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”

The catch is the comparison. Jalapeño’s benchmarks are measured against an Nvidia Blackwell system, which is the current generation. By the time Jalapeño reaches meaningful deployment, Nvidia will almost certainly have moved forward. Ho placed full-scale rollout in 2027, with only “very small volumes” shipping by end of 2026. That’s a long window for a competitor like Nvidia, or Google with its TPUs, or Amazon with Trainium, to close any gap.

Still, the underlying architecture tells a more interesting story than the headline numbers. Jalapeño was designed specifically to reduce friction at the prefill and communication phases of inference, two spots that regularly create bottlenecks in production systems. OpenAI says the chip keeps model state, including the KV cache, local and explicitly placed, so the system can activate the right mix of compute, memory, and networking for each inference phase without unnecessary data movement. That’s not a general-purpose design choice. It’s targeted at the specific workloads OpenAI runs at scale.

That full-stack thinking is what separates this from a typical chip announcement. OpenAI built Jalapeño as a multigenerational platform, with models, chips, memory, and AI products all developed together. This mirrors what Google has done with TPUs for years, and what Meta is attempting with MTIA. The difference is OpenAI is doing this while also being one of the largest external customers of Nvidia. Jalapeño isn’t a replacement yet. It’s a hedge, and a signal of where OpenAI wants to be in three to five years.

For developers and infrastructure teams, the near-term impact is limited. But the direction matters. If Jalapeño delivers on efficiency at scale, it gives OpenAI more control over its cost structure and latency profile than any API pricing negotiation ever could. That’s the real play here.

Share

Related news

Stability AI raises $76M from Sony, Universal, and Warner in a bet on creative AI
August 25, 2026

Stability AI raises $76M from Sony, Universal, and Warner in a bet on creative AI


Read more
Claude’s memory now spans chats and Cowork sessions, with a new sensitive topics toggle
August 25, 2026

Claude’s memory now spans chats and Cowork sessions, with a new sensitive topics toggle


Read more
Chinese state hackers are using DeepSeek to double their attack output
August 25, 2026

Chinese state hackers are using DeepSeek to double their attack output


Read more

Recent Posts

  • Stability AI raises $76M from Sony, Universal, and Warner in a bet on creative AI
  • Claude’s memory now spans chats and Cowork sessions, with a new sensitive topics toggle
  • OpenAI’s Jalapeño chip beats Nvidia Blackwell on inference benchmarks, but the real test comes later
  • Chinese state hackers are using DeepSeek to double their attack output
  • Anthropic’s most powerful model is losing the price war before the company even goes public
Best AI Tools

Discover the best AI tools for any use case

Explore
  • Tool Categories
  • AI Use Cases
  • AI Events
  • AI News
  • Saved Tools
Company
  • About Us
  • Contact Us
  • Media & Partnerships
  • Suggest a Tool
Legal
  • Privacy Policy
  • Terms of Service
Copyright © 2026 Best AI Tools 415 Mission Street, 37th Floor, San Francisco, CA 94105