logo-darklogo-darklogo-darklogo-dark
  • Tool Categories
    • 🎨Art & Creative Design505
    • 🏢Business Management644
    • 💻Coding & Development514
    • 👮Detection83
    • 🧠General Use728
    • 🏥Health & Wellness55
    • 📷Image & Photo Analysis100
    • 🖼️Image Generation & Editing618
    • 📐Interior & Architectural Design37
    • 🎓Learning & Education483
    • ⚖️Legal & Finance90
    • 🎭Lifestyle & Entertainment236
    • 📢Marketing & Advertising627
    • 🎧Music & Audio138
    • 👔Office & Workplace1,014
    • 🔬Research & Data Analysis373
    • 👥Social Media245
    • 🎥Video Generation & Editing426
    • 👧🏻Virtual Companion135
    • 🎤Voice Generation & Editing381
    • ✍️Writing & Editing808
    • All Categories
    • AI Use Cases
  • News
  • Events
    • Academic Conferences
    • Developer Conferences
    • Expos / Trade Shows
    • Industry Summits
    • Workshops / Training
    • All Events
    • Past Events
  • Saved Tools
  • Suggest a Tool
✕
Home › News › OpenAI’s Jalapeño chip posts numbers that should worry Nvidia

OpenAI’s Jalapeño chip posts numbers that should worry Nvidia

August 28, 2026
OpenAI’s Jalapeño chip posts numbers that should worry Nvidia

OpenAI built a chip, and it actually works. That’s not a given in this industry. Custom silicon projects from big tech companies have a long history of ambitious announcements followed by quiet shelving. So when OpenAI published the first real benchmark results for Jalapeño, its custom inference accelerator, the numbers deserved a hard look rather than a press release skim.

The headline figures are striking. Tested against GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T on the InferenceX public benchmark from SemiAnalysis, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput. End-to-end latency came in 1.7 to 3.6 times lower than the comparison systems. For highly interactive workloads, the performance gap widened to 2.1 to 4.1 times. These aren’t cherry-picked internal metrics. The team used a public benchmark across three models, two of which were built outside OpenAI entirely.

What the architecture actually does differently

The core insight behind Jalapeño’s design is that inference has two very different phases. Prefill, where the system processes an input prompt, is compute-heavy. Decode, where it generates tokens one by one, is bottlenecked by memory bandwidth. Most existing hardware is optimized for one or the other, or forces engineers to make tradeoffs between throughput and latency. Jalapeño was designed to handle both without that compromise.

The key is minimizing data movement. Model state, including the KV cache generated during decode, stays local rather than bouncing between chips or memory pools. The network is built into the architecture at the rack level, so the entire workload stays within a single connected system. The result is what OpenAI describes as a fungible accelerator that can shift resources between prefill and decode as workloads demand, which matters a lot for agentic tasks where those demands shift constantly mid-run.

The chip is rated at 700 watts but ran at or below 550 watts sustained on the tested workloads. That gap between rated and actual power draw directly improves the efficiency numbers being reported, and it’s a meaningful real-world advantage for anyone paying data center power costs.

How AI helped build the chip that runs AI

One of the more interesting details here is the development process itself. OpenAI used earlier model generations to help design and bring up Jalapeño, and is now using its latest models to optimize and program it. The team went from initial design to tapeout in nine months, a fast timeline for custom silicon. AI helped explore chip implementations, shorten verification loops, and optimize arithmetic circuits to fit more compute into the design on schedule.

The programming model was built with AI in mind from the start. Engineers and AI systems alike can describe work through local tensors, explicit communication, and predictable synchronization. That structure gives AI a tractable surface for parallel programming optimization, which has historically been one of the hardest problems in chip software.

Why this matters for the market

Jalapeño is aimed squarely at the economics of inference at scale. The relevant comparison isn’t just Nvidia H100s or B200s. It’s also Google’s TPUs, which have been running internal inference workloads for years, and AWS Trainium, which is still finding its footing in serving workloads. OpenAI’s angle is vertical integration: designing models, serving software, chips, memory, networking, and systems together in a single stack. That’s the same playbook Google used to make TPUs formidable, and it works when you have enough of your own workload to tune against.

  • 1.5 to 1.9x higher throughput per watt at peak load
  • 1.7 to 3.6x lower end-to-end latency than existing accelerators
  • 2.1 to 4.1x better performance on interactive, low-latency workloads
  • Sustained power at or below 550W against a 700W rating
  • Nine-month design-to-tapeout timeline with AI-assisted development

OpenAI says Jalapeño is the beginning of a multigenerational platform, not a one-off project. The internal results on frontier OpenAI models apparently show an even wider advantage than the public benchmark numbers, which suggests the chip is tuned for the workloads OpenAI actually runs. For developers and enterprises buying inference capacity through OpenAI’s API, this could mean lower costs and faster response times as Jalapeño scales. For Nvidia, it means its biggest customer is now also a serious silicon competitor.

Share

Related news

X says a Chinese bot farm used ChatGPT to stoke fears about AI data centers
August 28, 2026

X says a Chinese bot farm used ChatGPT to stoke fears about AI data centers


Read more
A federal judge just handed Anthropic a major legal win against the Pentagon
August 27, 2026

A federal judge just handed Anthropic a major legal win against the Pentagon


Read more
OpenAI, Google, and 100+ companies sign open letter warning AI cyber threats are getting worse
August 27, 2026

OpenAI, Google, and 100+ companies sign open letter warning AI cyber threats are getting worse


Read more

Recent Posts

  • OpenAI’s Jalapeño chip posts numbers that should worry Nvidia
  • X says a Chinese bot farm used ChatGPT to stoke fears about AI data centers
  • A federal judge just handed Anthropic a major legal win against the Pentagon
  • OpenAI, Google, and 100+ companies sign open letter warning AI cyber threats are getting worse
  • Anthropic poaches Google’s TPU architect to build its own AI chips
Best AI Tools

Discover the best AI tools for any use case

Explore
  • Tool Categories
  • AI Use Cases
  • AI Events
  • AI News
  • Saved Tools
Company
  • About Us
  • Contact Us
  • Media & Partnerships
  • Suggest a Tool
Legal
  • Privacy Policy
  • Terms of Service
Copyright © 2026 Best AI Tools 415 Mission Street, 37th Floor, San Francisco, CA 94105