logo-darklogo-darklogo-darklogo-dark
  • Tool Categories
    • 🎨Art & Creative Design505
    • 🏢Business Management644
    • 💻Coding & Development514
    • 👮Detection83
    • 🧠General Use728
    • 🏥Health & Wellness55
    • 📷Image & Photo Analysis100
    • 🖼️Image Generation & Editing618
    • 📐Interior & Architectural Design37
    • 🎓Learning & Education483
    • ⚖️Legal & Finance90
    • 🎭Lifestyle & Entertainment236
    • 📢Marketing & Advertising627
    • 🎧Music & Audio138
    • 👔Office & Workplace1,014
    • 🔬Research & Data Analysis373
    • 👥Social Media245
    • 🎥Video Generation & Editing426
    • 👧🏻Virtual Companion135
    • 🎤Voice Generation & Editing381
    • ✍️Writing & Editing808
    • All Categories
    • AI Use Cases
  • News
  • Events
    • Academic Conferences
    • Developer Conferences
    • Expos / Trade Shows
    • Industry Summits
    • Workshops / Training
    • All Events
    • Past Events
  • Saved Tools
  • Suggest a Tool
✕
Home › News › OpenAI’s Ultrafast mode hits 750 tokens per second with GPT-5.6 Sol

OpenAI’s Ultrafast mode hits 750 tokens per second with GPT-5.6 Sol

August 13, 2026
OpenAI’s Ultrafast mode hits 750 tokens per second with GPT-5.6 Sol

Most speed improvements in AI come with a trade-off: you get a faster model, but a less capable one. OpenAI is trying to break that pattern. The company announced Ultrafast, a new service tier that runs GPT-5.6 Sol at up to 14 times the speed of standard processing, generating up to 750 output tokens per second. That number puts it in a different category from anything else on the API market today.

Ultrafast is powered by Cerebras, the chip company whose wafer-scale processors are built specifically for high-throughput inference. This is an extension of an existing partnership, but the scope has expanded. Before, Cerebras was supporting smaller or specialized models. Now it’s handling GPT-5.6 Sol, which OpenAI positions as its most capable model in the current lineup. Pushing a frontier model through that kind of silicon and hitting 750 tokens per second is a meaningful technical step.

Why speed at this level changes what’s possible

There’s a real difference between a model that responds quickly and one that keeps pace with human thought in real time. At standard speeds, even capable models introduce enough latency that certain workflows are just impractical. Incident response is a good example. When a production system is failing, engineers are working in a fast-moving, high-pressure environment. A model that takes several seconds per response breaks the loop. One that answers in near real time can sit alongside an engineer as they read logs, trace errors, and form hypotheses.

OpenAI says its internal teams have already tested Ultrafast on exactly that use case, and also on research workflows where experiment cycles that used to run overnight can now be iterated during a single workday. These aren’t hypothetical examples. They’re the kinds of workflows where latency has historically been the limiting factor, not model quality.

Target use cases

OpenAI has outlined several areas where Ultrafast is expected to create the most value:

  • Incident response and reliability, where logs, traces, and engineer notes need to be synthesized while an outage is still active
  • Financial research and transaction monitoring, where market conditions and risk signals change by the second
  • Customer support and voice interfaces, where multi-step answers need to arrive without breaking conversational flow
  • Commerce, where product questions, inventory checks, and checkout issues need resolution before a shopper leaves the page
  • Live research and experimentation, compressing overnight batch runs into interactive working sessions

Early preview customers include Jane Street, Podium, Basis, and Rogo. That’s a mix of fintech, commerce, and developer tooling, which maps neatly onto the use cases above. John Crepezzi from Jane Street’s AI Assistants team noted that the speed makes it practical to work alongside models in a more focused way. That’s a quiet but telling observation: speed doesn’t just make the same workflows faster, it changes how developers and knowledge workers interact with models at all.

How this stacks up against the competition

Groq has been the go-to name in high-speed inference for a while, offering fast token generation on open-weight models. But Groq doesn’t run GPT-5.6 Sol. Anthropic’s Claude models are fast in standard tiers but aren’t positioned around raw token throughput. Google’s Gemini Flash is the closest competitor on the speed-with-intelligence angle, but OpenAI is now directly challenging that positioning by running its flagship model at speeds that were previously only associated with smaller distilled variants.

The Cerebras partnership is also strategically important. OpenAI is building a dependency on specialized inference hardware that could give it a structural speed advantage that’s hard for API resellers or smaller providers to replicate. That matters if OpenAI wants Ultrafast to become a durable tier, not just a preview feature.

Availability and access

GPT-5.6 Sol on Ultrafast mode is currently in limited preview through the OpenAI API. Access is restricted to a select group of customers while OpenAI builds out capacity. A sign-up option is available for teams that want to be notified when broader access opens. Pricing has not been publicly disclosed. Given the infrastructure involved, expect a meaningful premium over standard API rates.

For developers building latency-sensitive applications, this is worth watching closely. If OpenAI expands access at a reasonable price point, Ultrafast could shift what’s considered acceptable response time across the industry.

Share

Related news

Writer launches Palmyra X6 and updated agentic infrastructure to cut enterprise AI costs by 50%
August 13, 2026

Writer launches Palmyra X6 and updated agentic infrastructure to cut enterprise AI costs by 50%


Read more
IBM and OpenAI strike enterprise deal, and it’s bigger than it looks
August 13, 2026

IBM and OpenAI strike enterprise deal, and it’s bigger than it looks


Read more
Anthropic let AI agents fight over the same codebase — here’s what happened
August 13, 2026

Anthropic let AI agents fight over the same codebase — here’s what happened


Read more

Recent Posts

  • OpenAI’s Ultrafast mode hits 750 tokens per second with GPT-5.6 Sol
  • Writer launches Palmyra X6 and updated agentic infrastructure to cut enterprise AI costs by 50%
  • IBM and OpenAI strike enterprise deal, and it’s bigger than it looks
  • Anthropic let AI agents fight over the same codebase — here’s what happened
  • Microsoft merges its Copilot apps and cuts the features nobody used
Best AI Tools

Discover the best AI tools for any use case

Explore
  • Tool Categories
  • AI Use Cases
  • AI Events
  • AI News
  • Saved Tools
Company
  • About Us
  • Contact Us
  • Media & Partnerships
  • Suggest a Tool
Legal
  • Privacy Policy
  • Terms of Service
Copyright © 2026 Best AI Tools 415 Mission Street, 37th Floor, San Francisco, CA 94105