in

OpenAI Previews Ultrafast Mode, Running GPT-5.6 Sol Up to 14 Times Faster on Cerebras Hardware

OpenAI has opened a limited preview of Ultrafast mode, a new API service tier that runs its GPT-5.6 Sol model at up to 14 times the speed of standard processing. The tier is powered by Cerebras hardware and delivers up to 750 output tokens per second, and OpenAI says it runs the same model with the same intelligence rather than a smaller or distilled version. Access is currently waitlist only through the OpenAI API, and neither pricing nor a general availability date has been announced.

What Ultrafast actually is

The important detail is what has not changed. Ultrafast is not a new model. It is the existing GPT-5.6 Sol, OpenAI’s flagship reasoning and agentic model, running on different silicon. Cerebras builds wafer scale processors with a very different memory architecture to conventional GPU clusters, and the result is a latency profile that GPU based serving struggles to match for token by token generation.

Hosting 75% off

OpenAI published the announcement on its own site, Previewing Ultrafast mode, and Cerebras described its side of the engineering in a companion post. Coverage from Help Net Security put the practical framing simply: same answers, arriving much sooner.

Where speed changes the product, not just the benchmark

Frontier models have mostly competed on capability. Speed has been treated as a cost line rather than a feature. That works while AI is used the way a search box is used, where a few seconds of waiting is acceptable. It stops working the moment the model sits inside a live interaction.

OpenAI says preview customers are testing Ultrafast in coding, commerce, financial research and customer support. Those four have something in common. Each involves either a human waiting on screen or an agent chaining many model calls together, where per call latency compounds into something users notice. An agent that makes twenty calls at two seconds each feels broken. The same agent at 750 tokens per second feels like software.

The quiet story is the hardware

For most of the last few years, serving frontier models has effectively meant Nvidia GPUs. A flagship OpenAI model running in production on Cerebras hardware is a meaningful data point about how that might diversify, even if the tier stays small. Specialised inference silicon has been promising this for a while, and a limited preview is not a market shift, but it is a real deployment rather than a benchmark.

It is worth being precise about what is unconfirmed. OpenAI has not published pricing, has not committed to a launch date, and has not said how broadly the tier will scale. Fast inference on unusual hardware is often expensive inference. Until pricing appears, the economics are an open question.

What it means if you build or sell with AI

If you build products on top of these APIs, the takeaway is that latency is becoming a competitive dimension you can actually buy. Features you shelved because they felt sluggish, live coding assistance, real time support agents, interactive research tools, deserve a second look once a fast tier is generally available.

If you sell services rather than software, the shift is subtler. Faster models make more of the workflow interactive, which raises client expectations about turnaround. The freelancers and agencies who benefit are the ones who use the speed to compress delivery cycles rather than to produce more undifferentiated output.

For now, Ultrafast is a preview with a waitlist. The direction it points, that model quality and model speed are becoming separate things you choose between, is the part worth planning around.

Hosting 75% off

Written by Hisham Sarwar

WorkChest.com

That is all you ever need to know about me but let me warn you, freelancing for me is a journey, certainly not a destination :)

Samsung’s August 2026 Security Update Patches 56 Vulnerabilities Across Galaxy Devices

CodeRabbit Raises $143 Million Series C at a $1.5 Billion Valuation to Review AI-Written Code