AI

OpenAI’s fastest tier doesn’t run on GPUs

Ultrafast mode serves GPT-5.6 Sol at up to 750 output tokens per second — up to 14x standard processing — on Cerebras wafer-scale chips. It’s a limited preview for a handful of customers. The speed is the pitch. The supplier is the story.

N Noah · The Sharp Brief · August 14, 2026 · 4 min read

OpenAI previewed a new API service tier on Thursday called Ultrafast, running GPT-5.6 Sol at up to 14 times the speed of standard processing. Cerebras confirmed the same day that it is the silicon underneath, delivering up to 750 output tokens per second. Access is limited to a selected group of API customers and will widen as capacity allows.

Cerebras’ own benchmarks put some shape on it. Against output speeds reported by Artificial Analysis, it claims Sol on Ultrafast runs 11x faster than Claude Fable 5 and 5x faster than Claude Opus 4.8 on Fast mode. Running all 2,500 questions of Humanity’s Last Exam, it says Ultrafast finished in 11 hours 11 minutes where Fable 5 needed 78 hours 27 minutes — roughly 7x, at comparable accuracy. On GDP-Val, a benchmark built around economically valuable knowledge work, it reports a 5.6x end-to-end speedup with no quality degradation.

Every one of those comparisons was run by Cerebras. Treat them as a vendor claim until third parties reproduce them.

Why wafer-scale, and why now

Fast inference on a large model is a data-movement problem. On GPUs, weights shuttle between on-chip memory and off-chip storage for every token generated, and memory bandwidth becomes the ceiling. Cerebras packs 44GB of SRAM onto a wafer-sized chip so the weights stay put and tokens flow through layers pipelined across wafers. It is an architecture that has been looking for a marquee customer for years. It just got the biggest one available.

Our take: Latency is separating from intelligence and price as a third axis you can buy on, and OpenAI just validated a non-Nvidia architecture inside its own product line to serve it. That matters more than the token counts. Nvidia’s moat has always been that nobody else could serve frontier models at production scale; a limited preview isn’t a breach, but it is the first time a top lab has told its own customers that the fastest way to run its best model is on someone else’s chips. Watch the pricing when it appears — that is where you’ll learn whether this is a real product or a very good demo.

Where the seconds actually pay

Speed is worth nothing for an overnight batch job and worth a great deal on the critical path. Cerebras points at incident response, security operations under live attack, financial research, voice, commerce and interactive coding — all cases where a model that answers before you context-switch changes the shape of the work rather than just the cost of it. That is also the honest limit of the pitch: most enterprise AI spend today sits in asynchronous workloads where 14x buys nothing.

What to watch

Advertisement

Get the day, decoded — at 7 PM ET

The Sharp Brief: AI, money, business & performance in five sharp minutes. Free.

Free bonus: subscribe today and The 2026 AI Playbook (PDF) lands with your welcome email.

Recommended by 5+ newsletters across AI, markets & business.