OpenAI previewed a Cerebras-powered Ultrafast tier for GPT-5.6 Sol.
OpenAI said on August 13 that Ultrafast, a limited-preview API tier for GPT-5.6 Sol, runs up to 14 times faster than Standard processing and can generate up to 750 output tokens per second; Cerebras said it is powering the service, and TechCrunch reported access is initially limited to a small customer group.
What happened
OpenAI framed Ultrafast as a new speed class for frontier intelligence, aimed at workflows where response time determines whether a model can stay inside an operational loop. Cerebras said its system powers the tier and repeated the 750-output-token-per-second claim. TechCrunch independently reported that the preview is limited now and that OpenAI plans to expand access as capacity grows.
Why it matters
The practical model race is no longer only about benchmark ceilings. If high-end reasoning can be served fast enough, more agent workflows can move from asynchronous reports into interactive products, call-center tooling, coding loops, incident response and other time-sensitive work.
What to watch
Whether independent users see the advertised latency under real workloads; whether quality changes under speed pressure; how OpenAI prices scarce Ultrafast capacity; and whether Cerebras turns the partnership into broader evidence that wafer-scale inference can compete with GPU clusters.
The caveat
OpenAI and Cerebras are interested primary sources for the speed and quality claims. The independent evidence so far confirms the product preview and limited access, not broad customer performance.
