OpenAI just shipped a speed tier that turns “waiting for the model” into a thing of the past. The new Ultrafast mode runs GPT-5.6 Sol at up to 750 tokens per second — roughly 14x faster than standard processing. Crucially, this is not a smaller, cheaper model wearing a fast-mode costume. It is the same flagship model, running on radically different hardware.
What changes at 750 tokens per second
At that throughput, the AI stops being a query-response tool and starts behaving like a colleague reading over your shoulder. OpenAI’s internal demos are all latency-critical: engineers reading logs and correlating recent code changes while an alert is still firing; trading desks scanning transactions and flagging anomalies before the market moves; e-commerce assistants answering product questions while a shopper is still deciding. Research tasks that used to run for minutes now unfold as a live, watchable conversation.
Early users back this up. Alex Wang, applied AI lead at financial research firm Rogo, says complex problems now feel like talking to a human in real time. Podium’s voice AI product lead Courtland Lykins notes that support calls no longer stall while the model spins.
The trick: put the weights on the chip
The speedup does not come from clever prompting. It comes from Cerebras’s wafer-scale engine. Model weights are loaded directly into 44GB of on-chip SRAM, so inference never has to shuttle data in and out of HBM the way it does on conventional GPUs. That bypasses memory bandwidth — the bottleneck that has quietly capped inference speed for years. Compute stopped being the constraint a while ago; moving data was.
This is also the first visible payoff of the OpenAI-Cerebras deal signed earlier this year: a multi-year agreement to deploy 750MW of wafer-scale systems, worth over $10 billion. Ultrafast is what that contract looks like when it ships.
A second feature: Computer History
The same release adds Computer History to the ChatGPT desktop app. It records interaction events — clicks, typing, shortcuts, window switches — without screenshots, audio, or screen recording. Ask “where was I working earlier?” and it answers from a timeline of your actual machine activity. It is a rebuild of the Chronicle research preview from April, and it can fold recurring workflows into reusable skills.
Why this matters
Two structural signals here. First, speed is becoming a product axis, not a spec-sheet footnote — and the winners are not shrinking models, they are changing the silicon underneath them. Second, memory is shifting from “what you told the chatbot” to “what you did on the computer,” which redefines how assistants earn context. Both push the industry toward the same end: AI that feels present rather than summoned.
What to do about it
- If you build latency-sensitive agents (support, trading, ops), benchmark at the 750 tok/s end — flows that were unusable at 50 tok/s become viable.
- Watch Cerebras-class inference pricing; wafer-scale is no longer a curiosity.
- Reconsider your memory UX: Computer History-style interaction logging will become table stakes, and privacy controls (no screenshots, per-item deletion) will be the differentiator.