OpenAI's Ultrafast Mode Runs GPT-5.6 Sol at 750 Tokens Per Second
OpenAI added a new processing mode to GPT-5.6 Sol on August 13, 2026. It is called Ultrafast. The mode delivers up to 750 output tokens per second, which OpenAI says is 14x the speed of standard processing. Access is in preview, limited for now.
The Cerebras Connection
Ultrafast runs through OpenAI's partnership with Cerebras, a chipmaker. That is the full extent of what OpenAI has disclosed about the hardware arrangement. The speed figures are real in preview conditions. Real-world performance at scale is a separate question.
Who Actually Needs 750 Tokens Per Second
OpenAI's stated use cases: incident response, customer service, financial market analysis, e-commerce. These share a common trait. They all have latency requirements where slower responses are not just inconvenient but functionally useless. A financial analysis tool that processes market data in near-real-time is a different product category than one that does not. Ultrafast is aimed at that gap.
The Expansion Caveat
OpenAI says it will expand access as capacity grows. That phrasing suggests the infrastructure behind Ultrafast is constrained. What exactly is constrained, and how quickly capacity will grow, OpenAI has not specified. Preview-to-general-availability timelines in AI infrastructure tend to stretch.
Source: Techcrunch