OpenAI has introduced an early preview of Ultrafast, a new API service tier that runs GPT-5.6 Sol as much as 14 times faster than Standard processing. Powered by Cerebras, the service can generate up to 750 output tokens per second and is initially available to a limited group of customers.
Ultrafast changes how OpenAI serves GPT-5.6 Sol rather than introducing a separate model. The company is using Cerebras infrastructure to reduce the time required for its most capable model to respond, targeting applications where latency can determine whether AI fits naturally into an existing workflow.
The approach is designed to address a longstanding compromise between model capability and responsiveness. Developers seeking real-time performance have typically had to use smaller or more specialized models. OpenAI is positioning Ultrafast as a way to retain GPT-5.6 Sol's capabilities while substantially increasing inference speed. That distinction could be particularly important in workflows requiring repeated interaction with a model. OpenAI identified incident response, financial research, security, customer support, voice applications, commerce and live experimentation as early areas where it sees potential for the faster tier.
In incident response, for example, engineers can use Ultrafast to process logs and traces, review conversations and identify possible next steps while an issue is still developing. OpenAI said its own developers have been testing the service for this type of work, including helping prepare and validate potential fixes while leaving deployment decisions with engineers.
The company is also experimenting with Ultrafast in research workflows. OpenAI said its teams can use the service to search information sources, query data and organize findings across connected tools more quickly. The company sees the added speed potentially compressing experimentation cycles that previously involved running batches overnight and reviewing the results later.
OpenAI has also been testing GPT-5.6 Sol through Ultrafast with companies working across coding, commerce, financial research, customer support and other interactive applications. Jane Street, Podium, Basis and Rogo are among the early participants identified by OpenAI.
“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them,” John Crepezzi of Jane Street's AI Assistants team said.
Cerebras provides the infrastructure behind the performance increase. OpenAI described Ultrafast as an expansion of its existing partnership with the chip company, with Cerebras now supporting GPT-5.6 Sol at speeds reaching 750 output tokens per second.
For OpenAI, the preview also provides a way to study what changes when its highest-end intelligence becomes significantly more responsive. Rather than treating inference speed solely as a performance metric, the company is evaluating whether faster responses alter how developers design products and how users interact with them.
The initial focus on business applications gives OpenAI production environments in which to test that question. Feedback from early customers will help inform how the service develops as additional capacity becomes available.
GPT-5.6 Sol Ultrafast is currently in limited preview for select customers through the OpenAI API. OpenAI said it plans to broaden availability as capacity increases, while businesses interested in future access can register for updates.
About this article: This article was generated with AI assistance and reviewed by our editorial team to ensure it follows our editorial standards for accuracy and independence. We maintain strict fact-checking protocols and cite all sources.
Word count: 547Reading time: 0 minutes
Explore More AI Resources
Continue with high-value guides related to this topic.
Join thousands of weekly readers — the latest AI news, free in your inbox.
🤖 3 times a week📊 Industry analysis💡 Breaking news
Enjoying this article?
Get a free month of ChatAI Plus or Pro
🎁 Limited time offer
🔒 Once per user
ChatAI Plus & Pro unlock multi-model chat (GPT, Claude, Gemini & more), our productivity tools, and ad-free reading. Subscribe to our newsletter and take a 2-minute survey to claim one month free — no strings attached.