AI

OpenAI previews Ultrafast mode for GPT-5.6 Sol in API

Thursday, August 13, 2026Read Original

Details

  • OpenAI is previewing Ultrafast, a new service tier for GPT-5.6 Sol that can run up to 14x faster than Standard processing.[2]
  • The first launch is in the OpenAI API and is limited to a select group of customers, with broader access planned as capacity grows.[2]
  • OpenAI says Ultrafast is powered by Cerebras and can generate up to 750 output tokens per second.[2]
  • The company is positioning the mode for products and workflows where lower latency creates a measurable business advantage.[2]
  • OpenAI says it is working with an initial customer group to learn where the speed matters most and to shape future product decisions.[2]
  • The announcement links this speed tier to GPT-5.6 Sol, OpenAI’s most intelligent model in this preview, rather than a separate lightweight model.[2]

Impact

Ultrafast pushes OpenAI further into the latency race, where speed becomes a product feature rather than just a backend metric. At up to 750 tokens per second, it raises the bar for enterprise workflows that need rapid model turns, and it likely pressures rivals offering faster inference tiers or low-latency model variants. The limited API preview suggests OpenAI is testing demand before broadening access, which could make high-end frontier intelligence more practical in interactive and agentic applications.

Rift Dispatch