An NVIDIA post dated October 1 says GPT-6 Astra Ultrafast runs on NVIDIA Blackwell GPUs and is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. The speedup comes from inference optimizations OpenAI made on Blackwell. The post says token generation can be up to 8 times faster than Astra Standard mode. It does not say which prompts, context lengths, or batch sizes were measured, and it does not give a typical speed beside the “up to” figure.

[1]
Silverpoint: a short scrap with one thick black stroke, and a long strip with a pale line.
The upper scrap is short and holds one thick black stroke. The lower strip is long and carries a pale line. That is the same kind of generation stopping sooner. An illustration, not a benchmark chart., AI-generated illustration, not a news photograph

The post places the faster generation in a loop that repeats: an agent writes code, calls a tool, checks the result, and decides the next step. It says Ultrafast puts Astra into those time-sensitive loops, and that faster generation can shorten an edit-test-debug cycle and the wait between tool calls. That is a description of the use, not a published timing on a named task.

Philippe Tillet, inference lead at OpenAI, says NVIDIA’s work on tooling and documentation has made their models unusually good at programming Blackwell and Rubin GPUs, and that Astra can turn that knowledge into high-performance kernels across latency, throughput, and cost. Rubin in that sentence is hardware they program. The post says Ultrafast is running on Blackwell. It does not say Ultrafast is already running on Rubin.

[1]

The post says performance does not stop at deployment. OpenAI is using its own models to refine the inference software on NVIDIA GPUs, using the platform’s programmability to try changes. Uday Ruddarraju, chief technology officer of compute at OpenAI, says they used internal models to optimize inference on NVIDIA GPUs, and that programmability helped deliver the acceleration behind Ultrafast. That is how the two companies describe work already done. The post does not say how much faster a later round will be.

A programmable platform, the post says, lets the same infrastructure move across training, inference, and reinforcement learning, so compute can be reassigned as demand changes. No utilization figure is attached.

Developers can use GPT-6 Astra Ultrafast through the API today. Access, pricing, and implementation details are in an Ultrafast guide. This post does not state a price.

[1]

要点

  • GPT-6 Astra Ultrafast is in the OpenAI API and available to eligible ChatGPT Work and Codex users, running on Blackwell.
  • Token generation is described as up to 8 times faster than Astra Standard. The post gives no task, context length, or batch size.
  • Tillet mentions kernels for Blackwell and Rubin. The post says the mode now running is on Blackwell, not that it already runs on Rubin.
  • Price and setup are in a separate guide. This post states no price and no further speedup.