OpenAI's GPT-6 Astra Ultrafast is now officially available on NVIDIA Blackwell GPUs. Developers can access it via API, and eligible ChatGPT Work and Codex users can use it as well.

This version is all about speed. According to NVIDIA, Astra Ultrafast can generate tokens up to 8 times faster than the standard version. It is primarily aimed at agentic use cases such as coding and tool calling, with the goal of significantly reducing wait times when AI is writing code, calling tools, and handling complex tasks simultaneously. It is worth noting that the 8x figure comes from the company’s own claims, and no independent third-party tests have been reported yet. Specific pricing, latency across different tasks, and eligibility requirements for ChatGPT Work users have also not been disclosed.

The speed boost comes from OpenAI’s inference optimizations for the Blackwell architecture. One particularly interesting point is that OpenAI is using its own models to help improve inference software running on NVIDIA hardware, including writing high-performance compute kernels. OpenAI’s head of inference said that, with the help of NVIDIA’s tools and documentation, they have optimized their models for Blackwell and the next-generation Rubin GPUs.

Why does this matter? Over the past year or two, the race among large models was about “who’s smarter.” Now, it’s increasingly about “who’s faster and cheaper.” Agentic tasks often require dozens or even hundreds of model calls, so even a small speedup on each call can make a big difference to the overall experience. For businesses, speed directly affects cost and availability.

For NVIDIA, this is also a great showcase: the top model companies are using its latest hardware to squeeze out more performance, further strengthening its software and hardware ecosystem moat. This is consistent with NVIDIA’s shares closing at a record high on Monday.

My take: “inference as a service” is becoming AI’s main battleground. Whoever can bring down inference costs and latency will win the gateway to the agent era. Next, keep an eye on the pricing details OpenAI announces and whether competitors launch similar high-speed versions.

The content above is for informational purposes only and does not constitute investment advice.

#OpenAI #英伟达 #AI