How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
NVIDIA Blog Dion Harris ● Covered by 17 sources
NVIDIA’s Blackwell GPUs are now powering OpenAI’s GPT-6 Astra Ultrafast in the API. OpenAI says it can generate tokens up to 8x faster, which matters when agents are looping through code and tools.
Based on reporting by NVIDIA Blog, Dion Harris — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI has switched GPT-6 Astra Ultrafast onto NVIDIA Blackwell GPUs, and developers can use it now through the OpenAI API. Eligible ChatGPT Work and Codex users can access it too. The pitch is simple: faster output when the model has to keep up with the pace of a live workflow.
OpenAI says Ultrafast can generate tokens up to 8x faster than Astra Standard mode. That speedup is aimed squarely at the slow, repetitive parts of agent work — writing code, calling tools, checking results, then deciding what comes next. Shave time off each step, and the whole loop feels less like waiting around and more like a conversation.
NVIDIA says the gains come from inference optimizations that tap into Blackwell’s architecture. Philippe Tillet, OpenAI’s inference lead, said NVIDIA’s tooling and documentation helped the company make its models good at programming Blackwell and Rubin GPUs, and that Astra turns that into high-performance kernels. In his framing, the payoff is better latency, throughput and cost at the same time. That’s the kind of sentence hardware people write when they want to sound like they slept under a server rack.
The other interesting bit is that the work doesn’t stop once the model is shipped. OpenAI is using its own models to refine the inference software running on NVIDIA GPUs, leaning on the platform’s programmability to test and roll in improvements. Uday Ruddarraju, OpenAI’s chief technology officer of compute, said that helped deliver the acceleration behind Astra Ultrafast. NVIDIA also says its programmable platform lets teams reuse infrastructure across training, inference and reinforcement learning as models change, which can improve utilization and reduce overprovisioning.
For developers, the message is that the API is not just getting faster in a one-off way. NVIDIA and OpenAI are treating performance as something to keep tuning after launch, which is exactly how this stuff should work when agents are expected to do more than spit out a single answer. The Ultrafast guide has the access, pricing and implementation details.
My take — AI-written commentary, not fact-checked reporting
This is the real AI arms race: not bigger demos, just less waiting between tool calls. The model that feels snappiest inside a workflow will often beat the one that sounds smartest on a slide deck. Everyone loves frontier chatter; nobody misses staring at a spinner.
Read more about this at: NVIDIA Blog