Skip to content

OpenAI's New 'Ultrafast' Mode Shrinks Agent Response Times by Up to 8x

Short answer

OpenAI launched GPT-6 Astra Ultrafast, a faster inference mode on NVIDIA Blackwell GPUs offering up to 8x quicker token generation than Astra Standard. It's live in the OpenAI API and for eligible ChatGPT Work and Codex users, aimed at shortening the wait time inside agent loops that write code, call tools, and act on results.

What this means for operators

If your support or ops stack runs AI agents that chain multiple tool calls per customer interaction — pulling a CRM record, checking inventory, drafting a reply — the delay between each step compounds into real wait time for the end user. Ultrafast mode targets exactly that loop: faster generation between tool calls means a multi-step agent finishes its reasoning before a human notices the pause. Before switching production workflows over, check the Ultrafast guide for which plans and endpoints currently have access and what the pricing difference is versus Astra Standard — the source doesn't specify cost, so that's a question to put to your API rep, not an assumption to build on.

OpenAI has released GPT-6 Astra Ultrafast, a faster inference mode of its Astra model running on NVIDIA Blackwell GPUs. It's available now through the OpenAI API and to eligible ChatGPT Work and Codex users.

The headline figure: Ultrafast delivers up to 8x faster token generation than Astra's Standard mode. OpenAI frames the gain around agentic workflows — coding agents running edit-test-debug cycles, and any application where a model writes code, calls a tool, checks the result, and decides what to do next. Shortening each generation step in that loop shortens the whole cycle, since the delay repeats every time the agent takes an action.

Philippe Tillet, OpenAI's inference lead, said NVIDIA's tooling and documentation let OpenAI's models become "exceptionally good at programming Blackwell and Rubin GPUs," turning that into high-performance kernels that improve latency, throughput and cost on NVIDIA hardware. Uday Ruddarraju, OpenAI's chief technology officer of compute, said the company used its own models to optimize the inference software running on NVIDIA GPUs, with NVIDIA's programmable platform enabling the acceleration behind Ultrafast.

OpenAI also notes that performance work continues after deployment: it's using its own models to keep refining inference software on the same GPU platform, and that a programmable NVIDIA stack lets teams reuse infrastructure across training, inference and reinforcement learning rather than provisioning separately for each.

Developers can access GPT-6 Astra Ultrafast through the API now; OpenAI points to its Ultrafast guide for access, pricing and implementation details.

Source: NVIDIA Blog

Next step

Visibility Analyzer

This is what the Visibility Analyzer measures on a real site: which answers cite you, which pages an engine cannot retrieve, and what to fix first. Free to run.

Run a free visibility audit

Free to run. No card.

Fee
Free
Length
One run, minutes

Free tier: two analyses a day, no card required.