TLDRocket
Sign in

As AI Increases Demands on Memory, Storage Steps Up

NVIDIA Jason Hardy

NVIDIA's opening up its cuFile storage APIs and pushing new GPU-to-storage tech at this week's FMS conference. AI agents now hammer storage directly, so the old rules about memory vs. disk are getting rewritten in microseconds, not minutes.

Storage used to be the boring part of computing. You wrote data down, you left it there, and eventually something fetched it. That model is breaking under the weight of AI agents, which don't wait patiently for data — they demand it in bulk, all at once, straight from the GPU. NVIDIA's pitch at this week's Future of Memory and Storage conference in Santa Clara is that the industry has to rethink storage as an active participant in the AI pipeline, not a warehouse sitting off to the side.

The numbers behind this shift are blunt. GPUs can now fire off thousands of concurrent storage requests on their own, and every one of those requests needs encrypting, compressing, verifying, and reconstructing on the fly. NVIDIA says its Vera CPU, part of the new Vera BlueField-4 STX platform, handles a two-stage compression-and-encryption pipeline up to 3.21 times faster than a comparable x86 chip. That's not a marginal tweak — it's the difference between storage keeping pace with AI factories or becoming the thing that throttles them.

The more interesting move, though, is NVIDIA open sourcing cuFile, the API layer (and the storage stack beneath it) that lets GPUs read and write directly to storage without routing everything through a CPU. It's part of GPUDirect Storage, and it's being folded into a broader effort called the Open Secure AI Alliance, with Google, Intel and Meta signed on as co-maintainers alongside NVIDIA. The logic here is straightforward: if every vendor builds a proprietary shortcut between GPU and disk, nothing talks to anything else, and security becomes a patchwork. Open APIs, in theory, fix that.

Alongside cuFile, NVIDIA is rallying more than 40 storage and flash companies — DDN, KIOXIA, Micron among them — into an initiative called Storage-Next, meant to hammer out shared standards for how GPU-driven storage should actually behave. The technical centerpiece is SCADA, a framework letting GPUs pull only the exact data they need straight into high-speed memory, skipping the usual detours. DDN is already wiring SCADA into its Infinia platform, and CTO Sven Oehme framed the goal less as buying more infrastructure and more as squeezing real productivity out of what's already deployed.

NVIDIA is careful to note that speed without guardrails is a liability, not a feature — letting applications talk directly to a drive can just as easily corrupt someone else's data as accelerate your own. SCADA's answer is to split the fast path from the trusted, privileged path that sets up permissions, keeping raw speed available without handing every application unsupervised access to the hardware. It's a decades-old tradeoff — speed versus safety, memory versus storage — just now playing out at GPU speed instead of spinning-disk speed.

My take

Open sourcing cuFile is the smart part of this announcement; standards fights over GPU-storage plumbing help nobody but the vendor who loses the standards war, and NVIDIA clearly knows that. The Storage-Next coalition, though, is really NVIDIA writing the rulebook and inviting 40 companies to sign it after the fact — which is a fine strategy, just not the neutral industry effort it's dressed up as. Fast, agentic AI storage is coming whether anyone likes it or not; the only real question is whether the security model keeps up, and history says security usually loses that race first.

Read more about this at: NVIDIA

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.