TLDRocket
Sign in

Vast uses tiered storage to ease AI agent memory demands

SiliconANGLE Sloane Kali Faye

Vast says AI agents need more memory than GPUs can hold. Its fix is tiered storage, so long sessions don’t keep redoing work.

Based on reporting by SiliconANGLE, Sloane Kali Faye — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AI agents are starting to stress the plumbing underneath them. According to Vast co-founder and CTO Alon Horev, these systems are running longer, holding onto more context, and needing shared memory that survives across sessions. That creates a new problem: not just how to answer, but how to remember.

Horev laid out the issue at Fully Connected 2026 in a conversation with theCUBE Research’s John Furrier and Dave Vellante. The pressure shows up in inference first. A long session keeps its KV cache in GPU memory, and Horev said a session with half a million tokens can eat up one-tenth to one-twentieth of a GPU’s memory. That is not a side detail. It becomes the bottleneck.

Vast’s answer is to move memory through tiers. GPU memory comes first, then CPU memory on the same machine, then persistent media that can store petabytes of KV cache. Nvidia’s Dynamo software helps orchestrate the handoff. In Horev’s telling, that lets a session move from a busy GPU to a less busy one, with the cache read over the network or pulled from Vast instead of being rebuilt from scratch.

That matters because the workload is no longer a single chat window. Enterprises are deploying thousands of agents, and those agents may pause while code compiles, software is tested, or a user steps away for coffee. If the memory can survive that gap, the system avoids repeat calculation and gets more room to schedule work across a fleet of machines.

There is also a governance angle. Companies need to record what their agents do and keep it for a set period, especially when the agents are handling sensitive data or acting on customers’ behalf. Horev called those conversations “gold” because they can feed fine-tuning, training, or purpose-built models later. Vast has also launched a confidential computing service for sensitive workloads, which fits the same theme: memory is now both an operational problem and a control point.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI infrastructure people keep pretending is glamorous when it’s really storage wearing a nicer jacket. The model hype gets the attention, but the winners will be the vendors who make memory behave across GPUs, CPUs, and whatever comes after the GPU budget panic. Open or closed, it doesn’t matter much if the agent forgets everything the minute the tab stops moving.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.