Multi-tier storage rewrites the economics of AI inference
SiliconANGLE Mark Albertson
AI inference is pushing companies toward multi-tier storage instead of one giant bucket. The real win is cheaper GPU time and less data shuffling.
Based on reporting by SiliconANGLE, Mark Albertson — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
AI training used to get most of the attention. Now the money is in inference, and that shift is forcing a rethink of storage from top to bottom. The industry answer, at least from the panel on Supermicro’s Open Storage Summit, is not one silver bullet but a stack: flash, object storage and disk each doing the job they’re good at.
Paul McLeod of Supermicro said the old problem of huge monolithic systems looks different once the workload gets smaller and more distributed. That’s where software-defined storage partners come in, he said, especially for keeping AI agents’ key-value cache around long enough to be useful again later. The basic goal is blunt: keep the data close enough, long enough, and cheaply enough that the GPU isn’t sitting around waiting.
Intel’s Angela Gill made the case for pushing more work into hardware, pointing to QuickAssist Technology in select Intel processors. The idea is to move compression and encryption out of software, free up CPU cycles, cut latency and improve power efficiency. She put it plainly: storage can’t stay in the slow lane if it has to move at compute speed.
Western Digital’s Marc Tanguay framed the problem as one of real estate as much as capacity. Enterprises can’t just keep adding racks, he said. His Ultrastar line is at a 28 terabyte CMR hard drive that’s already been qualified for Supermicro’s products, with 30 terabytes coming before the end of this year.
The more interesting part is how the vendors are trying to keep GPUs busy without throwing more hardware at the wall. Scality is combining flash with S3 and GPU-direct storage access so training and inference pipelines can stream straight into GPU memory. Samsung Semiconductor is tackling the KV cache bottleneck with memory expansion, including the PM1723, a Gen 6 drive that the company says reaches up to 28.4 gigabytes per second of sequential read throughput and up to 6.6 million random read IOPS. WekaIO’s Anthony Lembo cut through the noise: the hard part is balancing cost, performance and the simple fact that data movement is expensive.
My take — AI-written commentary, not fact-checked reporting
This is the boring truth of AI infrastructure: the winners won’t be the teams that buy the most GPUs, but the ones that stop treating storage like an afterthought. Everyone loves shiny accelerators until the cache spills and the bill arrives. The sector has spent years worshipping scale; now it’s being forced to respect plumbing.
Read more about this at: SiliconANGLE