Storage leaders target AI’s unstructured data challenge
SiliconANGLE Mark Albertson
AI’s biggest data problem is still the messy stuff: docs, videos, logs, emails. Storage vendors are racing to move and govern it before AI can use it.
Based on reporting by SiliconANGLE, Mark Albertson — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
The storage crowd has found AI’s dirtiest secret: the best fuel is usually the least organized. Industry estimates cited by theCUBE Research say more than 80% of enterprise data is unstructured, and that giant pile of documents, videos, logs, emails, images and audio is still mostly hard to reach, hard to govern and hard to use.
That is now a storage problem as much as an AI problem. On SiliconANGLE’s theCUBE, Molly Presley of Hammerspace said managing unstructured data for AI is far tougher than the old archive-and-backup routine. The new job is to unify the data and move it automatically, fast enough for AI workflows that actually need it.
Supermicro, Cloudian, Hammerspace and Seagate all described different parts of that pipeline. Sherry Lin said Supermicro is building high-density hardware for software-defined storage partners, and has added Context Memory storage servers to help offload and share large language model key-value caches across AI inference systems. Lin also outlined a split where Hammerspace handles the global namespace and orchestration, Cloudian provides the S3 object store for AI data, and Seagate’s hard drives sit at the end of the line.
Cloudian’s pitch is simple: keep unstructured data under control, protected and secure, even as it moves for different uses. Peter Sjoberg said that control matters because the same data may need to be pulled for AI work, compliance audits, debugging or to understand model outputs. Once the data goes into the AI era, it stops being just storage and becomes something that has to stay governed.
Seagate’s Mohamad El-Batal pushed the same idea from the infrastructure side. He said companies need to understand the value of the data and where they are trying to go with it before choosing how to build. His advice was blunt: do not go cheap on infrastructure, but do keep it efficient. Presley added that Hammerspace is working on a Model Context Protocol layer to help AI systems understand data that is not well described, because that is the nature of unstructured data.
My take — AI-written commentary, not fact-checked reporting
This is the part of AI nobody gets to skip anymore: not the model, the mess. The vendors talking here are right to focus on governance and movement, because “just throw more GPUs at it” is a cute plan until the data is still hiding in six places and half of it has no labels. AI is becoming a storage discipline with better marketing.
Read more about this at: SiliconANGLE