TLDRocket
Sign in

Storage leaders target AI’s unstructured data challenge

SiliconANGLE Mark Albertson

AI’s biggest data problem is still the messy stuff: docs, videos, logs, emails. Storage vendors are racing to move and govern it before AI can use it.

Based on reporting by SiliconANGLE, Mark Albertson — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

The storage crowd has found AI’s dirtiest secret: the best fuel is usually the least organized. Industry estimates cited by theCUBE Research say more than 80% of enterprise data is unstructured, and that giant pile of documents, videos, logs, emails, images and audio is still mostly hard to reach, hard to govern and hard to use.

That is now a storage problem as much as an AI problem. On SiliconANGLE’s theCUBE, Molly Presley of Hammerspace said managing unstructured data for AI is far tougher than the old archive-and-backup routine. The new job is to unify the data and move it automatically, fast enough for AI workflows that actually need it.

Supermicro, Cloudian, Hammerspace and Seagate all described different parts of that pipeline. Sherry Lin said Supermicro is building high-density hardware for software-defined storage partners, and has added Context Memory storage servers to help offload and share large language model key-value caches across AI inference systems. Lin also outlined a split where Hammerspace handles the global namespace and orchestration, Cloudian provides the S3 object store for AI data, and Seagate’s hard drives sit at the end of the line.

Cloudian’s pitch is simple: keep unstructured data under control, protected and secure, even as it moves for different uses. Peter Sjoberg said that control matters because the same data may need to be pulled for AI work, compliance audits, debugging or to understand model outputs. Once the data goes into the AI era, it stops being just storage and becomes something that has to stay governed.

Seagate’s Mohamad El-Batal pushed the same idea from the infrastructure side. He said companies need to understand the value of the data and where they are trying to go with it before choosing how to build. His advice was blunt: do not go cheap on infrastructure, but do keep it efficient. Presley added that Hammerspace is working on a Model Context Protocol layer to help AI systems understand data that is not well described, because that is the nature of unstructured data.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI nobody gets to skip anymore: not the model, the mess. The vendors talking here are right to focus on governance and movement, because “just throw more GPUs at it” is a cute plan until the data is still hiding in six places and half of it has no labels. AI is becoming a storage discipline with better marketing.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.