TLDRocket
Sign in

The AI inference race moves beyond GPUs to reshape data center infrastructure

SiliconANGLE Victoria Gayton Covered by 3 sources

AI inference is turning data centers into whole-system puzzles, not just GPU contests. Storage, networking and power now shape cost, speed and how fast tokens show up.

Based on reporting by SiliconANGLE, Victoria Gayton — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AI inference is no longer just a fight over faster GPUs. As generative and agentic apps move into production, the bottlenecks spread out: storage latency, network bandwidth, data movement and power all start deciding how quickly and cheaply tokens get produced.

That shift changes the whole design problem. Interactive chat wants low latency. Batch inference wants throughput. Agentic systems keep growing their context windows. Add retrieval-augmented generation and multi-tenant AI factories, and the stack gets even more demanding, because the data has to be fresh, reachable and available without dragging everything into one giant pile first.

Ka Wai Leung of IBM put it simply: understand the workload first, then build around it. That matters in enterprise settings, where structured, unstructured and multimodal data can be scattered across mainframes and other systems. The challenge is not only access. It is also data gravity and data sovereignty, which make it awkward to haul huge datasets around just to feed an AI factory.

Power is part of the same story. Kioxia’s Anders Graham said the company saw a 76% improvement in random read and more than 100% improvement in random write IOPS per unit of power when comparing its newer BiCS8-based CM9 drives with the older CM7 generation. In a world where data centers are running into energy and space limits, that kind of gain matters as much as raw speed.

The reference design discussed by Supermicro, IBM and Kioxia ties those pieces together: Nvidia HGX B300 compute, IBM Storage Scale Erasure Code Edition, and Kioxia drives. The setup uses a high-performance storage tier for latency-sensitive work and can add a capacity tier for colder data. The companies also tested IBM Storage Scale as a shared KV cache, which let previously computed context be reused without filling GPU memory or system RAM.

And the test held up under stress, at least to a point. With heavy traffic added to the Spectrum-X network, the uncached advantage dropped from 4.8 requests per second to about 3.6, while still showing roughly 18 times better efficiency instead of 22 times. Real networks are noisy. Apparently, AI infrastructure is now expected to be noisy too.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI that deserves more airtime: the expensive plumbing underneath the demo. Everyone loves talking about models; the boring truth is that storage and networking decide whether those models are a lab trick or a production bill. The GPU gets the headline, and the data center gets the invoice.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.