Robotics and edge AI put new pressure on computing infrastructure
SiliconANGLE Chad Wilson ● Covered by 4 sources
Robots and edge devices are forcing companies to rethink how AI computing gets built and paid for. Turns out chatbots were the easy part — physical AI eats way more tokens and needs hardware everywhere, not just the cloud.
Everyone got comfortable with the idea that AI infrastructure meant a few hyperscale data centers stuffed with GPUs. Physical AI is breaking that assumption. Once you put intelligence into a robot, a car, or a factory sensor, the math changes completely — and a group of infrastructure vendors speaking at theCUBE's Robotics & AI Infra Leaders event laid out just how much.
Start with tokens. Max Kan of SemiAnalysis pointed out that agentic systems burn through tokens exponentially faster than last year's chatbots, because every follow-up question or tool call forces the model to reprocess the entire conversation history. That's not a minor efficiency tax — it reshapes what infrastructure needs to prioritize. Some workloads need raw compute for prefill, others need memory bandwidth for decode, and companies like Positron AI are now building air-cooled inference boxes specifically tuned for tokens-per-dollar and tokens-per-watt, aimed at data centers that were never designed for liquid cooling in the first place.
The security and cost problem is just as pressing. Rafay Systems CEO Haseeb Budhani described how his company slices GPUs into confidential, isolated chunks so cloud providers can sell capacity to multiple enterprises without handing over an entire system to each one. That's good for providers, who squeeze more revenue out of the same hardware, and good for enterprises, who get a lower total cost of ownership without sacrificing data security. It's the kind of unglamorous plumbing work that determines whether AI actually pencils out financially for a mid-size manufacturer, not just a hyperscaler with unlimited capex.
Meanwhile the intelligence itself is shrinking to fit the edge. Liquid AI's Ramin Hasani argued that small, fine-tuned models won't be generally intelligent, but they don't need to be — customizing them is cheap, and running them locally on a laptop or a vehicle cuts latency and keeps sensitive data off the network. Axelera AI is pushing similar edge-first architecture into servers and cloud systems using familiar developer tools, betting that frictionless software adoption matters as much as chip performance. Layer on top of that new financial infrastructure — Carmen Li's Silicon Data and Compute Exchange are literally building hedging instruments for GPU capacity, treating compute like a commodity with futures and volatility, because neoclouds and AI companies now have exposure that looks a lot like a trading desk's balance sheet.
What ties all of this together is a quiet admission: the cloud-centric AI model doesn't scale to robots and physical systems without serious rearchitecting. Power management chips from Axiado, telemetry-heavy networking from Aria Networks, human-gated automation from H Company — these aren't flashy model announcements, but they're the scaffolding that decides whether physical AI ships products or stays stuck in pilot programs.
My take
The industry spent two years hyping frontier models while the actual bottleneck was always going to be boring infrastructure — cooling, GPU slicing, token economics, compute futures markets. Nobody puts that on a keynote slide, but it's the stuff that decides whether a robotics startup survives its first production deployment or drowns in cloud bills. Watch the vendors solving power and cost problems, not the ones promising another million-token context window; that's where the real money and the real constraints actually live.》
Read more about this at: SiliconANGLE
Related stories
Running AI on mixed hardware for speed and affordability
IBM Research · 1 month ago ·
23
At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI
NVIDIA · 2 weeks ago ·
23