TLDRocket
Sign in

Rafay Systems targets the operating layer of the AI infrastructure boom

SiliconANGLE John Furrier

Rafay Systems wants to be the software layer that turns raw GPU clusters into actual cloud services. Because buying chips is easy now — making money off them fast is the real bottleneck.

There's a specific kind of anxiety spreading through the AI infrastructure world right now, and it has nothing to do with chip shortages. It's about what happens after the GPUs land. Companies have spent the last two years racing to secure Nvidia hardware, but a rack full of silicon isn't a business. It's a very expensive liability until someone wires up orchestration, multitenancy, security, and a way for customers to actually consume it without picking up the phone.

Haseeb Budhani, who co-founded and runs Rafay Systems, has built his company around that gap. In a recent conversation with SiliconANGLE, he laid out a blunt test for what counts as a cloud versus what's just custom infrastructure with a fancy name: can a customer hit an API or click a button in a portal and get an AI service running in a shared, multitenant environment, with zero phone calls? If yes, you're a cloud. If not, you're a very costly server closet.

That distinction matters because AWS, Microsoft, and Google spent well over a decade and thousands of engineers building the control planes that make this look easy. The neoclouds, sovereign providers, and telecoms now racing into AI infrastructure don't have a decade. Budhani says some of them have months, not years, and depreciation on GPU hardware doesn't wait for anyone to catch up. Rafay's pitch is to sit between the bare metal and the customer, packaging Kubernetes, VMs, serverless, open-source models, and token-based consumption into one operable layer, because different customers — a model lab, an enterprise dev team, a reseller — all want something different, and margins improve the further up that stack a provider can climb.

Sovereignty is accelerating all of this. Governments and regions want AI compute and data kept local, and Budhani says that demand started with model builders and agentic app developers before spreading to ordinary enterprises, who show up expecting the quotas, audit trails, and attestation they're used to from hyperscalers. He's also candid that software alone doesn't close deals — what he calls Rafay's real investment is the

My take

The GPU land-grab was always going to be the easy part; wiring up multitenancy, billing, and audit trails at hyperscaler-grade reliability with a skeleton crew is the actual hard problem, and most neoclouds are underestimating it. Sovereign AI is real demand, not just political theater, but plenty of regional providers will discover that promising sovereignty without operational maturity just means slower, more expensive AWS knockoffs. Watch the vendors who sell the plumbing, not the chips — that's where the consolidation pressure will land first.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.