TLDRocket
Sign in

Kubernetes can run AI inference. But can it count the real cost?

The New Stack Bill Doerrfeld

Kubernetes is getting better at AI inference, but the bill is getting harder to read. That’s why a bank’s big utilization wins and a WEKA exec’s warning land in the same week.

Based on reporting by The New Stack, Bill Doerrfeld — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Kubernetes keeps widening its job description. This week’s Road to KubeCon roundup puts AI inference front and center, alongside fresh security work, a new Gartner quadrant, and a few ecosystem updates that matter to the people actually running this stuff.

The biggest practical example comes from China Merchants Bank. At a CNCF event in China, the bank’s infrastructure team won the CNCF End User Case Study Contest with a setup that combines Kubernetes with Kueue, KEDA, Prometheus, HAMi, and Fluid. The bank says it has nearly 10,000 accelerator cards in its AI pool, and that the architecture unified management of 99% of its AI compute resources. It also pushed average utilization from 35% to more than 60% and cut the cost of processing 1 million tokens by 60% under comparable conditions.

That kind of result is exactly why the CNCF is adding an AI Inference + Agentic track to KubeCon + CloudNativeCon North America 2026. The event in Salt Lake City, Utah, runs November 9-12, and the new track is meant to cover the shift from training models to serving them in production. It also points to the tooling now orbiting inference: MCP, A2A, and AI gateways.

But the economics are still messy. Val Bercovici, chief AI officer at WEKA, argues that Kubernetes was never built to understand the things that drive inference cost: request mix, KV cache occupancy, prefill versus decode, and how memory and bandwidth behave inside the accelerator after a pod is already running. His point is blunt. Kubernetes may stay in the stack, but without a better resource model, it risks becoming a tax on inference economics.

Elsewhere, the platform is getting harder in the literal sense. Red Hat described two Alpha storage security features in Kubernetes v1.37 — new bind mount options and emptyDir permissions — aimed at helping users apply native hardening controls such as noexec, nodev, and nosuid. OpenTelemetry also pushed its Kubernetes attributes processor to v1.0.0, giving collectors a cleaner way to attach Kubernetes metadata. And DigitalOcean’s Spot GPU node pools entered public preview on September 9, offering interruptible GPU capacity at a lower, variable rate than on-demand nodes.

My take — AI-written commentary, not fact-checked reporting

The industry loves to talk about Kubernetes as the universal control plane, then acts surprised when the control plane can’t price a token. That’s the real story here: inference is dragging accounting, scheduling, and cache awareness into the room whether platform teams like it or not. The next fight isn’t about who can run AI on Kubernetes; it’s about who can explain the invoice without laughing.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.