TLDRocket
Sign in

Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots

The New Stack Adrian Bridgwater

Red Hat AI 3.5 adds tighter control for shared GPUs and AI tenants. It’s aimed at turning messy pilots into governed enterprise systems.

Based on reporting by The New Stack, Adrian Bridgwater — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Red Hat has pushed out Red Hat AI 3.5, and the pitch is simple: enterprise AI should be run with the same discipline as the rest of the critical stack. The new release leans hard into multi-tenancy, hardware-to-software isolation, and priority-aware service requests on shared GPU infrastructure. That matters because the easy part of AI has been building pilots. The hard part is keeping them predictable when they hit real workloads, sensitive data, and regulated environments.

Tushar Katarki, Red Hat’s senior director of product for Red Hat AI, frames the problem in blunt terms: without safety controls, enterprise AI is “like driving a supercar blindfolded.” His argument is that Red Hat AI 3.5 gives platform teams the guardrails to move from isolated experiments to something closer to a governed architecture. The release ties together pre-deployment safety benchmarking, real-time observability, and GPU resource management, which is exactly the sort of plumbing that gets ignored right up until a queue starts backing up.

The GPU queue is the real story here. Red Hat says its new shared GPU controls let it run native multi-tenancy on shared hardware, with fair-share GPU scheduling and priority-aware serving deciding which workloads get served first. Higher-priority inference can move ahead, while lower-priority jobs can use spare capacity instead of forcing separate GPU resources to spin up. On top of that, the isolation layer is meant to keep one tenant from touching another tenant’s data, models, or compute environment.

There’s also a practical set of tools layered on top. EvalHub is there for model verification before deployment, with safety benchmarking and compliance certifications in mind. New observability dashboards show inference health, GPU utilization, and model performance, while non-admin users get token-consumption showback and distributed inference visibility. Red Hat AI Hub adds agent templates and starter kits for code review, document processing, and research workflows.

The release also reaches into memory management. CPU offloading is generally available, and storage offloading is in developer preview, both aimed at handling longer conversations and larger documents without buying more GPU hardware. Red Hat’s bet is that AI infrastructure is moving from a pile of experiments to a policy-controlled pool. That sounds less glamorous than the usual AI sales pitch, which is exactly why it may be the more useful one.

My take — AI-written commentary, not fact-checked reporting

This is the unsexy truth of enterprise AI: the winner won’t be the company with the loudest demo, but the one that can stop one team from wrecking another team’s day on shared hardware. Open systems and strong controls beat theatrical chaos, every time. GPUs are expensive enough without turning them into a company-wide traffic jam.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.