Exclusive: Iterate.ai’s Lifeboat runs up to six times more AI agent sessions per GPU
SiliconANGLE Duncan Riley
Iterate.ai launched Lifeboat, a new AI engine that packs more agent sessions onto each GPU. It says one card can handle up to six times as many, and with built-in confidential computing.
Based on reporting by SiliconANGLE, Duncan Riley — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Iterate.ai has a new pitch for companies staring at GPU bills: squeeze harder before you buy more hardware. The enterprise AI software company launched Lifeboat, an inference engine for large language models that also bakes in confidential computing, and says it can fit two to six times as many concurrent AI agent sessions on each graphics processing unit.
That matters because agents are messy in a way chatbots usually aren’t. A single chat request often means one model call. An agent working through a task can rack up dozens, and each round adds to its context window. At scale, that pushes the key-value cache hard, and Iterate.ai says standard inference engines can bog down once they hit four or five long-context requests at the same time.
Lifeboat’s answer is a mix of fair scheduling, admission control and cache tuning. Each session gets its share of the GPU, so a heavy document-processing agent does not shove aside more interactive work. The company says its cache optimizations double effective capacity while model weights stay at full precision. On mixture-of-experts models, it loads only the experts a given agent needs. Each session also runs in its own security capsule with filtering, token budgets and sandboxed execution.
The company’s own testing used a single Nvidia RTX PRO 6000 Blackwell GPU running a Qwen 30B-A3B model. Lifeboat handled 2,048 concurrent sessions there, with every request completing. With the optimizations switched off, the same engine topped out at half that number and processed 4,965 tokens per second instead of 8,714. In a memory-pressure test with 128 sessions sending 18,000-token requests, Lifeboat kept 99th-percentile time to first token at 1.5 seconds; the baseline needed 189.
Iterate.ai is clearly aiming at banks, insurers and health systems that want agents running on their own data inside their own walls. Jon Nordmark, the company’s chief executive, said buyers should figure out what their existing GPUs can do before ordering more. The argument is simple, and a little brutal: if one card can do the work of several, the data center doesn’t need new racks just yet.
My take — AI-written commentary, not fact-checked reporting
This is the kind of AI story that actually matters: not bigger demos, but better utilization. The industry loves selling more GPUs; Iterate.ai is trying to make that impulse look lazy. That’s healthier, and a lot less glamorous, which is usually how the useful stuff arrives.
Read more about this at: SiliconANGLE
Related stories
Why AI infrastructure is more than “more GPUs” (long-running agents)
YouTube · 1 week ago ·
17
How Together AI Uses AI Agents to Automate Complex Engineering Tasks: Lessons from Developing Efficient LLM Inference Systems
Together AI · 1 year ago ·
9