TLDRocket
Sign in

Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models

MarkTechPost Michal Sutter

Prime Intellect launched Prime Inference for open models. It mixes serverless and reserved GPUs, and says its GLM-5.3 endpoint is already fast and stable.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Prime Intellect has put Prime Inference into the world as a serving layer for frontier open-source models. The pitch is simple enough: use serverless endpoints when demand jumps around, or reserve capacity on Prime’s own GPUs when the workload is steady. Under the hood, the platform already has some real mileage. Before launch, Prime says it was handling nearly a trillion tokens a day internally.

That traffic wasn’t just idle chat. It came from RL rollouts, synthetic data generation, evaluations, and long-running coding agents. Prime is trying to make serving part of the same loop as training, not a separate afterthought. The company already ships post-training tools such as prime-rl, verifiers, and sandboxes, so this release fills in the missing piece: deploy a model, collect production traces, feed them back into training, repeat.

The stack leans on NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, and Prime says it built the system with Inferact and NVIDIA. Traffic can fail over automatically across datacenters, and the service is OpenAI compatible, so existing SDKs can point at the Prime Inference endpoint without a rewrite. Prime also says the GLM-5.3 endpoint ranks among the fastest on OpenRouter, with near-zero tool-call errors and 100% uptime since launch.

The engineering details are aimed squarely at agentic workloads, where prompts are huge and each turn adds more tokens on top. Prime says it separates prefill and decode onto different GPU groups, uses a KV-aware router to keep sessions on the same decoder, and adds a second KV tier in host DRAM through Mooncake. In tests on GLM-5.3 running on GB200 NVL72, it says a 1:4 prefill/decode ratio served the most users, while cutting the prefill budget from 8K tokens to 4K dropped median queue wait from 550 ms to 110 ms.

There are more wins in the numbers. Prime says NVFP4 KV compression increased cached tokens per decoder from 1.09 million to 1.63 million, and its sparse-MLA kernel was quicker in one workload-specific test. It also says it shrank transfer descriptors from 19,559 to about 1,940, cutting mean transfer time from 146 ms to 78 ms. The pricing docs are still incomplete, though, so anyone hoping for a neat bill right now gets the usual startup special: architecture first, invoice later.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of OpenAI-compatible move: less demo theater, more actual serving plumbing. The open-model crowd has spent too long pretending training is the whole game, when the boring bits like routing, cache handling, and failover decide whether anything is usable. Prime Intellect seems to get that, which is rarer than it should be.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.