TLDRocket
Sign in

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

GitHub anerli

Magnitude launched: a self-optimizing inference engine for agents. It says local models run faster on Macs, Linux and Windows, up to 2x over llama.cpp.

Based on reporting by GitHub, anerli — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Two software engineers are betting that local agents need a different kind of inference engine. Anders and Tom have launched Magnitude, an open-source engine built to optimize itself for the machine it’s running on, instead of assuming a datacenter setup or a single hardware target.

Their pitch is blunt: most inference stacks make a tradeoff somewhere. Some are tuned for batched datacenter workloads, others try to work across lots of hardware, and some go hard on one chip or model family without being a full engine. Magnitude is meant to sit in the awkward middle that agent workflows actually live in: long-running sessions, multiple sessions at once, and a user who still wants the computer to feel usable.

The way it does that is by tuning kernels on the device before the model runs, reserving only enough memory up front for weights, and then growing the memory heap as sessions expand. When agents go idle, it frees memory back up. It also borrows the shared-prefix idea from engines like SGLang, but tries to keep placement friendly to single-session speed rather than sacrificing that for concurrency.

Magnitude ships as a desktop app and can hook into agents people already use, including Pi, OpenCode, Hermes and Codex. It automatically starts models when they’re needed and shuts them down after inactivity. The company says it’s fully open source under Apache 2.0 and written in Rust, with a custom GPU kernel runtime and autotuner.

On the benchmark side, the team compared it with llama.cpp using Qwen 3.6 35B A3B at 4-bit, with a 64k context and no speculative decoding. On a Mac M4 Pro with 48 GB of unified memory, they report 92% faster decode, 9% faster prefill and 28% lower per-agent memory use. On CUDA with DGX Spark, they report 19% faster decode, 23% faster prefill and 27% lower memory use.

My take — AI-written commentary, not fact-checked reporting

This is the right instinct: agents don’t need another engine that only behaves when the workload is polite. They need software that knows the laptop is also for email, Slack and whatever else people call “light multitasking” now. Open source helps here, because performance claims without code are just expensive fan fiction.

Read more about this at: GitHub

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.