TLDRocket
Sign in

BaseRT provides faster runtime for serving AI models with less infrastructure

basecompute.co

A new runtime called BaseRT claims to be the fastest way to run AI models on Apple Silicon. It beats llama.cpp and MLX at their own game, up to 6.4x faster on prefill.

Based on reporting by basecompute.co — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

BaseRT showed up this week with a simple pitch: your Mac can run local AI models faster than the tools you're already using, and you don't need a GPU cluster to prove it. The startup behind it, Basecompute, is positioning the runtime as a direct challenger to Apple's own MLX framework and the venerable llama.cpp, both of which have become default choices for anyone running open-weight models on M-series chips.

The numbers Basecompute is publishing are not subtle. On prefill — the phase where a model chews through your prompt before it starts generating anything — BaseRT claims up to 6.4x the throughput of llama.cpp and 3.9x that of MLX, tested on models like Qwen3 30B-A3B and Gemma 4 E2B running on an M5 Pro. Decode speed, the token-by-token generation most people actually notice, sees a smaller but still real bump of up to 1.33x. Prefill gains matter more than they sound: for long documents, big codebases, or anything with a hefty system prompt, that's where wall-clock time actually goes.

Installation is a one-line curl script, which tells you who this is built for. No enterprise sales call, no dashboard tour. Basecompute is also shipping a plugin so BaseRT can slot straight into coding agents — serve a model locally, point your agent's config at it, and the whole loop stays on your machine. No API key, no request logs sitting on someone else's server. For developers who've gotten used to sending every prompt to OpenAI or Anthropic, that's a meaningfully different workflow, not just a speed bump.

What's notable is the timing. Apple Silicon has quietly become the preferred hardware for local inference, thanks to unified memory letting big models fit without a discrete GPU. MLX was Apple's answer to owning that stack. BaseRT betting it can out-optimize Apple's own framework is a bold claim, and one that will get tested hard by anyone with a Mac and five minutes to spare — which is exactly the crowd Basecompute is courting with a Discord built for people already building on-device.

My take — AI-written commentary, not fact-checked reporting

I'll believe the 6.4x number once independent benchmarks confirm it, because runtime vendors love picking the workload that flatters them most. But the bigger story here isn't the multiplier — it's that local, no-API-key inference keeps getting faster and easier, which is exactly the trend that eventually makes cloud-only model APIs look like a tax rather than a necessity.

Read more about this at: basecompute.co

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.