TLDRocket
Sign in

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

Simon Willison’s Weblog Simon Willison ● Covered by 4 sources

Opinion — commentary, not a factual news event.

Bonsai 2 27B fits in about 5.95 GB, but only with Prism’s llama.cpp fork. That’s a huge shrink for a 27B model, if the runtime weirdness is sorted.

Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Simon Willison says the Bonsai 2 27B GGUF files can be tried from Prism’s Hugging Face repo, but there’s a catch: they need Prism’s own llama.cpp fork, not the stock build most people would reach for first. He even lays out the exact path, from grabbing the macOS runtime to downloading the model file and launching the server on port 8331.

The model file he points to is about 5.95 GB. For a 27B model, that’s the real headline here. Bonsai 2 27B is being framed as near-lossless compression in a footprint that is roughly nine times smaller, which is the sort of thing that makes local inference feel less like a hobbyist stunt and more like a practical option.

Willison says the bundled llama-server web UI is “very good,” and he also shows how to hit the local server through an OpenAI-style endpoint using the llm CLI. So this is not just a compressed model sitting on a shelf. It’s meant to be used.

Performance, though, looks a bit messy. On an M5 Pro, he saw around 20 tokens per second, then 44 after a server restart, and he says he’s not sure why the jump happened. On startup, the server reported that the tensor API wasn’t supported in that environment and disabled it. That’s a reminder that clever compression is one thing; a clean runtime is another.

My take — AI-written commentary, not fact-checked reporting

This is exactly where open-model progress gets interesting: not in giant benchmark posters, but in making a 27B model small enough to actually run. The annoying part is the usual one — the model is impressive, the tooling is the tax collector. Prism deserves credit for pushing the trick forward, and a little side-eye for making the path this specific.

Read more about this at: Simon Willison’s Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.