Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Simon Willison’s Weblog Simon Willison ● Covered by 4 sources
Opinion — commentary, not a factual news event.
Bonsai 2 27B fits in about 5.95 GB, but only with Prism’s llama.cpp fork. That’s a huge shrink for a 27B model, if the runtime weirdness is sorted.
Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Simon Willison says the Bonsai 2 27B GGUF files can be tried from Prism’s Hugging Face repo, but there’s a catch: they need Prism’s own llama.cpp fork, not the stock build most people would reach for first. He even lays out the exact path, from grabbing the macOS runtime to downloading the model file and launching the server on port 8331.
The model file he points to is about 5.95 GB. For a 27B model, that’s the real headline here. Bonsai 2 27B is being framed as near-lossless compression in a footprint that is roughly nine times smaller, which is the sort of thing that makes local inference feel less like a hobbyist stunt and more like a practical option.
Willison says the bundled llama-server web UI is “very good,” and he also shows how to hit the local server through an OpenAI-style endpoint using the llm CLI. So this is not just a compressed model sitting on a shelf. It’s meant to be used.
Performance, though, looks a bit messy. On an M5 Pro, he saw around 20 tokens per second, then 44 after a server restart, and he says he’s not sure why the jump happened. On startup, the server reported that the tensor API wasn’t supported in that environment and disabled it. That’s a reminder that clever compression is one thing; a clean runtime is another.
My take — AI-written commentary, not fact-checked reporting
This is exactly where open-model progress gets interesting: not in giant benchmark posters, but in making a 27B model small enough to actually run. The annoying part is the usual one — the model is impressive, the tooling is the tax collector. Prism deserves credit for pushing the trick forward, and a little side-eye for making the path this specific.
Read more about this at: Simon Willison’s Weblog
Related stories
Deploying a 1-Bit Bonsai-27B Model with PrismML llama.cpp and OpenAI-Compatible Local Inference Workflows
MarkTechPost · 2 months ago ·
35
PrismML Releases Bonsai 27B: 1-bit and Ternary Builds of Qwen3.6-27B That Run on Laptops and Phones
MarkTechPost · 2 months ago ·
21
Someone Fine-Tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5 Traces to Ship a 657MB Local Thinking Model
MarkTechPost · 2 months ago ·
35