TLDRocket
Sign in

PrismML launches Bonsai 2 27B, a high-intelligence AI model so small it fits on consumer hardware

SiliconANGLE Kyt Dotson Covered by 3 sources

PrismML just shrank a big AI model to 5.9GB so it can run on some PCs and phones. The catch: it keeps about 98.2% of the original, which is the part that should make cloud vendors sweat.

Based on reporting by SiliconANGLE, Kyt Dotson — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Prism ML Inc. said Thursday it has launched Bonsai 2 27B, a second-generation multimodal generative model built to be tiny enough for consumer hardware. The headline trick is size: the company says it took a Qwen3.8 27B-based model that sits at about 56 gigabytes in full 16-bit form and compressed it to around 5.9 gigabytes.

That shrink job comes from ternary scaling, which reduces the model’s weights to +1, 0 and -1 instead of the usual 16-bit representation. PrismML says that approach preserves about 98.2% of the model’s capabilities. That matters because ordinary quantization can save space too, but often at the cost of accuracy, knowledge and other capabilities that make an AI model useful in the first place.

The company backed up the pitch with benchmark numbers. Bonsai 2 landed close to Qwen3.8 on agentic and tool-calling tasks, within 3 points, at 77.6 and 79.8. It also posted aggregate coding scores of 81.6 and 82.2 across HumanEval+, LiveCodeBench v6, MBPP+ and BigCodeBench, plus 82.7 and 81.3 for knowledge and reasoning across MMLU-Redux, GPQA Diamond and AA-LCR.

Bonsai 2 can run without quantization on an Nvidia GeForce GTX 5090, where PrismML says it reaches 143 tokens per second, and on Apple’s M5 Max chip at 46.8 tokens per second. The company says it uses extremely low power per token at 0.714 megawatt-hours, which it says makes it 40% more energy-efficient than other 8B models running in full precision.

And that is the real point here. Small models are no longer just for toy demos or narrow tasks. PrismML is betting that a local model can handle translation, summarization and search organization on-device, while heavier cloud systems stay reserved for the messy, high-touch jobs. The weights are available now under Apache 2.0, and the model runs on Nvidia GPUs through CUDA and on Apple devices including Mac, iPhone and iPad through MLX.

My take — AI-written commentary, not fact-checked reporting

The industry keeps acting like bigger is automatically better, which is a lovely way to sell expensive server time. PrismML is making the more annoying argument: if a model can stay intelligent after being squeezed this hard, a lot of cloud-first AI swagger starts looking like a convenience tax.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.