PrismML releases Bonsai 2 27B, a compact multimodal LLM designed to run on consumer devices with weights released under Apache 2.0
Model release ● Confirmed 90% confidence first seen
PrismML launched Bonsai 2 27B, an ultra-compact multimodal language model intended to run locally on PCs and some high-end mobile devices. The release compresses a Qwen3.8 27B variant to about 5.9 GB while aiming to preserve roughly 98% of aggregate benchmark performance, and the weights are provided under Apache 2.0.
Decision brief
- What changed
- PrismML released Bonsai 2 27B, a compact multimodal version of Qwen3.8 27B sized at about 5.9 GB and intended to run locally on PCs and some high-end mobile devices. The model weights were released under Apache 2.0, with reported retention of about 98% to 98.2% of the parent model’s aggregate benchmark performance.
- Why it matters
- This gives companies a new open-weight option for on-device AI deployments where privacy, offline use, latency, or inference cost matter, because the reported model size and hardware requirements are much lower than the original FP16 form. For decision-makers, the practical implication is not just cheaper local inference but potential control over distribution and product embedding because Apache 2.0 licensing is permissive. However, adoption decisions should account for reported runtime constraints, since one outlet said the model currently requires PrismML’s llama.cpp fork or its MLX runtime rather than stock llama.cpp.
- Evidence
- All three cited outlets reported the same core facts: PrismML launched Bonsai 2 27B, reduced the model to roughly 5.9 GB, targeted consumer hardware, and claimed about 98% performance retention versus Qwen3.8 27B. SiliconANGLE and MarkTechPost both specified Apache 2.0 licensing, and MarkTechPost added an implementation detail about needing PrismML-specific runtime support.
- What remains uncertain
- The coverage relies on PrismML’s reported benchmark retention and deployment claims; it does not independently verify real-world latency, power draw, quality on specific enterprise tasks, or compatibility across actual consumer devices. It is also unclear from the coverage how broad multimodal support is in production settings, how mature the tooling is outside PrismML’s runtimes, and whether licensing or redistribution has any practical constraints beyond Apache 2.0.
- Monitor next
- Watch for independent benchmark and deployment tests showing real-world performance, device compatibility, and whether support expands beyond PrismML’s current llama.cpp fork or MLX runtime.
Analytical support, not advice — assumptions and open questions stated above.