Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling
MarkTechPost Michal Sutter
Underdog shrank Qwen3.8-27B into a 7.89 GB file that runs in stock llama.cpp. It still beats the full model at tool calling, which is the awkward part most small models lose.
Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Conway Research’s Underdog has put out Saluki 27B, an Apache 2.0 release built from Qwen3.8-27B. The pitch is blunt: take a 27B model, compress it hard enough to fit in 7.89 GB, and keep the part that matters for agents. That part is tool calling, the bit that turns a chat model into something that can actually do work.
The full BF16 version of Qwen3.8-27B needs 54 GB, so this is not a gentle trim. Saluki is a 2-bit, mixed-precision GGUF, tuned further by Underdog after an earlier compression pass from ISTA-DASLab. The company says it did not publish the full recipe for its own pass, only that the result is an IQ2-mix file with an imatrix tag.
The interesting twist is where the model holds up. On Underdog’s internal tool-use test, Saluki scored 88 against 84 for the full model, and on parallel tool calls it hit 42 versus 35. Across nine benchmarks, Underdog reports 96% average retention. But the cracks show up where you’d expect: AIME 2025 drops to 79.2 from 96.7, and AIME 2026 lands at 80.0 versus 94.6.
Underdog says Saluki runs on stock llama.cpp and apps built on it, with full GPU offload. That matters because some competing compact builds still need custom forks. Saluki also comes with optional 629 MB or 928 MB vision add-ons, while the base model keeps Qwen3.8-27B’s 262,144-token native context. The message here is pretty clear: if you want a smaller local agent model that still behaves well on tool use, this one is aimed straight at that job.
My take — AI-written commentary, not fact-checked reporting
This is the kind of compression story that actually matters: not “small for small’s sake,” but smaller while keeping the agent bit alive. Hype loves giant model scores; real users love models that fit on disk and still call tools without drama. That’s the useful benchmark, not another victory lap over math olympiad trivia.
Read more about this at: MarkTechPost