BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost
MarkTechPost Michal Sutter
BottleCap AI released a Qwen3.8-27B fine-tune that thinks less. It cuts reasoning tokens 37.2% on average, with accuracy down just 0.86 points.
Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
BottleCap AI has released ThinkingCap-Qwen3.8-27B, the second model in its ThinkingCap line. It is a fine-tune of Qwen3.8-27B with a very specific aim: make the model reason with fewer tokens, not with a different personality or a bigger knowledge base.
Across 12 benchmarks, that gamble mostly works. The model uses 37.2% fewer thinking tokens on average, while macro-average accuracy slips from 86.65% to 85.79%. In other words, BottleCap is trading a little raw score for a lot less internal chatter. On the company’s numbers, the pooled mean thinking token count falls from 15,735 to 12,144.
Some of the cuts are dramatic. MMMLU drops 65.5% in thinking tokens, MMLU-Pro falls 57.3%, and GPQA-Diamond is down 43.1%. IFBench gets 46.4% fewer tokens with accuracy almost unchanged, while long-context retrieval actually improves: AA-LCR rises 2.25 points to 84.00% and still uses 38.6% fewer thinking tokens. LiveCodeBench v6 also edges up slightly, even as thinking shrinks by 20.3%.
The trade-offs show up where you’d expect them to. AIME 2026 loses 3.85 points, from 98.13% to 94.27%, for a 30.2% reduction in thinking. Agentic tasks stay fairly close to the base model: τ²-bench gives up 1.01 points for a 30.9% cut, and Terminal-Bench 2.1 loses 0.56 points, inside its reported interval, with a 10.7% cut.
BottleCap says the model is meant to drop into Qwen3.8-27B workflows on vLLM or SGLang, with FP8, NVFP4, GGUF and MLX builds available. The repository is gated, and commercial use beyond the small-business license needs a BottleCap agreement. The company also says the compression stacks with Qwen3.8-27B’s reasoning-effort setting, and it recommends xhigh for the best accuracy-to-token balance.
My take — AI-written commentary, not fact-checked reporting
This is the sort of AI work that actually deserves attention: less drama, fewer tokens, and a clear trade-off. The industry loves models that talk a lot and call it “reasoning”; BottleCap is basically telling them to shut up and answer. That’s a healthier benchmark than another glossy demo with no shipping story.
Read more about this at: MarkTechPost
Related stories
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
Simon Willison’s Weblog · 1 month ago ·
7