TLDRocket
Sign in

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost

MarkTechPost Michal Sutter

BottleCap AI released a Qwen3.8-27B fine-tune that thinks less. It cuts reasoning tokens 37.2% on average, with accuracy down just 0.86 points.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

BottleCap AI has released ThinkingCap-Qwen3.8-27B, the second model in its ThinkingCap line. It is a fine-tune of Qwen3.8-27B with a very specific aim: make the model reason with fewer tokens, not with a different personality or a bigger knowledge base.

Across 12 benchmarks, that gamble mostly works. The model uses 37.2% fewer thinking tokens on average, while macro-average accuracy slips from 86.65% to 85.79%. In other words, BottleCap is trading a little raw score for a lot less internal chatter. On the company’s numbers, the pooled mean thinking token count falls from 15,735 to 12,144.

Some of the cuts are dramatic. MMMLU drops 65.5% in thinking tokens, MMLU-Pro falls 57.3%, and GPQA-Diamond is down 43.1%. IFBench gets 46.4% fewer tokens with accuracy almost unchanged, while long-context retrieval actually improves: AA-LCR rises 2.25 points to 84.00% and still uses 38.6% fewer thinking tokens. LiveCodeBench v6 also edges up slightly, even as thinking shrinks by 20.3%.

The trade-offs show up where you’d expect them to. AIME 2026 loses 3.85 points, from 98.13% to 94.27%, for a 30.2% reduction in thinking. Agentic tasks stay fairly close to the base model: τ²-bench gives up 1.01 points for a 30.9% cut, and Terminal-Bench 2.1 loses 0.56 points, inside its reported interval, with a 10.7% cut.

BottleCap says the model is meant to drop into Qwen3.8-27B workflows on vLLM or SGLang, with FP8, NVFP4, GGUF and MLX builds available. The repository is gated, and commercial use beyond the small-business license needs a BottleCap agreement. The company also says the compression stacks with Qwen3.8-27B’s reasoning-effort setting, and it recommends xhigh for the best accuracy-to-token balance.

My take — AI-written commentary, not fact-checked reporting

This is the sort of AI work that actually deserves attention: less drama, fewer tokens, and a clear trade-off. The industry loves models that talk a lot and call it “reasoning”; BottleCap is basically telling them to shut up and answer. That’s a healthier benchmark than another glossy demo with no shipping story.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.