Qwen3.8 27B addition in words
Simon Willison’s Weblog Simon Willison
Qwen3.8-27B was tested on a tricky sum-in-words task and got 167 of 169 right with reasoning on. With reasoning off, the results were much messier.
Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Simon Willison replayed a neat old experiment: give a model two large numbers, ask for their sum, and force it to return the answer in words. The original version, posted by Colin Frasier on Bluesky more than two years ago, used GPT-4o and tracked how well it handled bigger and bigger additions. Willison pulled the idea into a local setup and ran it on a DGX Spark.
He used a Codex Remote session running GPT-6 Astra to drive the test, with Qwen3.8-27B-Q4_K_M.gguf doing the actual work. First came a 30-attempt run for each combination with reasoning turned off. Then he repeated the experiment with reasoning enabled, but only one attempt per square because each run took much longer.
That second pass was the standout. Qwen3.8-27B got 167 of 169 attempts right. Because each cell was only tested once, the resulting heatmap looked rougher than the earlier one, with each square landing at either 100% or 0% instead of showing a smoother spread.
Willison also shared a report that included some of the model’s reasoning traces from the larger sums. The snippets read like a model talking itself through grade-school arithmetic: aligning digits, adding from right to left, and checking carries one position at a time. For a task this simple on paper, that’s exactly the point — the difference between a model that guesses and one that can keep its place.
My take — AI-written commentary, not fact-checked reporting
This is the kind of test that cuts through a lot of model marketing fluff. Local models don’t need to be magical; they need to be boringly correct, and Qwen3.8-27B mostly was. The bigger lesson is that “reasoning” isn’t just theater when the results jump that sharply.
Read more about this at: Simon Willison’s Weblog
Related stories
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
Simon Willison’s Weblog · 1 month ago ·
7
Qwen 3.6 27B is the sweet spot for local development
Quesma · 3 months ago ·
33
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Simon Willison’s Weblog · 1 month ago ·
44