TLDRocket
Sign in

Qwen3.8 27B addition in words

Simon Willison’s Weblog Simon Willison

Qwen3.8-27B was tested on a tricky sum-in-words task and got 167 of 169 right with reasoning on. With reasoning off, the results were much messier.

Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Simon Willison replayed a neat old experiment: give a model two large numbers, ask for their sum, and force it to return the answer in words. The original version, posted by Colin Frasier on Bluesky more than two years ago, used GPT-4o and tracked how well it handled bigger and bigger additions. Willison pulled the idea into a local setup and ran it on a DGX Spark.

He used a Codex Remote session running GPT-6 Astra to drive the test, with Qwen3.8-27B-Q4_K_M.gguf doing the actual work. First came a 30-attempt run for each combination with reasoning turned off. Then he repeated the experiment with reasoning enabled, but only one attempt per square because each run took much longer.

That second pass was the standout. Qwen3.8-27B got 167 of 169 attempts right. Because each cell was only tested once, the resulting heatmap looked rougher than the earlier one, with each square landing at either 100% or 0% instead of showing a smoother spread.

Willison also shared a report that included some of the model’s reasoning traces from the larger sums. The snippets read like a model talking itself through grade-school arithmetic: aligning digits, adding from right to left, and checking carries one position at a time. For a task this simple on paper, that’s exactly the point — the difference between a model that guesses and one that can keep its place.

My take — AI-written commentary, not fact-checked reporting

This is the kind of test that cuts through a lot of model marketing fluff. Local models don’t need to be magical; they need to be boringly correct, and Qwen3.8-27B mostly was. The bigger lesson is that “reasoning” isn’t just theater when the results jump that sharply.

Read more about this at: Simon Willison’s Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.