Qwen3.8 Scores 23.57% on Addition in Words in Willison’s No-Reasoning Test
A local benchmark separates correct formatting from correct arithmetic. A much smaller reasoning-enabled run performed better, but took longer and used only one example per comparison.
The experiment highlights a gap between producing an answer in the requested format and getting the underlying calculation right: Qwen3.8-27B achieved 96.17% word-only compliance but just 23.57% numeric accuracy without reasoning. Accuracy fell to 6.44% on operands with 10–13 digits, while a medium-reasoning run got 167 of 169 answers right. That comparison is promising but not conclusive: reasoning took longer, and each test cell had only one attempt, so Willison cautioned that results could vary on another run.
01
Willison tested 5,070 addition cases on a DGX Spark, with 30 attempts for each operand-length combination.
02
Without reasoning, accuracy was 97.04% for one-to-three-digit operands and 6.44% for ten-to-thirteen-digit operands.
03
The reasoning-enabled comparison used 169 attempts—one per test cell—so its 167 correct answers are based on a much smaller sample.
An AI model can obey the instruction to answer in English words while getting the sum wrong. In research published October 4, Simon Willison found that a local Qwen3.8-27B model answered just 23.57% of addition problems correctly with reasoning disabled. A smaller reasoning-enabled comparison produced a sharply different result: 167 correct answers in 169 attempts.
Willison ran the experiment on a DGX Spark using the local model file Qwen3.8-27B-Q4_K_M.gguf. The task was narrow: add positive integers and express the exact result solely in English words. That made success depend on two separate requirements—calculating the right number and returning it in the requested form.
The idea came from Colin Frasier’s earlier experiment with GPT-4o, which examined how well it could calculate sums and write the answers in words as the numbers grew. After Frasier shared a chart on Bluesky, Willison used it to guide a new experiment on local hardware, where he wanted a fully controlled environment.
The format held up; the arithmetic did not
The reasoning-disabled run covered 5,070 cases, with 30 attempts for each combination of operand lengths—the number of digits in the two numbers being added. The model followed the word-only output requirement in 96.17% of cases. Its much lower numeric accuracy shows why passing a formatting check would not have been enough to judge these answers.
Two different measures of success
96.17%Word-only format compliance
Most reasoning-disabled answers followed the required output format.
23.57%Numeric accuracy
Fewer than a quarter of answers represented the correct sum across the same 5,070-case run.
Number size made a substantial difference. Without reasoning, accuracy was 97.04% for operands with one to three digits, but just 6.44% for operands with ten to thirteen digits. The aggregate score therefore masks a pronounced split: strong results on short numbers and poor results on much longer ones.
More working, fewer samples
Willison then enabled reasoning and ran a medium-reasoning comparison. Each pair took much longer, so he reduced the test from 30 samples per grid cell to one, for 169 attempts altogether. The model missed only two. Willison cautioned that another run would produce different results because each cell contained a single attempt.
The published reasoning traces offer a concrete view of the model’s working. In one large calculation, it pauses to redo the sum, aligns the digits and adds from right to left. It explicitly records a carry after adding six and nine, rather than jumping straight to the final wording.
Editorial illustration for Qwen3.8 Scores 23.57% on Addition in Words in Willison’s No-Reasoning Test.Source: simonwillison.net.
Reader comments
Newest comments first. Replies stay oldest first.