Qwen3.8 27B addition in words

Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to test how well it could "compute the sum but return the answer in words" across increasingly large numbers. He shared a chart of those results.
Simon Willison wrote that he was inspired to run the experiment again on local hardware, a DGX Spark, to explore the effect in a fully controlled environment. He pasted Frasier's image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf.
The first run used 30 attempts per combination with reasoning disabled. He then ran it again with reasoning enabled. That took much longer per pair, so instead of 30 samples per square he ran just one, producing a less visually appealing heatmap in which each square is either 100% or 0%.
With reasoning enabled, the model got the right answer in 167 out of 169 attempts. Willison noted that because these were one-shot, a second run would likely produce different results.
He also shared a version of the report including reasoning traces from some of the larger calculations. One trace shows the model aligning two numbers and adding from right to left, tracking carries position by position.
Willison wrote that he is confident GPT-4o did not cheat and use a calculator, especially since it got so many of the calculations wrong.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controll