Why We Think
Lilian Weng published a post reviewing recent developments in test-time compute, or "thinking time," and why it improves model performance. She credits John Schulman for feedback and direct edits on the post.
The review frames longer model thinking through several lenses. One is an analogy to Daniel Kahneman's dual process theory in Thinking, Fast and Slow (2013), distinguishing fast, automatic System 1 thinking from deliberate System 2 thinking. Another treats computation as a resource, noting that in Transformer models the computation per generated token is roughly twice the number of parameters, while sparse models such as mixture of experts use only a fraction of parameters per forward pass.
The post also presents latent variable modeling, in which a hidden variable z is marginalized to express a distribution over visible variables y, and surveys work on thinking in tokens, including the AQUA-RAT dataset (Ling et al. 2017), the Grade School Math dataset (Cobbe et al. 2021), scratchpad intermediate tokens (Nye et al. 2021), and the term chain-of-thought (Wei et al. 2022).
Covered methods include parallel sampling techniques such as best-of-N, beam search, and self-consistency, and sequential revision, which the post says can fail without external feedback. Reinforcement learning for reasoning is discussed through DeepSeek-R1 (DeepSeek-AI, 2025), which uses cold-start supervised fine-tuning followed by reasoning-oriented RL with format and accuracy rewards. The post also addresses faithful thinking, optimization pressure on chain-of-thought, thinking in continuous space, and scaling laws for thinking time.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post. Test time compute ( Graves et al. 2016 , Ling, et al. 2017 , Cobbe et al. 2021 ) and Chain-of-thought (CoT) ( Wei et al. 2022 , Nye et al. 2021 ), have led to significant improvements in model performance, while raising many research questions. This post aims to review recent developments in how to effectively use test-time compute (i.e. “thinking time”) and why it helps.