Controlling Reasoning Effort in LLMs
Ahead of AI published an article titled "Controlling Reasoning Effort in LLMs," describing how reasoning models with multiple effort settings are developed. The article notes that OpenAI released the GPT-5.6 model family last week, with three sizes, each offering roughly five or six reasoning-effort settings.
The article traces the lineage of reasoning models to OpenAI's o1, released almost two years ago, and DeepSeek-R1, which followed about four months later with details of a reinforcement learning with verifiable rewards (RLVR) recipe. It defines a reasoning model as one that outputs an intermediate reasoning trace working through a task step by step, and notes that most of today's LLMs are effectively reasoning models trained in a similar fashion using a form of RLVR.
On implementation, the article states that OpenAI has not shared details of its effort settings. It cites the open-source gpt-oss models as evidence that OpenAI allows reasoning effort to be toggled via a system prompt reading "Reasoning effort: low/medium/high" prepended to each prompt, and says GPT-5 models presumably use a similar approach. It describes two possible training implementations: applying different length penalties during RLVR depending on the system prompt, or fine-tuning after RLVR via supervised fine-tuning on prompts paired with target responses of the desired reasoning length.
As a case study, the article cites the newly released Inkling technical report from Thinking Machine Labs. During large-scale RL, it says, the desired effort level was specified in the system message and the cost assigned to each generated token was adjusted, with low effort using a larger per-token cost and high effort a smaller one. At inference time, Inkling receives a system message such as "Thinking effort level: 0.8" and adjusts token usage. Unlike gpt-oss and GPT-5.6, Inkling's effort label is a continuous number between 0 and 1 rather than ordinal labels. The article states Inkling does not disclose its exact reward formula, token-cost coefficients, or whether effort conditioning was included in SFT.
The article also covers Qwen3's on/off thinking toggle, implemented primarily through supervised fine-tuning and reinforced during general RL.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes