RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback
Apple Machine Learning Research published a paper titled "RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback," authored by Michael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry, Aryo Lotfi, Abbas Kazerouni, Omar Attia, Sanjoy Chowdhury and Alexander Toshev. The work is listed under Methods and Algorithms and a NeurIPS conference paper, published October 2026.
The paper addresses reinforcement learning with verifiable rewards (RLVR), where agents make multiple attempts at a task and optimize toward successful ones. The authors state this becomes problematic in self-improvement settings where tasks are so difficult that the agent has a low or even no chance of success, and where no teacher models or example solutions are available to distill from.
RLTL;DR shows the policy the verifier outputs after each failed attempt and lets it write its own feedback as a single TL;DR insight. The next rollout is conditioned on all previous insights, and rollouts are sampled sequentially until a solution is found. The authors also enable backpropagation on the in-context insights to internalize a direct task-to-insight mapping.
On challenging tool-calling and coding datasets filtered to Pass@128 = 0, the paper reports that standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR achieves a Pass@1 of 14-31% with insights in context during training and 12-13% when no insight is in context at evaluation time. The authors identify task-to-insight internalization as the key.
The paper also describes a reduced approach, SFTL;DR, which trains only on (task, insight) tuples without showing or backpropagating on rollouts. Training on only 4k such tuples recovered almost the full performance of RLTL;DR and classical SFT on full rollouts, which the authors present as a compacted training paradigm and hope will inspire future research.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own f