Apple Researchers Introduce RLTL;DR to Enable Self-Improvement via Internalized Feedback
Key point
Apple researchers introduce RLTL;DR, a method that boosts Pass@1 scores from near-zero to 12–13% on difficult tasks by internalizing self-generated feedback.
Details
Apple researchers introduced RLTL;DR, a reinforcement learning method designed for self-improvement in scenarios where agents have low or zero chance of success and no teacher models are available. Standard Reinforcement Learning with Verifiable Rewards (RLVR) struggles when tasks are too difficult for the agent to solve even once, leaving no successful trajectories to optimize against.
How RLTL;DR Works
Instead of relying on successful attempts, RLTL;DR conditions the policy on self-generated feedback. After each failed attempt, the agent writes a concise TL;DR insight based on the verifier's output. Subsequent rollouts are conditioned on all previous insights, and the model backpropagates on these in-context insights to internalize a direct task → insight mapping.
Performance Results
On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy remained flat at 0% to 1% Pass@1. RLTL;DR broke this barrier, achieving:
- 14–31% Pass@1 with insights in context during training.
- 12–13% Pass@1 when no insight was in context at evaluation time.
The SFTL;DR Variant
The researchers identified that the key mechanism was the internalization of the task-to-insight mapping. They tested a simplified variant, SFTL;DR, which trains only on (task, insight) tuples without showing or backpropagating on any rollouts. Training on just 4k of these tuples recovered almost the full performance of RLTL;DR, suggesting a promising compacted training paradigm where agents learn to "keep this sort of thing in mind" for specific task types.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.