TLDRocket
Sign in

RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

Apple Machine Learning Research

RLTL;DR proposes a reinforcement-learning method that uses the verifier’s output after each failed attempt to generate self-written feedback for the agent. The paper’s approach is to have the agent write a single TL;DR insight per failure that conditions the next rollout. As a result, the agent can self-improve even when success is rare and there are no teacher models or example solutions available.

Why it matters

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned…

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.