TLDRocket
Sign in

Reinforcement learning with prediction-based rewards

OpenAI Blog

Researchers developed Random Network Distillation, a reinforcement learning method that uses prediction errors as curiosity rewards to drive exploration in agents. The approach achieved higher average scores than human players on Montezuma's Revenge, a notoriously difficult exploration-heavy video game. This enables RL agents to solve environments where traditional reward signals provide insufficient guidance for discovery.

Why it matters

We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments through curiosity, which for the first time exceeds average human performance on Montezuma’s Revenge.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.