Learning complex goals with iterated amplification
OpenAI Blog
Researchers proposed iterated amplification, an AI safety technique that breaks down complex tasks into simpler subtasks rather than relying on labeled data or reward functions to train AI systems. The approach has only been tested on toy algorithmic problems so far, with no concrete benchmarks or performance metrics published. If successful at scale, the method could address the challenge of specifying AI behavior for goals too complicated for humans to directly specify.
Why it matters
We’re proposing an AI safety technique called iterated amplification that lets us specify complicated behaviors and goals that are beyond human scale, by demonstrating how to decompose a task into simpler sub-tasks, rather than by providing labeled data or a reward function. Although this idea is in its very early stages and we have only completed experiments on simple toy algorithmic domains, we’ve decided to present it in its preliminary state because we think it could prove to be a scalable approach to AI safety.