TLDRocket
Sign in

Safety Research

15 summarised stories about Safety Research, each linking back to the original source. Browse all topics →

Tuesday, 13 June 2017

Learning from human preferences

OpenAI Blog 9 years ago

Researchers at Anthropic and DeepMind have developed an algorithm that learns human preferences by comparing pairs of proposed behaviors rather than requiring explicit goal functions to be written. The system evaluates which of two actions a human prefers, using this feedback to infer the underlying objective. This approach reduces risks from misaligned proxy goals or incorrectly specified objectives in AI systems.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.