Our approach to alignment research
OpenAI Blog
Anthropic is developing techniques for AI systems to learn from human feedback and help humans evaluate AI behavior. The company aims to create an AI system well-aligned enough to assist in solving remaining alignment challenges. This approach assumes that a sufficiently aligned AI could accelerate progress on broader AI safety problems.
Why it matters
We are improving our AI systems’ ability to learn from human feedback and to assist humans at evaluating AI. Our goal is to build a sufficiently aligned AI system that can help us solve all other alignment problems.