Interpretable machine learning through teaching
OpenAI Blog
Researchers developed a method where AI systems teach each other using examples selected to be interpretable to humans as well as informative to machines. The approach automatically identifies the most instructive examples—such as representative images—to convey a concept like "dogs." This technique enables AI models to learn concepts while maintaining human-understandable reasoning about what makes those examples representative.
Why it matters
We’ve designed a method that encourages AIs to teach each other with examples that also make sense to humans. Our approach automatically selects the most informative examples to teach a concept—for instance, the best images to describe the concept of dogs—and experimentally we found our approach to be effective at teaching both AIs