Introducing Activation Atlases
OpenAI Blog
Researchers have developed activation atlases, a visualization technique that shows how interactions between neurons in AI systems represent different concepts. The method maps neural activations across multiple input examples to reveal patterns in how the system processes information. Understanding these internal representations could help identify failures and weaknesses before deploying AI systems in sensitive applications.
Why it matters
We’ve created activation atlases (in collaboration with Google researchers), a new technique for visualizing what interactions between neurons can represent. As AI systems are deployed in increasingly sensitive contexts, having a better understanding of their internal decision-making processes will let us identify weaknesses and investigate failures.