TLDRocket
Sign in

Language models can explain neurons in language models

OpenAI

OpenAI used GPT-4 to write explanations for what individual neurons inside GPT-2 are actually doing. It's a step toward cracking open the black box, though the explanations are still pretty rough.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has a new party trick: pointing one language model at another and asking it to explain itself. Specifically, the company used GPT-4 to generate natural-language explanations for the behavior of neurons scattered throughout GPT-2, then had GPT-4 grade its own explanations for accuracy. The result is a full dataset covering every neuron in GPT-2, now released publicly.

The basic idea isn't new — researchers have been trying to interpret individual neurons in neural networks for years, usually by hand, one at a time, which does not scale when your model has hundreds of thousands of them. GPT-4 automates the tedious part. Feed it examples of when a neuron fires strongly, and it drafts a guess about what pattern the neuron is detecting — something like 'this neuron activates on mentions of Canadian cities' or 'this fires on movie references.' Then a second pass checks how well that explanation predicts the neuron's actual behavior on new text.

The catch, and OpenAI is upfront about it, is that most of these explanations are mediocre. GPT-4 tends to do fine on neurons with a single, obvious job and falls apart on neurons that seem to be doing several unrelated things at once — which, it turns out, describes a lot of neurons in a model like GPT-2. Superposition, the phenomenon where individual units encode multiple overlapping concepts, remains the core obstacle, and no amount of clever prompting fully gets around it.

Still, there's something notable in the method itself. This is one of the more concrete examples of using a more capable model to help audit a smaller one, a technique that starts to matter a lot once models get too large and too alien for humans to inspect line by line. GPT-2 is small and mostly harmless by 2024 standards, so the stakes here are low. But the pattern — bigger model interprets smaller model's internals, at scale, without a human reading every neuron — is the part worth paying attention to, regardless of how rough today's explanations are.

My take — AI-written commentary, not fact-checked reporting

I like this less as a GPT-2 curiosity and more as a preview of the only realistic path to interpretability at scale: models auditing models, because no research team will ever hand-label a trillion neurons. The explanations are mediocre now, sure, but mediocre-and-automatable beats perfect-and-impossible, and that's the trade every safety team is quietly betting on.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.