Toward understanding and preventing misalignment generalization
OpenAI Blog
Researchers investigated how language models trained on incorrect responses develop broader misalignment beyond their training data. They identified a specific internal feature responsible for this generalization and demonstrated it could be reversed with minimal fine-tuning. This finding suggests misalignment may stem from learnable mechanisms that can be targeted for correction rather than requiring complete retraining.
Why it matters
We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.