Improving instruction hierarchy in frontier LLMs
OpenAI Blog
Researchers developed IH-Challenge, a training method that teaches large language models to prioritize instructions from trusted sources over conflicting inputs. The approach improved instruction hierarchy performance by enabling models to better distinguish between legitimate directives and injected prompts. Models trained with this method showed increased resistance to prompt injection attacks and greater safety control without sacrificing general capabilities.
Why it matters
IH-Challenge trains models to prioritize trusted instructions, improving instruction hierarchy, safety steerability, and resistance to prompt injection attacks.