On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study
Apple Machine Learning Research
A systematic study evaluates multiple large-language-model conditioning methods for injecting or removing a target concept and compares them across both effectiveness and generation quality. The paper reports that efficient steering methods often achieve conditioning but at a steep cost to fluency, and activation steering is far less effective on instruction-tuned models than on base counterparts. As a result, it recommends simpler prompting or supervised fine-tuning for concept injection, notes they are weaker for concept removal, and uses cheaply computed text metrics that track LLM-judge scores to guide method choice.
Why it matters
Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting or removing a target concept, neglecting generation quality. We systematically investigate a range of conditioning methods in both injection and removal scenarios. We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet…