Can LLMs invent better ways to train LLMs?
Sakana AI
Sakana AI used LLMs to automatically discover new preference optimization algorithms for training other LLMs, a process they call LLM². They discovered Discovered Preference Optimization (DiscoPOP), which outperforms existing methods like DPO across multiple benchmarks. This approach reduces reliance on human researchers to manually design training algorithms and creates a self-referential feedback loop where AI improvements can accelerate future AI development.