Fine-tuning open LLM judges to outperform GPT-5.2
Together AI
Researchers fine-tuned open-source LLM judges using Direct Preference Optimization on 5,400 preference pairs from RewardBench 2, achieving higher accuracy than GPT-5.2 on human preference alignment. GPT-OSS 120B reached 62.63% accuracy compared to GPT-5.2's 61.62%, while costing 15.3 times less and running 14 times faster. This demonstrates that smaller open-source models can match or exceed closed-source judge performance when trained on preference data, enabling cost-effective and deployable evaluation systems.
Why it matters
Fine-tuned open-source LLM judges can outperform GPT-5.2 at evaluating model outputs. Using Direct Preference Optimization on just 5,400 preference pairs, we trained GPT-OSS 120B to beat GPT-5.2 on human preference alignment—at 15x lower cost and 14x faster inference speeds.