Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Hugging Face
LFM2.5-350M was fine-tuned with GRPO using TRL for structured-output compliance and then re-evaluated on IFStruct. The run used about 500 samples and 100 training steps, improving the IFStruct pass rate from 22.6% to 29.7%. JSON compliance increased from 18.0% to 31.9% while YAML stayed roughly flat, indicating task-specific fine-tuning shifts the model toward the targeted output formats.