TLDRocket
Sign in

Finetune Stable Diffusion Models with DDPO via TRL

Hugging Face Blog

Hugging Face's TRL library integrated DDPO (Denoising Diffusion Policy Optimization), a reinforcement learning method for fine-tuning diffusion models like Stable Diffusion to align outputs with human preferences. The method requires an A100 GPU minimum and uses a reward model trained on aesthetic preferences, with recommended hyperparameters including a learning rate of 3e-4 and training batch size of 3 for single-GPU setups. Users can now fine-tune Stable Diffusion models to generate images matching specific objectives such as visual aesthetics without needing the supervised fine-tuning step typically required in RLHF workflows.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.