Finetune Stable Diffusion Models with DDPO via TRL
Hugging Face
Hugging Face's TRL library now lets you fine-tune Stable Diffusion with reinforcement learning, using a method called DDPO. No human preference labels needed at the final training step, just a reward model.
Based on reporting by Hugging Face — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Diffusion models are great at making pretty pictures, but
Read more about this at: Hugging Face