TLDRocket
Sign in

Japanese Stable Diffusion

Hugging Face Blog

rinna Co., Ltd. fine-tuned Stable Diffusion on approximately 100 million Japanese-captioned images to create a text-to-image model that accepts Japanese prompts and generates culturally appropriate images. The Japanese dataset used was 1/20th the size of the English training data, requiring a two-stage training approach that first replaced the text encoder and then fine-tuned both the encoder and diffusion model jointly. Japanese Stable Diffusion can now interpret Japanese-specific terms like "salary man" and cultural expressions that don't translate well to English, enabling generation of images reflecting Japanese culture rather than Western aesthetics.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.