TLDRocket
Sign in

Welcome aMUSEd: Efficient Text-to-Image Generation

Hugging Face Blog

aMUSEd is an open-source text-to-image model that uses Masked Image Modeling instead of diffusion, requiring fewer inference steps to generate images. The model contains 800 million parameters and achieves significantly faster inference latencies compared to diffusion-based systems like SDXL, while also enabling zero-shot image inpainting without additional fine-tuning. The release encourages community exploration of non-diffusion approaches for image generation, with simplified fine-tuning possible on consumer GPUs using 11GB of VRAM or 7GB with LoRA optimization.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.