TLDRocket
Sign in

Taming Outlier Tokens in Diffusion Transformers

Apple Machine Learning Research

Researchers found that Diffusion Transformers for image generation develop outlier tokens—high-norm tokens that attract attention despite carrying little useful information—in both their encoder and denoising components. The outlier phenomenon appears particularly in intermediate layers of DiTs and in pretrained ViT encoders within RAE-DiT pipelines. Understanding and controlling these outliers could improve the efficiency and quality of transformer-based image generation models.

Why it matters

We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models remains underexplored. We show that this phenomenon appears in both the encoder and denoiser of modern Representation Autoencoder (RAE)-DiT pipelines: pretrained ViT encoders can produce outlier representations, and DiTs themselves can develop internal outlier tokens, especially in intermediate layers…

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.