Limits of Confidence in Diffusion
Apple Machine Learning Research
The paper shows that discrete diffusion steps only match the training distribution when the positions written are conditionally independent given the already-fixed tokens, and that products of per-position distributions cannot reproduce dependent token groups. On the ScanAndAdd task, it measures the generated distribution at 29× the sampling-noise floor total variation while per-sample metrics are 1.0. This limits how confidence rankings that write per-position distributions can correct dependence errors, even when the task’s joint distribution is known.
Why it matters
Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes, or words) there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no product of per-position distributions can match a dependent group, and that…