Limits of Confidence in Diffusion

Researchers Russ Webb, Amitis Shidani, Alice Bizeul and Dan Busbridge present a theoretical analysis of discrete diffusion, covering remasking and uniform-state samplers. Such samplers generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions.
The authors note that domains of general interest, including pixels, phonemes and words, contain inherent dependencies between tokens. They show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, and that no product of per-position distributions can match a dependent group. They further state that per-position distributions do not determine whether a group is dependent: two joint distributions can have identical per-position marginals while differing in which combinations of values occur.
The work reports experiments on ScanAndAdd, described as a synthetic task whose joint distribution is available in closed form. On that task, the researchers verify that every group of two or more undetermined positions written by a confidence ranking is dependent. They measure the generated distribution at 29 times the sampling-noise floor in total variation, while per-sample metrics are 1.0.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write from those same distributions. For domains of general interest (pixels, phonemes, or words) there are inherent dependencies between tokens. We show that a step matches the training distribution only when the positions it writes are conditionally independent given the tokens already fixed, that no prod