Apple Research: Confidence Sampling in Diffusion Models Cannot Match Training Distribution Due to Token Dependencies
Key point
Apple research shows confidence-based sampling in discrete diffusion models selects dependent token groups, causing a 29x deviation from the training distribution despite perfect per-sample metrics.
Details
Discrete diffusion models, such as remasking and uniform-state samplers, generate sequences by writing multiple token positions per step. These models draw each token from a per-position distribution and use those same distributions to determine which positions to write. However, domains like pixels, phonemes, and words contain inherent dependencies between tokens.
Theoretical Limitations
The research demonstrates that a sampling step only matches the training distribution if the positions written are conditionally independent given the already fixed tokens. No product of per-position distributions can match a dependent group. Furthermore, per-position distributions do not determine whether a group is dependent; two joint distributions can have identical per-position marginals while differing in which combinations of values occur.
Empirical Evidence from ScanAndAdd
On ScanAndAdd, a synthetic task with a known joint distribution, the authors verified that every group of two or more undetermined positions selected by a confidence ranking is dependent. The generated distribution's total variation was measured to be 29× the sampling-noise floor. This distributional error persists even though per-sample metrics remain at 1.0, indicating that standard evaluation methods may miss these fundamental distributional mismatches.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.