Local Mechanisms of Compositional Generalization in Conditional Diffusion
Key point
This work showed on CLEVR and SDXL that compositional generalization in Conditional Diffusion corresponds exactly to local conditional scores.
Details
The study explored whether Conditional Diffusion models handle conditioner combinations outside the training distribution, that is, whether they actually learn compositional generalization. This was concretized as length generalization, testing on CLEVR the ability to generate more objects than seen during training.
Results differed across models. Successful models showed local conditional scores that were sparsely connected with respect to input pixels and conditioners, while failing models did not. In other words, compositional generalization was not a default property of all models, but appeared only when a specific structure was learned.
Theoretically, it was proven that a specific compositional structure, conditional projective composition, is exactly equivalent to local conditional scores. This framework extends to cases of composing concepts in feature-space, such as style+content.
- When causal intervention was used to forcibly inject local conditional scores, models that had previously failed were able to perform length generalization.
- In the pixel-space of SDXL, spatial locality was observed, but locality with respect to conditioners was largely weak.
- In contrast, quantitative evidence of local conditional scores was confirmed in the network's feature-space.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.