
ZW
Zitong Wang, Zijun Shen, Haohao Xu, Zhengjie Luo, Weibin Wu
· 1 min read
ResearcharXiv cs.CV
Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation
arXiv:2603.10210v2 Announce Type: replace
Abstract: While Diffusion Models excel in text-to-image synthesis, they frequently suffer from catastrophic concept omission when generating complex multi-instance scenes. Existing training-free methods attempt to resolve this by rescaling attention maps, which merely exacerbates unstructured noise without establishing coherent semantic representations. To address this, we propose Delta-K, a backbone-agnostic, plug-and-play inference framework that resolves omission by operating directly in the shared cross-attention Key space. Utilizing a lightweight Vision-Language Model (VLM) preview, we isolate a differential key ($\Delta K$) capturing the pure semantic signature of missing concepts, and proactively inject it during the early semantic planning phase. Governed by a dynamically optimized scheduling mechanism, Delta-K grounds diffuse noise into stable structural anchors while naturally preserving existing concepts via the inherent orthogonality of $\Delta K$. Extensive experiments validate its universal applicability, demonstrating that Delta-K significantly improves compositional alignment across both modern DiT and foundational U-Net architectures without requiring spatial masks, auxiliary training, or structural modifications.
Original source
This story was published by arXiv cs.CV and written by Zitong Wang, Zijun Shen, Haohao Xu, Zhengjie Luo, Weibin Wu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


