
JY
Jiayi Yu, Xi Ye, Lina Wang, Yunkun Xia
· 1 min read
ResearcharXiv cs.CV
MDN-Control: Mask-Depth-Noise Guided Region Control for Multi-Subject Video Editing
arXiv:2609.16475v1 Announce Type: new
Abstract: Multi subject video editing modifies designated subjects while preserving non target content, but faces cross subject attribute leakage, and occlusion ambiguity. Existing approaches rely on masks and struggle to distinguish overlapping subjects or ensure consistent generation. To address these limitations, we propose MDN-Control, a training free framework jointly controlling target localization, occlusion geometry, and appearance initialization. Specifically, mask-guided localization provides consistent target localization, while depth-aware occlusion control resolves ambiguous boundaries between overlapping subjects. We further introduce noise latent prompting, which retrieves Gaussian initializations from a noise library for prompt relevant priors. Experiments on MSVBench show that MDN-Control achieves the lowest CM-Err and the highest Q-Edit, while maintaining competitive text alignment and temporal consistency, demonstrating the effectiveness of combining spatial, geometric, and latent priors for multi subject video editing.
Original source
This story was published by arXiv cs.CV and written by Jiayi Yu, Xi Ye, Lina Wang, Yunkun Xia. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


