SyncAI.news, a Varaisys broadcasting
MDN-Control: Mask-Depth-Noise Guided Region Control for Multi-Subject Video Editing
JY

Jiayi Yu, Xi Ye, Lina Wang, Yunkun Xia

· 1 min read

ResearcharXiv cs.CV

MDN-Control: Mask-Depth-Noise Guided Region Control for Multi-Subject Video Editing

arXiv:2609.16475v1 Announce Type: new Abstract: Multi subject video editing modifies designated subjects while preserving non target content, but faces cross subject attribute leakage, and occlusion ambiguity. Existing approaches rely on masks and struggle to distinguish overlapping subjects or ensure consistent generation. To address these limitations, we propose MDN-Control, a training free framework jointly controlling target localization, occlusion geometry, and appearance initialization. Specifically, mask-guided localization provides consistent target localization, while depth-aware occlusion control resolves ambiguous boundaries between overlapping subjects. We further introduce noise latent prompting, which retrieves Gaussian initializations from a noise library for prompt relevant priors. Experiments on MSVBench show that MDN-Control achieves the lowest CM-Err and the highest Q-Edit, while maintaining competitive text alignment and temporal consistency, demonstrating the effectiveness of combining spatial, geometric, and latent priors for multi subject video editing.

Original source

This story was published by arXiv cs.CV and written by Jiayi Yu, Xi Ye, Lina Wang, Yunkun Xia. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News