
ZL
Zhixuan Li, Yujia Liu, Chen Hui, Chenyue Song, Weisi Lin
· 1 min read
ResearcharXiv cs.CV
Single Point, Full Mask: Velocity-Guided Level Set Evolution for End-to-End Amodal Segmentation
arXiv:2508.01661v2 Announce Type: replace
Abstract: Amodal segmentation aims to recover complete object shapes, including occluded regions, serving as an essential technique for user-centric multimedia authoring and object-level visual manipulation. Existing methods typically rely on informative prompts, such as bounding boxes or dense visible masks, which heavily degrade the user experience and interaction efficiency in real-world multimedia applications. While recent interactive paradigms (e.g., the Segment Anything Model) support lightweight point-based interactions, they often perform direct mask regression. Crucially, the opaque nature of these direct-regression models offers no visual explainability regarding how occluded structures are inferred, conflicting with the growing demand for interpretable multimedia systems. To address these limitations, we propose VELA, an end-to-end VElocity-driven Level-set Amodal segmentation method that enables explicit and transparent contour evolution driven by simple point clicks. VELA constructs an initial level set function from visual features and the user's point input, which then progressively evolves into the final amodal mask under the guidance of a shape-specific motion field predicted by a fully differentiable network. This mechanism learns to generate evolution dynamics at each step, ensuring that the spatial reasoning process is geometrically grounded, topologically flexible, and visually explainable to the user. Extensive experiments on COCOA-cls, D2SA, and KINS benchmarks demonstrate that VELA outperforms existing methods that use bounding-box or dense visible-mask prompts while requiring only a single-point prompt, validating the effectiveness of explainable geometric modeling for interactive multimedia tasks.
Original source
This story was published by arXiv cs.CV and written by Zhixuan Li, Yujia Liu, Chen Hui, Chenyue Song, Weisi Lin. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


