SyncAI.news, a Varaisys broadcasting
Targeted Visual Counterfactual Explanations for Contrastive Vision-Language Model
VB

Van Bach Nguyen, J\"org Schl\"otterer, Christin Seifer

· 1 min read

ResearcharXiv cs.CV

Targeted Visual Counterfactual Explanations for Contrastive Vision-Language Model

arXiv:2609.37638v1 Announce Type: new Abstract: Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textbf{A}daptive \textbf{C}ounterfactual \textbf{E}xplanations (\mace), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. \mace constructs an editable region from either source attribution or source--target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate \mace on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel-level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.

Original source

This story was published by arXiv cs.CV and written by Van Bach Nguyen, J\"org Schl\"otterer, Christin Seifer. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News