SyncAI.news, a Varaisys broadcasting
Adapting Vision-Language Models for Human-Readable XAI in Industrial Object Detection
SS

Sarvenaz Sardari, Freddy Fernandes, Samarth Yelvande, Jose Moises Araya-Martinez, Alina Roitberg

· 1 min read

ResearcharXiv cs.CV

Adapting Vision-Language Models for Human-Readable XAI in Industrial Object Detection

arXiv:2609.31690v1 Announce Type: new Abstract: Explainable Artificial Intelligence (XAI) solutions are essential for building trust in AI technologies and their integration in real manufacturing lines. However, most existing methods are tailored to technical experts, limiting their accessibility to diverse user groups such as blue-collar workers in manufacturing lines who use AI for quality control. In this work, we introduce an XAI interface for object detection in industrial manufacturing based on a fine-tuned vision-language model, designed to generate intuitive explanations for non-expert users. We benchmark existing vision-language models and demonstrate that out-of-the-box models often fall short in delivering clear, context-relevant explanations for non-expert users. To address this, we fine-tune a vision-language model and integrate it into our interface, enabling contextualized, accessible explanations for non-expert users. We demonstrate improvements in explanation clarity, instruction adherence, image groundedness, and contextual awareness over GPT 4o-mini on proprietary and public robotics dataset. This approach advances the accessibility and usability of AI explanations, making them more intuitive and applicable in manufacturing domain.

Original source

This story was published by arXiv cs.CV and written by Sarvenaz Sardari, Freddy Fernandes, Samarth Yelvande, Jose Moises Araya-Martinez, Alina Roitberg. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News