SyncAI.news, a Varaisys broadcasting
Vision And Text Transformer For Predicting Answerability On Visual Question Answering
TL

Tung Le, Huy Tien Nguyen, Le Minh Nguyen

· 1 min read

ResearcharXiv cs.CV

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

arXiv:2609.16565v1 Announce Type: new Abstract: Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.

Original source

This story was published by arXiv cs.CV and written by Tung Le, Huy Tien Nguyen, Le Minh Nguyen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News