
TL
Tung Le, Huy Tien Nguyen, Le Minh Nguyen
· 1 min read
ResearcharXiv cs.CV
Vision And Text Transformer For Predicting Answerability On Visual Question Answering
arXiv:2609.16565v1 Announce Type: new
Abstract: Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.
Original source
This story was published by arXiv cs.CV and written by Tung Le, Huy Tien Nguyen, Le Minh Nguyen. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


