SyncAI.news, a Varaisys broadcasting
Image Recognition with Vision and Language Embeddings of VLMs
IV

Illia Volkov, Nikita Kisel, Klara Janouskova, Jiri Matas

· 1 min read

ResearcharXiv cs.CV

Image Recognition with Vision and Language Embeddings of VLMs

arXiv:2509.09311v2 Announce Type: replace Abstract: Vision-language models (VLMs) have enabled strong zero-shot classification through image-text alignment. Yet, their purely visual inference capabilities remain under-explored. In this work, we conduct a comprehensive evaluation of both language-guided and vision-only image classification with a diverse set of dual-encoder VLMs, including both well-established and recent models such as SigLIP 2 and RADIOv2.5. The performance is compared in a standard setup on the ImageNet-1k validation set and its label-corrected variant. The key factors affecting accuracy are analysed, including prompt design, class diversity, the number of neighbours in k-NN, and reference set size. We show that language and vision offer complementary strengths, with some classes favouring textual prompts and others better handled by visual similarity. To exploit this complementarity, we introduce a simple, learning-free fusion method based on per-class precision that improves classification performance. The code is available at: https://github.com/gonikisgo/bmvc2025-vlm-image-recognition.

Original source

This story was published by arXiv cs.CV and written by Illia Volkov, Nikita Kisel, Klara Janouskova, Jiri Matas. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News