
Hugging Face Blog
· 1 min read
A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality
With the release of the Aya Vision family, our new 8B and 32B parameter vision-language models (VLMs), we are addressing one of the biggest challenges in AI: bringing multilingual performance to multimodal models.
Aya Vision is Cohere For AI's latest open-weight multilingual and multimodal model family, designed to be a strong foundation for language and vision understanding across 23 languages. It builds on the success of Aya Expanse, state-of-the-art multilingual language models, and extends it using a combination of advanced techniques. These include synthetic annotations, scaling up multilingual data through translation and rephrasing, and multimodal model merging – key methods that improve both language and vision understanding in a multilingual setting.
As a result, our models perform well in a variety of tasks, including image captioning, visual question answering, text generation, and translating both text and images into clear, natural-language text. We evaluated Aya Vision models on a set of datasets, including our new open-ended vision-language benchmark AyaVisionBench and a multilingual version of Wild Vision Bench (mWildVision) that is translated into 23 languages, which we release both of them for research.
In pair-wise comparison, Aya Vision 32B outperforms models more than 2x of its size, such as Llama-3.2 90B Vision, Molmo 72B, and Qwen2.5-VL 72B by win rates ranging from 50% to 64% on AyaVisionBench and 52% to 72% on mWildVision average across 23 languages.
Our compact and more efficient model Aya Vision 8B achieves the best performance in multilingual multimodal in its parameter class, outperforming leading models such as Qwen2.5-VL 7B, Pixtral 12B, Gemini Flash 1.5 8B, Llama-3.2 11B Vision, Molmo-D 7B, and Pangea 7B by up to 79% win-rates on AyaVisionBench and 81% on mWildBench.
Aya Vision Architecture and Training
Training process
Multimodal Data Enhancement and Expanding Language Coverage
Multimodal Model Merging
Scaling up to 32B
To get started:
Original source
This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on huggingface.co


