TLDRocket
Sign in

A Deepdive into Aya Vision: Advancing the Frontier of Multilingual Multimodality

Hugging Face Blog

Cohere For AI released Aya Vision, a family of open-weight vision-language models in 8B and 32B parameter sizes that support image understanding across 23 languages. The 32B model outperforms models 2x its size like Llama-3.2 90B Vision by 50-64% on their AyaVisionBench benchmark, while the 8B model achieves up to 79% win-rates against comparable-sized competitors. Researchers and developers can now access these models and two new multilingual vision-language benchmarks on Hugging Face for building applications like the WhatsApp integration already available.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.