TLDRocket
Sign in

Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?

Hugging Face Blog

Researchers at UCLA created ConTextual, a benchmark dataset with 506 instructions designed to evaluate how well multimodal AI models can reason jointly about text and images in text-rich scenes like maps, shopping interfaces, and infographics. The dataset covers 8 real-world visual scenarios, and initial experiments tested 13 models including GPT-4V, Gemini Vision Pro, and open-source alternatives like LLaVA-1.5-13B. Current models substantially underperform humans on the benchmark, with even the best proprietary model GPT-4V struggling on time-reading and infographic tasks, suggesting the need for better image encoders and vision-language alignment techniques.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.