TLDRocket
Sign in

PaliGemma 2 Mix - New Instruction Vision Language Models by Google

Hugging Face Blog Covered by 2 sources

Google released PaliGemma 2 Mix, a set of fine-tuned vision language models trained on multiple tasks including optical character recognition, image captioning, and visual question answering. The model family includes three sizes (3B, 10B, 28B parameters) with three resolution options (224x224, 448x448, 896x896 pixels). Users can now apply these models to document understanding, text recognition, object detection, and image segmentation without requiring task-specific prefixes in most cases.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.