Google releases PaliGemma 2 Mix and SigLIP 2 vision-language models
Model release Provisional 95% confidence first seen
Google released PaliGemma 2 Mix, a family of fine-tuned vision-language models in three sizes (3B, 10B, 28B parameters) supporting multiple visual understanding tasks including OCR, image captioning, and VQA. Simultaneously, Google released SigLIP 2, an improved multilingual vision-language encoder family with up to 1 billion parameters, featuring enhanced capabilities for semantic understanding, localization, and dense feature extraction.