Google expands EmbeddingGemma beyond text to images, audio and video
SiliconANGLE Duncan Riley ● Covered by 5 sources
Google released EmbeddingGemma 2, expanding its open embedding model from text to shared multimodal embeddings for images, audio, and video. The model is 740 million parameters and is designed to run on a smartphone. It enables on-device retrieval like matching a voice memo to a video moment without sending data off the phone, with options to shrink embedding vectors down to 128 numbers.
Why it matters
Google LLC today released EmbeddingGemma 2, an open multimodal embedding model small enough to run on a smartphone. The release takes the EmbeddingGemma line beyond text, which was all the first version handled when Google introduced it in September 2025. Images, audio and video now share one embedding space with text. An app built on […] The post Google expands EmbeddingGemma beyond text to images, audio and video appeared first on SiliconANGLE.