TLDRocket
Sign in

Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4

MarkTechPost Asif Razzaq ● Covered by 6 sources

Google DeepMind dropped EmbeddingGemma 2, a 740M open model for text, code, images, video and audio. It can run on-device, so search and retrieval don’t have to leave the device.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google DeepMind has shipped EmbeddingGemma 2, an open multimodal embedding model that packs text, code, images, video and audio into a single 768-dimensional space. It comes with 740 million parameters, an 8,192-token context window, and an Apache 2.0 license. The pitch is simple: keep retrieval local, fast and private.

This is not just a bigger version of a text embedder with a few extras bolted on. The model is modular. The text and code backbone uses 270 million parameters, the vision encoder adds 170 million, and the audio encoder adds another 300 million. Developers can load only the pieces they need, so the footprint ranges from 270 million parameters for text alone to the full 740 million for everything.

That shared vector space is the real trick. A text query can match an image. A voice memo can pull up a video clip. Mixed inputs like a product page with text, pictures and a demo video still end up in one embedding, which makes the model useful for on-device search, classification and privacy-first RAG. Google says the weights are already live on Hugging Face and Kaggle, with Ollama, llama.cpp GGUF and LiteRT builds available now.

On the benchmark side, Google’s research team says EmbeddingGemma 2 leads among sub-1 billion parameter multimodal embedders on MTEB Code and MAEB. The model card shows code retrieval jumping from 68.76 on EmbeddingGemma 1 to 78.68 here. Multilingual text quality stays basically flat at 61.36 on MTEB multilingual v2, while image, audio and overall multimodal scores are also listed in the model card.

The practical angle matters just as much as the scores. With quantization on a Pixel 11 Pro, Google reports about 191MB of active RAM for text-only use and about 567MB for the full multimodal model. Matryoshka Representation Learning lets developers truncate vectors to 512, 256 or 128 dimensions, which can cut storage sharply. At 256 dimensions, multilingual text barely slips. At 128 dimensions, Google says the model is mainly for text-only workloads. And yes, Android support with NPU acceleration is said to be coming within weeks.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI that deserves more attention than another chat demo: smaller open models that actually fit on devices. Big cloud models make the headlines, but local embeddings are what turn privacy from a slogan into a setting. The industry keeps selling magic; meanwhile, the useful stuff is getting quieter and cheaper.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.