TLDRocket
Sign in

Google launches private multimodal search powered by ‘EmbeddingGemma 2’

Google ● Covered by 6 sources

Google’s new EmbeddingGemma 2 can search text, images, audio, and video on-device. It’s built for private offline retrieval, not cloud-heavy AI plumbing.

Based on reporting by Google — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google has launched EmbeddingGemma 2, a new multimodal embedding model that pushes search and retrieval beyond text. The model maps text, images, audio, video, and code into one shared embedding space, so a voice memo can surface a specific clip or a text query can search through hours of recordings without sending data off the device.

The company is pitching it as an on-device tool first. EmbeddingGemma 2 has 740 million parameters, is built on the Gemma 4 architecture, and is released under the Apache 2.0 license. Google says that size makes it a good fit for local inference, where privacy and latency matter more than brute-force scale.

It also arrives with a more modular setup than the usual one-model-does-everything story. Text-only workloads can use as little as 270 million parameters, with optional vision and audio encoders for full multimodal support. Google says the model can run with output vectors trimmed from 768 dimensions down to 512, 256, or 128, which can cut storage use by up to 6x in local vector databases.

The performance pitch is just as direct. Google says EmbeddingGemma 2 leads sub-1B multimodal embedders on benchmarks including MTEB Code and MAEB, and that it improves code scores by 9.92 points over EmbeddingGemma, from 68.76 to 78.68. On a Google Pixel 11 Pro, with quantization, it can use about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model.

There’s also a bigger context window here: 8K tokens, four times more than EmbeddingGemma 1. Google says that lets the model handle up to 5.5 minutes of audio, 29 images, or 58 video frames directly on local hardware, and pair cleanly with Gemma 4 for on-device RAG pipelines. The model weights are on Hugging Face and Kaggle now, with Gemini Enterprise Agent Platform Model Garden support coming later.

My take — AI-written commentary, not fact-checked reporting

This is the kind of AI release that actually deserves the word useful. Big models still get the headlines, but private multimodal search on a phone is the sort of thing people will remember because it works without a data-exfiltration side quest. The industry keeps selling magic; Google is selling plumbing, which is usually where the real money and the real adoption live.

Read more about this at: Google

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.