TLDRocket
Sign in

EmbeddingGemma 2: an open, lightweight multimodal embedding model

Google ● Covered by 3 sources

Google DeepMind launched EmbeddingGemma 2, a small model that turns text, images, audio and video into one shared space. It’s built for phones and other devices, so search and retrieval can happen offline and with less data exposure.

Based on reporting by Google — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google DeepMind has launched EmbeddingGemma 2, a multimodal embedding model built for on-device use. It takes text, code, images, audio and video and maps them into a shared embedding space, so a single model can handle cross-modal search instead of stitching together separate systems.

The company is pitching it as the most capable model in this category for local hardware. EmbeddingGemma 2 has 740 million parameters, is built on the Gemma 4 architecture, and comes under the Apache 2.0 license. DeepMind says it can also be trimmed down for lighter jobs: text-only workloads can use as little as 270M parameters, with optional vision and audio encoders for the full setup.

There’s a clear practical angle here. DeepMind says the model can search through hours of audio from a text query, or pull out a specific video clip from a voice memo, all on-device. It also uses Matryoshka Representation Learning so output vectors can be cut from 768 dimensions to 512, 256 or 128, which the company says can reduce local storage needs by up to 6x.

Performance is part of the pitch too. DeepMind says EmbeddingGemma 2 gets leading scores among sub-1B multimodal embedders on benchmarks including MTEB Code and MAEB, and that it improves code performance by 9.92 points over EmbeddingGemma, from 68.76 to 78.68. The model also has an 8K token context window, which DeepMind says is enough for up to 5.5 minutes of audio, 29 images or 58 video frames on local hardware.

The privacy story is doing a lot of work here, and fairly so. Keeping embeddings on the device means less data moving around, lower latency, and fewer excuses to ship your voice notes to a server because “the cloud” was convenient. The bigger point is simpler: open, local multimodal models are getting good enough that the old centralised default is starting to look lazy.

My take — AI-written commentary, not fact-checked reporting

This is the kind of release that makes the on-device crowd look sensible and the “just send everything to the cloud” crowd look a bit dated. Open models keep winning the unglamorous battles that actually matter: privacy, cost, and control. The hype machine can keep chasing giant models; the useful work is happening in smaller ones that fit in a pocket.

Read more about this at: Google

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.