Your phone’s vector index might be bigger than the AI model running it
The New Stack Amanda Caswell ● Covered by 3 sources
Google’s new EmbeddingGemma 2 can search text, images, audio and video on a phone. The twist: the index can eat more memory than the model itself.
Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google rolled out EmbeddingGemma 2 on Tuesday, and the pitch is simple: one open model for searching text, code, images, video and audio on a device. In Google’s testing on a Pixel 11 Pro, the full setup used about 567MB of active RAM with quantization. The weights are out under Apache 2.0, and on-device use is already available through LiteRT and MediaPipe Tasks.
The model sits on Gemma 4 and maps all five input types into the same 768-dimensional vector space. That matters because images do not need captions and audio does not need transcripts before they can be searched alongside text. Google says an Android ML Kit integration with NPU acceleration is coming in the next few weeks.
The modular setup is the practical part. Developers only load the encoders they need, so text and code run on a 270-million-parameter base, while adding vision takes it to 440 million. Audio alone brings the total to 570 million, and loading both pushes the model to the full 740 million parameters. Google measured the text-and-code setup at about 191MB of active RAM on the same phone.
Google also stretched the context window from 2,048 to 8,192 tokens. The company says that can cover as much as 5.5 minutes of audio, 29 images or 58 video frames in one input. In its Video Moments Finder demo, frames and audio chunks are indexed locally and then searched with plain text, with no captions or transcripts generated. Instant Media Search does the same for photos and videos on a phone, storing embeddings in SQLite and updating results as the user types.
The memory story gets sharper once the index enters the picture. Google says a million 768-dimensional bfloat16 vectors take about 1.5GB, so the index can rival the model for space. Matryoshka Representation Learning lets developers shrink those embeddings to 512, 256 or 128 dimensions without retraining. At 256 dimensions, that same million-vector index falls to about 500MB, with Google saying quality stays mostly intact for text and code and around 95% for image, video and speech retrieval.
Google also showed the model doing code search and classification on-device. It reports an MTEB Code score of 78.68, up from 68.76 for the original EmbeddingGemma. In one chess demo, MediaPipe Decision checked 500 options per turn in under 100 milliseconds. And for teams trying to build local RAG pipelines, the point is obvious: sometimes the expensive part is not the model. It is the memory hanging off the side of it.
My take — AI-written commentary, not fact-checked reporting
Open models keep winning the boring battles that actually matter, like memory and latency, while the hype crowd is still yelling about frontier scores. A phone that can search its own photos, audio and code without shipping everything to the cloud is the kind of upgrade people notice after five minutes and forget to tweet about. That usually means it’s the real thing.
Read more about this at: The New Stack