Google Releases EmbeddingGemma 2 for Offline Search Across Photos, Audio and Video
The downloadable model lets developers load only the media encoders they need. Google’s memory figures come from a Pixel 11 Pro; managed Android access is still weeks away.
EmbeddingGemma 2 gives developers an open-weight route to multimodal semantic search on-device, with a modular design that can omit vision and audio components when they are not needed. Google DeepMind released the 740-million-parameter model under Apache 2.0, alongside Android and iOS reference apps; weights and developer tools are available now. The launch makes local cross-media retrieval more practical, though Google’s reported memory figures are specific to a Pixel 11 Pro, and managed Android access through ML Kit is still planned for the coming weeks.
01
On a Pixel 11 Pro, Google reports about 191 MB of active RAM for quantized text-only weights and 567 MB for the full model; other devices may differ.
02
Developers can reduce output vectors from 768 to 128 dimensions, cutting local vector storage by up to sixfold, according to Google.
03
Instant Media Search stores embeddings in a local SQLite database, while Video Moments Finder locates relevant timestamps without transcription or captions.
A text query can find an audio recording, and a voice memo can help locate a video clip—all without sending the search to the cloud. Google DeepMind released EmbeddingGemma 2 on October 6, 2026, giving developers an open-weight model designed to connect text, code, images, audio and video on local devices.
The 740-million-parameter model is built on Gemma 4 and released under Apache 2.0, a commercially permissive license. Its weights are available on Hugging Face and Kaggle. Google’s launch announcement positions it as a successor to last year’s text-focused EmbeddingGemma, which the company says passed 20 million downloads.
Search without the transcription detour
EmbeddingGemma 2 converts different kinds of content into numerical representations called embeddings. Those representations share one space, allowing an app to compare a written query with an image or audio clip by similarity. Google says this reduces the latency and memory overhead of chaining separate image-captioning, speech-to-text and text-embedding models.
Two new showcases in Google AI Edge Gallery illustrate the approach. Instant Media Search embeds local media and stores the results in a local SQLite database. It ranks matches against text or example-image queries, updating results as the user types. Google says the embedding calculations and search work without an internet connection.
Video Moments Finder indexes video frames with audio chunks, then highlights timestamps matching a descriptive query. It does not need to transcribe the audio or generate intermediate captions. Google offers both showcases in the Gallery app on Android and iOS, giving developers reference implementations rather than only a downloadable model.
Choose what the device loads
The full parameter count is not a requirement for every workload. The model’s modular design lets developers load just the components their app needs, or activate an audio encoder when voice memos are added. Google describes three building blocks:
Text: a 270-million-parameter component for text-only workloads.
Vision: an optional 170-million-parameter encoder for visual inputs.
Audio: an optional 300-million-parameter encoder, bringing the full model to 740 million parameters.
Google reports about 191 MB of active RAM for quantized text-only weights and 567 MB for the full multimodal model on a Pixel 11 Pro. Quantization compresses the model’s weights. These are Google’s device-specific measurements, not a promise that every phone will deliver the same footprint.
Storage is adjustable too: developers can shorten output vectors from 768 dimensions to 512, 256 or 128. Google’s launch post says this can reduce local vector storage by up to sixfold. The model also has an 8K-token input window, accommodating up to 5.5 minutes of audio, 29 images, 58 video frames, or combinations of media.
Downloads now, managed Android access later
Google’s developer guide describes two integration routes. MediaPipe Tasks handles preprocessing and local retrieval for developers seeking packaged tools. LiteRT provides finer control over processing and hardware acceleration for custom applications. The release also supports familiar serving tools including Transformers, llama.cpp, Ollama and MLX.
Android developers waiting for a managed service will need to wait longer. Google plans ML Kit access in the coming weeks, with automatic model updates and neural processing unit acceleration where available. That planned integration is separate from the weights and development tools available at launch.
Sources
developers.googleblog.comBring multimodal semantic search to the edge with EmbeddingGemma 2- Google Developers Blog
blog.googleEmbeddingGemma 2: an open, lightweight multimodal embedding model
Reader comments
Newest comments first. Replies stay oldest first.