On-device AI & multimodal retrieval

EmbeddingGemma 2 brings multimodal search on-device: when local retrieval wins

•Make Better Editorial

Google’s 740M open-weight embedding model unifies text, code, images, video and audio and can run fully on-device. Here’s where local retrieval is useful—and where cloud embeddings still win.

Google DeepMind launched EmbeddingGemma 2 on October 6 as a compact open-weight model for multimodal retrieval. It maps text, code, images, video and audio into the same vector space, while its modular design lets developers load only the encoders they need. The practical shift is that semantic search, RAG and intent routing can run locally on consumer hardware instead of sending every query and document to a cloud embedding API.

What launched

EmbeddingGemma 2 at a glance

740M parameters
Full model
Text, vision and audio
270M parameters
Text/code only
Smallest modular configuration
768d
Vector size
Truncatable to 512d, 256d or 128d
~567MB
Full active RAM
Google measurement on Pixel 11 Pro

The model uses an 8,192-token context window and a shared 768-dimensional space across modalities. That means a text query can retrieve relevant images, video moments or audio without first converting every asset into captions or transcripts. Google is also integrating it with MediaPipe retrieval and decision tasks, with ML Kit support planned in the coming weeks.

Why local retrieval matters

Make Better analysis

The strongest use case is not simply avoiding an API bill. Local embeddings change the privacy, latency and offline constraints of a product. A meeting assistant can index private audio locally; a mobile app can search photos without uploading a library; an agent can search a local codebase before deciding what context to send to a larger model. The embedding layer can stay on-device even when the final reasoning step still uses a cloud model.

When local beats cloud—and when it does not

Choose the retrieval layer by constraint

Use local EmbeddingGemma 2 when…Prefer cloud embeddings when…
Sensitive data should remain on the deviceCentralized retrieval across many users or systems is required
Low latency or offline operation mattersServer-side scaling and managed infrastructure matter more than device privacy
The corpus is naturally local, such as files, media or a codebaseThe corpus is huge or changes continuously across distributed sources
You can validate quality on a bounded taskYou need the strongest available retrieval quality and can accept network dependency

Vector storage is another lever

EmbeddingGemma 2 supports Matryoshka truncation from 768 dimensions down to 128. Google says 256-dimensional vectors retain most text and code quality and about 95% of full-dimensional quality for image, video and speech retrieval while cutting vector storage by roughly three times. At 128 dimensions storage falls about six times, but Google reports a larger quality drop for multimodal retrieval. Those are model-level results, not a guarantee for a particular corpus.

Implementation caveat

Truncated vectors must be re-normalized, and queries and indexed documents must use the same dimension. More importantly, benchmark quality does not replace workload evaluation: retrieval accuracy, indexing time, battery use, device thermals and memory pressure can determine whether local inference is actually the better product choice.

A practical evaluation checklist

  1. Build a representative query set with known relevant results before comparing configurations.
  2. Test the smallest modality configuration that covers the product instead of loading all 740M parameters by default.
  3. Measure recall at 768d and 256d before trading vector quality for storage.
  4. Track end-to-end latency, memory and battery impact on the weakest device you intend to support.
  5. Keep sensitive retrieval local where useful, and send only the minimum retrieved context to a cloud reasoning model if one is still needed.
  6. Compare local and cloud options by cost and quality per successful retrieval, not by model size alone.
Bottom line

EmbeddingGemma 2 makes multimodal retrieval small enough to become an on-device architectural choice rather than a cloud-only feature. Its biggest value is the option to keep search, indexing and routing close to private data. Use local retrieval when privacy, latency or offline operation materially matter; keep cloud embeddings where centralized scale and maximum managed quality are the stronger constraint.

Sources & useful resources