Twenty million downloads is a number most open-source embedding projects would kill for. That’s what Google’s original EmbeddingGemma pulled in after launch, which tells you there’s real demand for on-device retrieval that doesn’t phone home to a server. Now Google has announced EmbeddingGemma 2, and the scope has expanded significantly beyond text.
The new model has 740 million parameters, runs on the Gemma 4 architecture, and ships under an Apache 2.0 license, meaning commercial use is fair game without licensing negotiations. The core pitch is a single model that embeds text, code, images, audio, and video into one unified vector space. So a voice memo can retrieve a video clip. A text query can surface a moment buried inside an hours-long audio recording. All of it processed locally, without a cloud API in the loop.
What actually changed from version 1
The most meaningful upgrade is multimodality. EmbeddingGemma 1 was text-only. Version 2 adds optional vision (170M parameters) and audio (300M parameters) encoders that bolt onto a 270M text-only base. Developers who only need text can skip the extra weight entirely. That modularity matters on constrained hardware.
Context window also grew from 2K to 8K tokens, a 4x increase. In practice, that means the model can process up to 5.5 minutes of audio, 29 images, or 58 video frames in a single pass. And on code retrieval, Google reports a 9.92-point improvement on the MTEB Code benchmark, moving from 68.76 to 78.68. For anyone building semantic code search or indexing a local codebase, that’s a material jump.
Why on-device embedding is the right comparison axis
The obvious competitors here are API-based embedding models: OpenAI’s text-embedding-3 series, Cohere’s Embed v3, and Voyage AI’s multimodal offerings. Those are strong, but they require data to leave the device. For enterprise apps handling sensitive documents, medical records, or private media libraries, that’s a non-starter. EmbeddingGemma 2 fits a different use case by design.
Among sub-1B on-device multimodal models, Google claims best-in-class scores across MTEB and MAEB benchmarks, with performance that matches or beats some models more than twice its size. That quality-per-parameter efficiency is the real story. It runs on a Pixel 11 Pro using roughly 191MB of active RAM for text-only and about 567MB for the full multimodal build.
Storage and deployment details worth knowing
EmbeddingGemma 2 uses Matryoshka Representation Learning, which lets developers truncate output vectors from 768 dimensions down to 512, 256, or 128. That’s up to a 6x reduction in local vector database storage, which adds up fast in mobile or edge deployments.
For deployment, the model is available through:
- Hugging Face and Kaggle for model weights
- Google AI Edge MediaPipe for cross-platform on-device apps
- LiteRT for custom model integration and browser via WebGPU
- Standard inference tools including vLLM, Ollama, llama.cpp, and sentence-transformers
- Qdrant for vector storage
When paired with Gemma 4 for generation, both models share a text tokenizer and audio encoder, so running them together in a RAG pipeline costs less total memory than running two independent models. That’s a practical engineering advantage, not just a marketing claim. So for developers building private, offline-capable search over mixed media, this is probably the most complete on-device option available right now.



