Google DeepMind has released EmbeddingGemma 2, built on the Gemma 4 architecture with 740 million parameters, mapping text, code, images, video, and audio into a shared vector space. Distributed under the Apache 2.0 license and optimized for consumer hardware, the model slashes reliance on pricey cloud infrastructure and external API calls.

EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference.

Infrastructure Economics and Deployment Flexibility

As Google DeepMind explained, the model uses a modular design requiring just 270 million parameters for pure text processing, supplemented by optional 170 million parameter visual encoders and 300 million parameter audio encoders.

This provides up to 6x storage reduction for local vector databases and memory usage.

Running on a Google Pixel 11 Pro with quantization, the model consumes roughly 191MB of active RAM for text weights and 567MB for the full multimodal setup. This shifts heavy computational lifting away from expensive cloud clusters straight to local hardware.

Performance and Enterprise Integration

With the context window expanded to 8K tokens—a fourfold increase over the first iteration—the model processes up to 5.5 minutes of audio, 29 images, or 58 video frames locally. On the MTEB Code benchmark, the score jumped by 9.92 points, moving from 68.76 to 78.68, setting the pace for models under one billion parameters. With weights already live on Hugging Face and Kaggle, and an upcoming integration with the Gemini Enterprise Agent Platform Model Garden, it reads as a deliberate push into enterprise infrastructure.

Building on more than 20 million downloads of the first version, developers are turning to solutions like this to build local search tools and privacy-focused RAG pipelines.

Artificial IntelligenceLarge Language ModelsRAG and Vector SearchOn-Device AIGoogle DeepMind