Embeddings on Local Machines: EmbeddingGemma 2
EmbeddingGemma 2 is an open multimodal embedding model designed for on-device use. It maps text, code, images, video, and audio into a shared embedding space, allowing applications to organize, search, and connect different kinds of information locally. Its local design supports privacy-first retrieval and applications that can continue working without an internet connection.
The model is built on the Gemma 4 architecture and released under the commercially permissive Apache 2.0 license. With 740 million parameters, it is intended for inference on consumer hardware. It can process cross-modal queries, such as using a voice memo to locate a video clip or using text to search audio recordings.
EmbeddingGemma 2 has a modular structure. Text-only workloads can use as little as 270 million parameters, while optional vision and audio encoders add 170 million and 300 million parameters respectively. This lets developers select components according to the type of data an application needs. The model also uses Matryoshka Representation Learning, which allows its output vectors to be reduced from 768 dimensions to 512, 256, or 128 dimensions. According to Google, this can reduce storage requirements for local vector databases and memory use by up to six times.
The model is optimized for constrained devices. With quantization, Google reports active RAM requirements of about 191 MB for text-only weights and about 567 MB for the complete multimodal model on a Google Pixel 11 Pro. Its 8,000-token context window is four times larger than that of the first EmbeddingGemma. It can handle up to 5.5 minutes of audio, 29 images, 58 video frames, or combinations of these inputs on local hardware.
Google reports leading results among multimodal embedding models with fewer than one billion parameters on benchmarks including MTEB Code and MAEB. EmbeddingGemma 2 also improves code performance over the first EmbeddingGemma, with its MTEB Code score rising from 68.76 to 78.68. The model is intended for tasks such as local codebase indexing, semantic code search, media retrieval, classification, routing, and on-device retrieval-augmented generation.

Developers can obtain the model weights through Hugging Face and Kaggle. Deployment options include Google AI Edge MediaPipe and LiteRT, while browser use is supported through tools such as transformers.js and WebGPU. The model can also be served with frameworks and applications including Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LM Studio.
Source:
