Most embedding models still think in text. Google DeepMind’s new open release is built to put text, images, audio and video into one vector space—and to run small enough for phones.
On 6 October 2026, Google’s developer blog announced EmbeddingGemma 2: a 740-million-parameter embedding model under Apache 2.0, built on the Gemma 4 architecture, and described as natively mapping combinations of text, images, audio and video into a unified embedding space. [1]
What Google says the model is
According to the company blog (Research Engineers Sahil Dua and Henrique Schechter Vera, Google DeepMind), EmbeddingGemma 2 is:
- Released under a commercially permissive Apache 2.0 licence.
- Modular: as little as about 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders for full multimodal support.
- Storage-efficient via Matryoshka Representation Learning (MRL), truncating output vectors from 768 dimensions down to 512, 256 or 128—claimed as up to 6× storage reduction for local vector databases.
- Optimised for on-device use: with quantisation, Google reports about 191MB active RAM for text-only weights and about 567MB for the full multimodal model on a Google Pixel 11 Pro.
- Extended context: an 8K token context window (stated as 4× EmbeddingGemma 1), said to allow up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations on local hardware. [1]
The same post says the model is built from the same technology as Gemini Embedding models and shares Gemma 4’s text tokenizer and audio encoder, which Google argues lowers combined memory when paired with Gemma 4 for on-device RAG. [1]
Benchmarks—company-reported
Google reports a 9.92-point improvement on MTEB Code, from 68.76 to 78.68, while saying multilingual text performance matches EmbeddingGemma. It also claims leading scores among sub-1B multimodal embedders on benchmarks such as MTEB Code and MAEB (Massive Audio Embedding Benchmark), and that the model matches or outperforms many larger models across text, vision and audio tasks. [1]
Those figures are company-reported. This draft does not re-run MTEB, MAEB or on-device memory measurements.
How Google wants developers to use it
The blog points to on-device semantic search, video-moment finding, multimodal RAG with Gemma 4, and MediaPipe decision/routing tasks. Weights are said to be on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability “coming soon,” plus tooling mentions spanning LiteRT, MediaPipe, transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, LM Studio and Qdrant. [1]
Google also cites more than 20 million downloads of the earlier EmbeddingGemma text model as community context—not as a quality proof for version 2. [1]
What we don’t know / What this does not prove
- “Best-in-class for its size” is a company framing; independent third-party leaderboards are not cited in the source used here beyond Google’s own summary.
- Apache 2.0 covers the released weights as described by Google; downstream app compliance (privacy, on-device data handling) still sits with the developer.
- Pixel 11 Pro RAM figures are Google measurements under quantisation; other devices and runtimes may differ.
- Model Garden timing remains “coming soon.”
The Bottom Line
EmbeddingGemma 2 is Google DeepMind’s Apache 2.0, 740M multimodal embedding model for text, images, audio and video, with company-reported code-benchmark gains and on-device memory footprints. Useful if you need local cross-modal retrieval—but treat benchmark and RAM numbers as Google’s until independently checked.
Sources
- Google (Sahil Dua & Henrique Schechter Vera, Google DeepMind), “EmbeddingGemma 2: an open, lightweight multimodal embedding model,” 6 October 2026 — https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/

