
Image to help understand the article
Google DeepMind introduces a multimodal open embedding model
Google DeepMind has released EmbeddingGemma 2, an open model designed to convert text, images, video and audio into a shared embedding format that artificial intelligence systems can use for search and retrieval.
The model was announced on Oct. 6, 2026, and is available through Hugging Face and Kaggle under the Apache 2.0 license, which allows commercial use. Google said the model was designed from the beginning to support multiple types of data in a single embedding system.
The model is identified as “google/embeddinggemma-2” in its repositories and is also expected to become available through Google’s Gemini Enterprise Agent Platform Model Garden. Developers can run it using tools including vLLM, llama.cpp and Ollama.
A smaller AI model designed for local devices
EmbeddingGemma 2 contains 740 million parameters in its full multimodal version. Its architecture includes a 270 million-parameter text backbone, a 170 million-parameter vision encoder and a 300 million-parameter audio encoder. Google also provides a text-only 270 million-parameter version.
The model supports an 8K-token context window and produces 768-dimensional output vectors. Using a training method known as Matryoshka representation learning, developers can reduce those vectors to 512, 256 or 128 dimensions depending on their needs.
Google also shared memory figures after applying quantization. On a Google Pixel 11 Pro, the text-only version uses about 191MB of memory, while the full multimodal version uses about 567MB.
The company highlighted a local-first Mac meeting notes application that uses the 740 million-parameter version of EmbeddingGemma 2. The app can organize meeting records from multiple applications and in-person conversations without requiring an internet connection.
Benchmark results and limitations
Google’s published comparisons focused on improvements over the previous EmbeddingGemma 1 model. In the MTEB Code benchmark, EmbeddingGemma 2 scored 78.68 on NDCG@10, compared with 68.76 for the earlier version.
For multilingual text evaluation, however, the MTEB Mean(Task) score showed only a small change, reaching 61.36 compared with 61.15 for EmbeddingGemma 1.
Google also released scores for image, video and audio-related evaluations, including 64.64 on MIEB (lite) for image tasks, 50.67 on MMEB v2 Video Hit@1 for video, and 69.54 on MSEB Retrieval MRR@10 for audio search. The company did not provide comparisons with other companies’ models for those categories.
Google said the model achieved leading results among multimodal embedding models with fewer than 1 billion parameters. However, the comparisons published by Google and its model documentation primarily compare EmbeddingGemma 2 with its own previous version rather than a broader ranking against competing models.
The company also noted several limitations. Reducing embeddings to 128 dimensions can lower quality, particularly for image, video and audio searches. Google recommends testing performance with real deployment data before using that setting. The company said 256 dimensions generally maintain quality with little loss.
Other restrictions include an audio processing limit of about 5.5 minutes per input and possible reductions in text embedding quality if recommended task prefixes are omitted. Google also noted that performance may vary across more than 100 supported languages and warned that models trained on large real-world datasets may reflect social and cultural biases.
Comments
Post a Comment