Google recently introduced the EmbeddingGemma 2 model, developed based on the Gemma 4 architecture and released under the Apache 2.0 license. EmbeddingGemma 2 features a total of 740 million parameters and leverages the same underlying technology as the Gemini embedding model, enabling the unified mapping of various data types—including text, code, images, videos, and audio—into a single embedding space. Among multimodal embedding models with sub-1B parameters, EmbeddingGemma 2 stands out for its leading performance. It adopts a modular design, supports dynamic truncation of output vector dimensions, and maintains low memory usage during end-side runtime. Additionally, the model is equipped with an 8K Token context window, significantly enhanced code performance compared to its predecessor, and exceptional performance in multimodal tasks. EmbeddingGemma 2 prioritizes data privacy protection, reduces pipeline latency, and supports the construction of fully offline cross-modal search and retrieval systems. When paired with generative models like its counterpart Gemma 4, it enables low-memory-footprint end-side multimodal RAG pipelines. The model supports multiple platforms and a wide range of commonly used development tools. Developers can download the model weights on Hugging Face and Kaggle, with subsequent availability planned for the Gemini Enterprise Agent Platform Model Garden.
