Model release

Google releases EmbeddingGemma 2 for multimodal embeddings on edge devices

The 740M-parameter Apache-2.0 model maps text, code, images, video and audio into one embedding space, with modular encoders, an 8K context window and published on-device deployment paths.

Open Model Weights published7 Oct 2026
Primary sourceGoogle DeepMind
Source published2026-10-06

Google DeepMind has launched EmbeddingGemma 2, extending the earlier text-focused EmbeddingGemma line into a natively multimodal embedding model. Google says the 740M-parameter release is built on the Gemma 4 architecture and maps text, code, images, video and audio into a shared representation space. The company publishes the model under Apache 2.0 and positions it for local semantic search, retrieval and routing on consumer hardware rather than only server-side embedding pipelines.

The architecture is deliberately modular. Google says text-only workloads can use roughly 270M parameters, while optional vision and audio encoders expand the model to full multimodal operation. The release also uses Matryoshka Representation Learning so 768-dimensional embeddings can be truncated to 512, 256 or 128 dimensions. Its stated context window is 8K tokens, with Google mapping that capacity to mixed media such as audio segments, image sets and sampled video frames. The company also publishes Pixel 11 Pro memory figures, but those numbers remain publisher-reported deployment measurements rather than Open Model Weights benchmarks.

For the open-weight ecosystem, the release is notable because one model can support cross-modal retrieval without forcing separate text, image and audio embedding stacks. Google lists weights through Hugging Face and Kaggle and names deployment paths including MediaPipe, LiteRT, Transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio. Open Model Weights should treat those runtime statements as source-derived compatibility signals and verify the exact repository artifacts, license evidence and configuration before any canonical registry record is created or updated.

OMW REGISTRY WATCH

What this changes for the evidence layer

Open-weight multimodal embedding candidate. Canonical model fields should come from the official weight repository and model card; benchmark, RAM and performance figures in this brief remain publisher-reported until independently reproduced.

PRIMARY SOURCE

Google DeepMind

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗