Google has released EmbeddingGemma 2, an open-weights multimodal embedding model published under the Apache 2.0 licence. Embedding models turn text, and here also audio, into numerical vectors, which is the step that makes search and retrieval work at all. This one is built to do it on a device rather than in a datacentre.
The model has 740 million parameters, with a modular design that drops to as little as 270 million for text-only work. Google reports roughly 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model, measured on a Pixel 11 Pro with quantisation applied. The context window is 8,000 tokens, which Google says is four times that of the first EmbeddingGemma.
The improvement Google puts a number on is code retrieval: a 9.92 point gain on MTEB Code, from 68.76 to 78.68. It also cites results on MAEB, an audio embedding benchmark. Both are the vendor’s own measurements.
The licence matters more here than the benchmark. A permissive Apache 2.0 release on a model that fits in a few hundred megabytes of working memory is the combination that lets retrieval run on hardware which never contacts a server, which is the deployment case that regulated and offline users keep asking for. Weights are on Hugging Face and Kaggle.
Source: Google, EmbeddingGemma 2: an open, lightweight multimodal embedding model.
