Google has released EmbeddingGemma 2, an open model that lets an app search across text, code, images, audio and video entirely on the user’s own phone or laptop. Sahil Dua and Henrique Schechter Vera, both research engineers at Google DeepMind, announced it on the Google blog on 6 October 2026. The pitch is that nothing has to be sent to a cloud service for the search to work.
An embedding turns a piece of content into a long list of numbers, so that a computer can judge which things are similar by comparing the lists. Two photos of the same street end up with close numbers, and so do a spoken description of that street and a typed one. That is the machinery behind search by meaning, and behind the retrieval step that feeds relevant documents to a chatbot. EmbeddingGemma 2 places all of its input types in one shared space, so a text query can find an audio clip or a video moment.
The input list in the post varies slightly between paragraphs. The opening summary names text, images, audio and video, while the launch paragraph adds code. Treat all five as supported, and check the model card before depending on any single combination.
Running locally changes the economics and the privacy picture. Google says generating embeddings on the device protects data, cuts latency and works offline. For a company handling medical notes, legal files or personal voice memos, skipping the round trip to a vendor’s server can also remove a per-request bill and a data-processing agreement. The model is released under the Apache 2.0 licence, which allows commercial use.
Size is what makes that plausible. The full model has 740 million parameters, but it is modular: text alone needs about 270 million, with optional vision and audio encoders of 170 million and 300 million. Google says that with quantization, on a Pixel 11 Pro, text-only weights need about 191MB of active memory and the full multimodal model about 567MB. Because it is built on Gemma 4, it reuses that model’s text tokenizer and audio encoder, so running both together costs less memory than running them apart.
Developers can also shrink the stored vectors. Using a technique called Matryoshka Representation Learning, the 768-number lists can be cut to 512, 256 or 128 numbers, which Google says saves up to six times the storage in a local vector database. The context window is 8,000 tokens, four times its predecessor’s, which Google says fits up to 5.5 minutes of sound, or 58 frames of video, or 29 pictures.
The first EmbeddingGemma, text only, passed 20 million downloads, according to Google. On code, the post reports the MTEB Code benchmark score climbing about ten points over that model (9.92, to be exact), ending at 78.68.
The comparison claims are Google’s own. The post calls the model best in its size class among multimodal embedders under 1 billion parameters, citing MTEB Code and a benchmark for audio called MAEB, and says it outperforms a few purpose-built models with over double the parameters. The text names no rival models and gives no rival scores, because the charts carrying the comparison are not reproduced in it. The post also does not say who runs either benchmark, and nothing indicates an outside group has reproduced the numbers.
Weights can be downloaded from Hugging Face or Kaggle, with Google’s cloud model catalogue to follow, and the model runs in tools including Ollama, llama.cpp and sentence-transformers. Developers building private search or on-device retrieval should test it on their own audio and video before committing, since the only quality evidence so far is the maker’s.
Sahil Dua and Henrique Schechter Vera, Google DeepMind, on the Google blog (blog.google), 6 October 2026.