
EmbeddingGemma 2 needs an indexing contract
In short: EmbeddingGemma 2 unifies multimodal retrieval, but production quality still depends on a versioned contract for what the index actually saw.
Google DeepMind released EmbeddingGemma 2 on 6 October 2026: a 740M-parameter model that maps text, code, images, video and audio into one 768-dimensional space. The pitch is simple — one vector space for cross-modal retrieval on a phone.
The useful answer is slightly less tidy. An EmbeddingGemma 2 indexing contract still has to declare what entered the index, which encoders ran, how video was sampled and which vector dimension reached storage. A shared space removes a chain of models; it does not remove the choices that decide recall.
EmbeddingGemma 2 replaces a chain, not the acceptance test
The 54-second source moves from separate modality pipelines at roughly 11 seconds to a shared cat example around 30 seconds: text, an image, video and audio land near one another. Google's developer guide shows the practical advantage. The same checkpoint can load text and code at 270M parameters, add vision at 440M, add audio instead at 570M, or load all modalities at 740M.
The storage choice is modular too. Google's guide says one million BF16 vectors need roughly 1.5 GB at 768 dimensions, versus 250 MB at 128 dimensions. The trade is not free: its reported multimodal retrieval quality falls from 59.01 on MMEB v2 at 768 dimensions to 45.65 at 128 dimensions.
| Configuration | Footprint or storage | Reported trade-off |
|---|---|---|
| Text-only encoders | 270M parameters | No image, video or audio path |
| Full multimodal encoders | 740M parameters | Text, image, video and audio |
| 768d vectors | 1.5 GB per million BF16 vectors | Full reported MMEB v2 score: 59.01 |
| 128d vectors | 250 MB per million BF16 vectors | Reported MMEB v2 score: 45.65 |
Google-reported configuration figures; I did not reproduce them locally.
My deployment judgement is to version the index recipe
I have not deployed EmbeddingGemma 2 at Trigent or for a client. My editorial judgement is that teams should version the index recipe beside the vectors: source eligibility, content hash, enabled encoders, frame rate, task prefix, vector dimension and model revision.
That record is what lets you explain a miss. Google's model card says video defaults to one frame per second and all modalities share an 8,192-token window. A moment between sampled frames is not a ranking failure. A paragraph excluded by the corpus rule is not an embedding failure. This advances my earlier rule to prove the missing item was indexed: for multimodal search, prove the relevant moment was sampled too.
The acceptance set should then follow the deployment corpus. Compare dimensions and encoder combinations on named slices: short clips, motion-heavy clips, noisy audio, internal acronyms, Indian languages and documents whose answer sits near the context boundary. Google's benchmark table is evidence that the model is worth testing, not evidence that your sharpest slice passed.
Where EmbeddingGemma 2 helps and where it breaks
The launch is strongest for local search where privacy, intermittent connectivity or interactive latency matter. Google reports about 191 MB active RAM for text-only weights and 567 MB for the full model on a Pixel 11 Pro, while its Edge Gallery stores media vectors locally in SQLite.
Those are vendor-reported launch measurements on named hardware. I did not install the weights, reproduce the benchmarks or test thermal behaviour, multilingual recall or long-running indexing. “On-device” describes execution location; it does not certify privacy, fairness, corpus permissions or retrieval quality.
What's in it for you
- You can replace several modality-specific retrieval stages without losing visibility into what each index build contained.
- You can trade vector storage against measured slice recall instead of accepting the 6x compression headline on faith.
- You get a clean debugging order: eligibility, sampling, representation, then ranking.
A unified embedding space is useful; the production artifact is the versioned recipe that decides what the space was allowed to remember.
Sources
- Sundar Pichai — EmbeddingGemma 2 launch video, 6 October 2026
- Google Developers Blog — Bring multimodal semantic search to the edge with EmbeddingGemma 2, 6 October 2026
- Google AI for Developers — EmbeddingGemma 2 model card, updated 6 October 2026


