LLM·

Google's EmbeddingGemma 2 — Matryoshka embedding truncation — What does it mean?

The news. On October 6, 2026, Google released EmbeddingGemma 2, an Apache 2.0 embedding model that maps text, code, images, audio and video into one 768-dimensional space. It has 740M parameters in total: a 270M text model plus optional 170M vision and 300M audio encoders. Google says that with MRL, developers can truncate the output from 768 dimensions to 512, 256 or 128, for up to 6× less storage in local vector databases, and that the quantized text-only weights need about 191 MB of active RAM on a Pixel 11 Pro. Read the release →

Picture a newspaper story written in the inverted pyramid. The lede says who, what and when; each later paragraph adds less important detail. Editors rely on this: when the page is short, they cut from the bottom, and the story still makes sense. MRL trains an embedding model to write its vector the same way, so the first numbers hold the core meaning and each later block adds finer detail. An ordinary embedding is a story whose key facts are spread through every paragraph. Nothing in its training says that dimension 1 matters more than dimension 700, so what each dimension holds is spread across the whole list, and cutting it at 256 can remove facts the ranking needed.

The training change is small. A normal embedding model computes its training loss on the full vector only. MRL computes the same loss several times, once on each of a chosen set of nested prefixes of the vector, and adds the results, so the model is penalized whenever its first 128 numbers, or its first 256, cannot do the job alone (the MRL paper adds a weight to each term). The paper reports that this adds no cost at inference: the model runs once and emits one vector, and a shorter embedding is simply the first m numbers of it. The paper also reports that these prefixes are at least as accurate as separately trained low-dimensional representations. Google's release does not say which nested sizes it trained on; the model card lists 768, 512, 256 and 128 as the supported cut points.

The cut has two rules that are easy to miss, and both come from the model card. First, a sliced vector must be L2-normalized again before it is used for cosine similarity. The full vector has length 1, but its first 128 numbers alone have a length below 1, and the shortfall differs from one vector to the next. A store that scores by dot product on the assumption of unit length then ranks documents partly by that leftover length instead of by meaning. The card warns that skipping this step produces plausible-looking scores rather than an error. Second, the query and the documents must use the same cut: a 768-number query cannot be compared with a 128-number corpus. In sentence-transformers both rules are one call, model.encode(text, truncate_dim=256, normalize_embeddings=True).

Output dimensionStorage vs 768dMTEB (eng, v2)MTEB (code, v1)MMEB (v2) overallSource
768d (full)1×68.4678.6859.01model card
512d1/1.568.4177.2458.38model card
256d1/367.7876.1856.24model card
128d1/665.6871.4145.65model card

Here is what the cut buys on a concrete index (our arithmetic on the model card's numbers). Hold three things fixed: 10 million documents, one vector each, stored as float32 (4 bytes per number). At 768 dimensions each vector takes 768 × 4 = 3,072 bytes, so the raw vectors, before any index overhead, take about 30.7 GB. At 256 dimensions each vector takes 1,024 bytes, so the same index is about 10.2 GB, one third of the size, and the English MTEB score drops from 68.46 to 67.78, a loss of 0.68 points. Cut again to 128 dimensions and the index is about 5.1 GB, the 6× saving Google quotes; English text still scores 65.68, but the multimodal MMEB score falls from 59.01 to 45.65, a loss of 13.36 points. The same knob that is nearly free for text is expensive for images, video and audio, which is why the model card calls quality close to lossless down to 256 dimensions and recommends 128 mainly for text-only workloads.

Back at the newspaper, the editor does not have to pick one length for every page. Because each supported cut point gives a usable vector, a system can use a short cut where speed matters and the full vector where precision matters, from the same model. A phone app can keep a small local index within its memory budget. A large search system can shortlist candidates with the first 128 numbers and re-rank only the shortlist with all 768; the MRL paper calls this adaptive retrieval and reports up to 14× real-world speed-ups for large-scale retrieval on ImageNet-1K and 4K. That is the paper's result, not an EmbeddingGemma 2 measurement. This is the same storage-versus-recall trade that approximate nearest-neighbour search makes, applied to the size of each vector instead of to the search algorithm.

Goes deeper in: LLM Internals → Embeddings → What Do Dimensions Mean?

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based