Cohere Embed 5 — Shared embedding space across model tiers — What does it mean?
The news. On September 30, 2026, Cohere released Embed 5 in two tiers: Pro, built for maximum retrieval quality, at $0.12 per million text tokens, and Fast, a lighter model for latency-sensitive search, at $0.08. Both accept text, images, or fused text-and-image inputs, cover 100+ languages and a 128K-token context, and output 256 to 2,048 dimensions as float, int8 or binary vectors. The structural feature this explainer covers: Pro and Fast share one embedding space, and Cohere recommends the pattern "index with Pro, query with Fast." Read the release →
Picture a library where a careful senior cataloguer reads every book and writes a shelf number on its spine. Later, a fast intern works the front desk: for each request they write a shelf number on a slip and walk to that shelf. The arrangement only works because both of them use the same numbering system. If the intern had learned a different system, slip "B-07" would send them to the wrong shelf, and the library would have to keep the old intern or re-shelve every book in the new system.
Ordinary embedding models are in exactly that situation. An embedding model turns text or an image into a vector, a point in a space where nearby points mean similar content. But the axes of that space are learned, not agreed on in advance. Two separately trained models produce two unrelated coordinate systems, so the cosine similarity between a vector from model A and a vector from model B carries no meaning, even when both vectors have 1,024 numbers. The same dimension count is not the same space, in the same way that two maps of the same size can use different grids.
Embed 5 changes this for its own two tiers. Pro and Fast are trained so that their vectors land in one shared space, so a document vector written by Pro can be compared directly with a query vector written by Fast. Cohere has not published how the two models were aligned. A common way to get this property is to train the smaller encoder to place each input where the larger encoder already places it, but treat that as general background, not as Cohere's documented recipe. The two vectors also need the same output dimension, because a similarity score is only defined between vectors of equal length.
Why does this matter? Because in a retrieve-then-generate pipeline the two sides of the index have opposite cost profiles. Documents are embedded once, in an offline batch job, where quality matters most and the embedding cost is paid once per document. Queries are embedded on every search, on the live request path, where each millisecond adds to the user's wait. Cohere points out that agents can issue dozens of searches per task, which multiplies that per-query cost inside an agent's cost profile. A shared space lets you choose the model for each side separately: the careful cataloguer for the books, the fast intern at the desk. It also makes a later change on the query side cheap, because moving queries from Pro to Fast, or back, needs no re-index of the corpus.
| Corpus indexed with | Query embedded with | Mean retrieval quality (all-Pro = 100) | Source |
|---|---|---|---|
| Pro | Pro | 100 (baseline) | Cohere |
| Pro | Fast | 98.4 | Cohere |
| Fast | Pro | 97.3 | Cohere |
| Fast | Fast | 96.6 | Cohere |
The table is Cohere's mean nDCG@10 across 40 development datasets, normalized so that Pro on both sides scores 100. Two things stand out. Every mixed pairing stays within about 3 points of the all-Pro baseline, and Cohere reports no dataset with a major failure. And spending the quality on the index side (98.4) beats spending it on the query side (97.3), which matches the library picture: a well-catalogued library forgives a hurried slip better than a careful slip rescues a badly catalogued library.
Worked example. Hold three things fixed: Cohere's normalized quality scores from the table, Cohere's list prices, and a query traffic of 5 billion text tokens a month (illustrative). An all-Fast system scores 96.6 and an all-Pro system scores 100, so the gap between the tiers is 3.4 points. Index with Pro and query with Fast and the system scores 98.4, which keeps 1.8 of those 3.4 points, about 53% of the quality gap. On the query side, 5 billion tokens cost $600 a month at Pro's $0.12 per million and $400 at Fast's $0.08, so the split saves $200 a month, one third of the query bill. Compared with all-Pro, the index-side bill does not change, because both systems use the same Pro index. The result: about half the quality gap closed, for a third off every query.
The shared space has clear limits. It is a property of this model family, not of embeddings in general. A Pro vector still cannot be compared with a vector from another vendor's model, and nothing in the release says Embed 5 vectors are compatible with Embed 4's, so plan for a re-index when moving across generations unless Cohere documents otherwise. The cross-model numbers are averages over Cohere's own development datasets, so measure a sample of your own queries both ways before you switch. And the shared space does not change the usual recall-versus-speed trade of the ANN index; it only separates the model that writes the document vectors from the model that writes the query vectors.
Goes deeper in: AI Agents → Retrieval & RAG → Embeddings as Coordinates
Related explainers
- The Undetected Damage of Quantization on Retrieval — Top-1 score-gap certificate — the other way to cut retrieval cost: shrink the vectors and check which answers silently change