Cohere Releases Embed 5 With Compatible Models for Indexing and Live Search
Pro and Fast share an embedding space, so developers can use different models for stored documents and incoming queries. Cohere recommends that split, but its release announcement does not quantify the performance tradeoff.
Cohere’s Embed 5 is a two-model embedding family aimed at enterprise retrieval across complex document collections, including PDFs, code and multilingual material. Its Pro and Fast variants share a vector space, allowing teams to build indexes with Pro and serve searches with Fast instead of sacrificing indexing quality for query latency. That separation lets retrieval systems tune offline preparation and live traffic independently; Cohere describes gains over Embed 4 but provides no benchmark figures in the supplied material to quantify them.
01
Embed 5 accepts text, images or mixed inputs, and Cohere lists support for more than 100 languages.
02
The models support a 128,000-token context window and selectable embedding dimensions from 256 to 2,048, with float, int8 and binary outputs.
03
Developers can access Embed 5 through Cohere’s Embed API, Microsoft Foundry and Amazon SageMaker; Model Vault is listed for single-tenant deployment.
Developers using Cohere’s new search models no longer have to choose the same variant for preparing a document collection and handling incoming queries. Released on September 30, 2026, Embed 5 comes in Pro and Fast versions that share an embedding space. Cohere recommends indexing documents with Pro and querying them with Fast—a pairing designed to put retrieval quality into preparation and speed into live search.
The release starts with difficult documents
Cohere positions Embed 5 as an upgrade for complex enterprise data, rather than only straightforward text. The company says it delivers major gains over Embed 4 on visually rich documents, financial filings, parsed PDFs, code and multilingual retrieval.
The response is a two-model family with different jobs. The quality-focused variant, embed-v5.0-pro, is optimized for offline indexing and searches where retrieval quality is critical. The speed-focused variant, embed-v5.0-fast, targets low latency and high throughput. Cohere identifies interactive search, repeated searches during an agent’s work and high-volume query traffic as uses for Fast.
Compatibility connects those roles. Both models produce embeddings—the vector representations used for retrieval—in the same space. Cohere says a collection indexed with either model can be queried with the other. Its recommended setup assigns Pro to the document collection and Fast to incoming search requests, rather than treating the two variants as mutually exclusive choices.
Text and images can stay together
The supported input and storage envelope
Language coverage: Cohere lists support for more than 100 languages, alongside the family’s text and image inputs.
Input length: the context window is 128,000 tokens, the units used to measure model input.
Vector size: developers can choose 256, 512, 768, 1,024, 1,536 or 2,048 dimensions—the number of components in each embedding.
Cohere calls the selectable sizes Matryoshka embeddings and lists float, int8 and binary output types.
Available now; the size of the tradeoff remains open
Embed 5 is available through Cohere’s Embed API, Microsoft Foundry and Amazon SageMaker. Cohere also lists Model Vault for single-tenant deployment.
Reader comments
Newest comments first. Replies stay oldest first.