Cohere has released Embed 5, a family of models that convert content into vectors, allowing developers to separate document indexing from search-query processing. Pro focuses on indexing quality, while Fast prioritizes low latency and high throughput for interactive search and agent workflows. Their shared vector space lets developers combine the two without rebuilding the index.
One index, different models for documents and queries
Developers can turn documents into vectors with Pro, then process subsequent queries with Fast against the same index. Compatibility requires matching settings across both paths: the vector dimensions, output representation and similarity configuration must remain the same.
Cohere tested every combination of document and query models across 40 development datasets. According to the company, combinations using different models performed 1.6% and 2.7% worse on average than their respective baselines, in which the same model processed both documents and queries.
Across typical context lengths, Fast achieved an average of 2.4 times Pro’s throughput in Cohere’s measurements. The company reports approximately 377 documents per second for Fast and 160 for Pro; these figures measure document-processing throughput, not the end-to-end latency of a user query.
Text, images and a 128,000-token context
Embed 5 accepts text and image inputs, as well as combinations such as PDF pages containing both text and images. It supports more than 100 languages and has a context window of 128,000 tokens.
Outputs can use 256, 512 or 768 dimensions, with 1,024, 1,536 and 2,048 dimensions also available. The models offer float, int8 and binary representations. Both variants support Matryoshka embeddings, allowing applications to store truncated vectors when smaller storage requirements take priority over maximum retrieval quality.
Vector sizes can be tailored to storage needs
The choice of dimensions and representation substantially changes the size of each vector. A 2,048-dimensional float32 vector occupies 8 KB, while a 1,024-dimensional int8 vector requires 1 KB. A 256-dimensional binary vector takes 32 bytes; these figures exclude index metadata, graph structures, replicas and database overhead.
Cohere recommends 1,024-dimensional int8 vectors as a balance between storage size and retrieval quality, reporting performance close to float representations for both models. The database index must support both storing and comparing vectors in the selected representation.
Results on enterprise documents and finance tasks
On ViDoRe V3, a benchmark focused on visually complex enterprise documents, Cohere reports average scores of 85.8 for Embed 5 Pro and 84.7 for Fast. In the same company comparison, Voyage 4 Large scores 83.7, Gemini Embedding 2 scores 83.2 and OpenAI text-embedding-3-large scores 75.5.
In Cohere’s finance-focused evaluations, Pro scored 80.1 on FinanceBench, 90.0 on FinQA and 85.0 on ViDoRe V3 Finance. The company says Fast ranked second in all three tests. On FinanceBench, Cohere reports that Pro scored 21.4 points above OpenAI text-embedding-3-large.
Compared with Embed 4, Cohere reports its largest improvement in Farsi, at 13 points, with reported gains of 12 points each in Telugu and Hindi.
A new metric evaluates retrieved documents for relevance
Cohere optimized Embed 5 using RCP-nDCG@10, a retrieval-evaluation metric designed to address incomplete relevance labels. It uses a calibrated AI judge to assess every retrieved document against a shared rubric.
In a Cohere study involving 46 human annotators, existing benchmark labels scored 0.65 on AUC, a measure of how well they predicted human relevance judgments. A calibrated judge based on a large language model reached 0.91. Human reviewers considered 28% of the documents labeled irrelevant by the benchmark to be useful.
In pairwise system comparisons, Cohere’s research found that RCP-nDCG@10 matched human reviewers’ choice of system in 77% of comparisons, while the conventional nDCG metric matched their choice in 52%. Crediting relevant documents missing from the reference list accounted for a 19-point improvement in that research.
Availability through the API and deployment platforms
Embed 5 is available through the Cohere Embed API, Microsoft Foundry and Amazon SageMaker. Model Vault is available for single-tenant deployments.
API list pricing is $0.12 per million tokens for Pro and $0.08 for Fast, while hosted and single-tenant deployments may have platform-specific terms.



