Cohere Embed 5 pairs Pro indexing with Fast search

Cohere Embed 5 pairs Pro indexing with Fast search

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 10. 2026
3 minutes reading · 1 views
Listen to the article
Audio version of the article
Cohere Embed 5 pairs Pro indexing with Fast search

Cohere has released Embed 5, a family of models that convert content into vectors, allowing developers to separate document indexing from search-query processing. Pro focuses on indexing quality, while Fast prioritizes low latency and high throughput for interactive search and agent workflows. Their shared vector space lets developers combine the two without rebuilding the index.

One index, different models for documents and queries

Developers can turn documents into vectors with Pro, then process subsequent queries with Fast against the same index. Compatibility requires matching settings across both paths: the vector dimensions, output representation and similarity configuration must remain the same.

Cohere tested every combination of document and query models across 40 development datasets. According to the company, combinations using different models performed 1.6% and 2.7% worse on average than their respective baselines, in which the same model processed both documents and queries.

Across typical context lengths, Fast achieved an average of 2.4 times Pro’s throughput in Cohere’s measurements. The company reports approximately 377 documents per second for Fast and 160 for Pro; these figures measure document-processing throughput, not the end-to-end latency of a user query.

Text, images and a 128,000-token context

Embed 5 accepts text and image inputs, as well as combinations such as PDF pages containing both text and images. It supports more than 100 languages and has a context window of 128,000 tokens.

Outputs can use 256, 512 or 768 dimensions, with 1,024, 1,536 and 2,048 dimensions also available. The models offer float, int8 and binary representations. Both variants support Matryoshka embeddings, allowing applications to store truncated vectors when smaller storage requirements take priority over maximum retrieval quality.

Vector sizes can be tailored to storage needs

The choice of dimensions and representation substantially changes the size of each vector. A 2,048-dimensional float32 vector occupies 8 KB, while a 1,024-dimensional int8 vector requires 1 KB. A 256-dimensional binary vector takes 32 bytes; these figures exclude index metadata, graph structures, replicas and database overhead.

Cohere recommends 1,024-dimensional int8 vectors as a balance between storage size and retrieval quality, reporting performance close to float representations for both models. The database index must support both storing and comparing vectors in the selected representation.

Results on enterprise documents and finance tasks

On ViDoRe V3, a benchmark focused on visually complex enterprise documents, Cohere reports average scores of 85.8 for Embed 5 Pro and 84.7 for Fast. In the same company comparison, Voyage 4 Large scores 83.7, Gemini Embedding 2 scores 83.2 and OpenAI text-embedding-3-large scores 75.5.

In Cohere’s finance-focused evaluations, Pro scored 80.1 on FinanceBench, 90.0 on FinQA and 85.0 on ViDoRe V3 Finance. The company says Fast ranked second in all three tests. On FinanceBench, Cohere reports that Pro scored 21.4 points above OpenAI text-embedding-3-large.

Compared with Embed 4, Cohere reports its largest improvement in Farsi, at 13 points, with reported gains of 12 points each in Telugu and Hindi.

A new metric evaluates retrieved documents for relevance

Cohere optimized Embed 5 using RCP-nDCG@10, a retrieval-evaluation metric designed to address incomplete relevance labels. It uses a calibrated AI judge to assess every retrieved document against a shared rubric.

In a Cohere study involving 46 human annotators, existing benchmark labels scored 0.65 on AUC, a measure of how well they predicted human relevance judgments. A calibrated judge based on a large language model reached 0.91. Human reviewers considered 28% of the documents labeled irrelevant by the benchmark to be useful.

In pairwise system comparisons, Cohere’s research found that RCP-nDCG@10 matched human reviewers’ choice of system in 77% of comparisons, while the conventional nDCG metric matched their choice in 52%. Crediting relevant documents missing from the reference list accounted for a 19-point improvement in that research.

Availability through the API and deployment platforms

Embed 5 is available through the Cohere Embed API, Microsoft Foundry and Amazon SageMaker. Model Vault is available for single-tenant deployments.

API list pricing is $0.12 per million tokens for Pro and $0.08 for Fast, while hosted and single-tenant deployments may have platform-specific terms.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Qwen-Image-2.1 combines image generation and editing with transparent outputQwen-Image-2.1 combines image generation and editing with transparent output
Alibaba has released a single checkpoint for image generation and editing. Qwen-Image-2.1 supports transparent PNGs and, according to the company, accepts up to 10 reference images. Commercial deployment requires a separate license.
3 min read
3. 10. 2026
CoreWeave launches Forge for training and continuously improving AI agentsCoreWeave launches Forge for training and continuously improving AI agents
Forge connects model training, evaluation and improvement with insights from production. It offers post-training without a dedicated cluster, experiment analysis and isolated environments for agents.
2 min read
3. 10. 2026
Microsoft releases MAI-Transcribe-2-Streaming for live speech transcriptionMicrosoft releases MAI-Transcribe-2-Streaming for live speech transcription
MAI-Transcribe-2-Streaming produces text as audio arrives. Artificial Analysis ranked it first for final-transcript word error rate among 38 models. It is available through Voice Live API in public preview.
2 min read
3. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok