Cohere Embed 5 splits indexing and querying to cut RAG latency
Cohere released Embed 5, allowing developers to index with a high-quality Pro model and query with a faster Fast model in the same vector space.
Cohere launched Embed 5 on Wednesday, introducing a dual-model architecture that separates data indexing from query execution. This update enables engineering teams to use the higher-fidelity Embed 5 Pro for building vector indexes while serving live queries with the lower-cost, higher-throughput Embed 5 Fast, all without maintaining separate data structures.
What happened
The core innovation in this release is the shared embedding space between the two models. Previously, switching to a more efficient query model often required re-indexing the entire corpus or maintaining parallel indexes, which doubled storage costs and operational complexity. With Embed 5, developers can ingest documents using the Pro model to ensure high retrieval quality during the indexing phase. When the system serves user queries, it switches to the Fast model, which operates within the same vector dimensions and compatibility standards.
Cohere’s internal testing indicates that this hybrid approach results in minimal degradation of retrieval quality. Across 40 datasets that included text, images, fused documents, and parsed documents, the combination of Pro indexing and Fast querying achieved a relative score of 98.4 compared to a baseline of 100 for Pro-to-Pro usage. Using the Fast model for both indexing and querying dropped the score further to 96.6. The company noted that no individual dataset exhibited a major drop in performance when using the mixed Pro-Fast configuration.
This architectural shift addresses a common bottleneck in Retrieval-Augmented Generation (RAG) and agent workflows. In these systems, data ingestion happens infrequently, but search queries occur repeatedly and must be low-latency. By optimizing the heavy query traffic with the Fast model, teams can significantly reduce response times and infrastructure costs without sacrificing the precision established during the initial indexing phase.
How it works
Both Embed 5 Pro and Fast produce compatible vectors at identical dimensions, allowing them to coexist in the same index. This compatibility extends to advanced compression techniques like Matryoshka truncation and int8 quantization. Developers can choose from six vector dimensions ranging from 256 to 2,048, and select from float32, int8, or binary formats depending on their storage and precision requirements.
The storage implications are substantial for large-scale deployments. A standard 2,048-dimensional float32 vector occupies 8 KB, meaning a corpus of 100 million chunks would require roughly 819 GB. Switching to a 1,024-dimensional int8 vector reduces this footprint to approximately 102 GB. For even greater efficiency, a 256-dimensional binary vector shrinks the same dataset to about 3.2 GB. Cohere recommends the 1,024-dimensional int8 format for most use cases, as it balances memory savings with near-full-precision retrieval quality. Binary representations are suggested for initial retrieval stages where speed is critical, followed by higher-precision reranking.
The models also support multimodal inputs, including text, images, and fused text-image combinations, across more than 100 languages. They feature a 128K-token context window, allowing them to process large documents or complex visual data directly. This capability enables the embedding of page images or combined inputs into a single vector, streamlining the handling of diverse document types in modern AI applications.
Key details
- Embed 5 Pro costs $0.12 per million tokens, while Embed 5 Fast costs $0.08 per million tokens.
- The Fast model delivers an average of 2.4 times the document throughput compared to the Pro model in Cohere’s tests.
- Pro indexing with Fast querying scored 98.4 relative to a Pro-to-Pro baseline of 100 across 40 datasets.
- Both models support six vector dimensions from 256 to 2,048, with float32, int8, and binary format options.
- The models handle text, images, and fused inputs across more than 100 languages with a 128K-token context window.
- Embed 5 is available via Cohere’s API, Model Vault, Microsoft Foundry, Amazon SageMaker, and supports private VPC and on-premises deployment through vLLM.
Why it matters
For software engineers building RAG systems, the ability to decouple indexing quality from query latency is a significant operational advantage. Most production environments are read-heavy, meaning the cost and speed of queries dominate the total cost of ownership. By using a cheaper, faster model for the high-volume query path, teams can reduce API costs and improve user experience without the overhead of managing multiple indexes or re-processing data.
However, the benchmark results come with important caveats. Cohere evaluated Embed 5 using RCP-nDCG@10, a metric that uses query-specific relevance criteria rather than fixed labels. While this method can identify relevant results missed by traditional benchmarks, it measures reranking over a fixed candidate set rather than first-stage retrieval from the full corpus. First-stage retrieval was evaluated separately using standard nDCG and Recall, meaning the reported scores are not directly comparable across all evaluation types.
Production teams must validate these findings against their own data. Retrieval errors in agent workflows can compound through multiple steps, leading to significant downstream issues. Therefore, while the average score drop is small, specific domains or query patterns may behave differently. Engineers should benchmark the Pro-to-Fast setup against their specific corpus and query distribution before fully migrating.
What you can do
- Benchmark your current RAG pipeline using Embed 5 Pro for indexing and Fast for querying to measure latency and cost savings.
- Evaluate the impact of different vector dimensions and quantization formats on your storage costs and retrieval accuracy.
- Test multimodal capabilities by embedding page images or fused text-image inputs if your application handles diverse document types.
- Review your infrastructure to ensure it supports the chosen deployment option, such as vLLM for on-premises or private VPC setups.
- Monitor retrieval quality closely in production, especially if your agent workflow involves multiple sequential retrieval steps.
- Consider using binary vectors for initial retrieval stages followed by reranking if storage constraints are tight.