AI Briefing
KOSign in

Cohere Releases Embed 5: Pro and Fast Models Share Embedding Space for Enterprise Search

·2026.09.30 18:51

Key point

Embed 5 Pro achieves an average score of 85.8 on ViDoRe V3, outperforming competitors like Voyage 4 Large and Gemini Embedding 2.

1 / 10

Details

Cohere has released Embed 5, a new family of embedding models comprising Pro and Fast tiers. The core architectural innovation is a shared embedding space, allowing users to index documents with the high-quality Pro model and query with the low-latency Fast model without re-indexing. This hybrid approach reduces costs by up to one-third for query operations while maintaining near-baseline search quality.

Performance and Benchmarks

Embed 5 Pro demonstrates significant improvements in enterprise document retrieval. On the ViDoRe V3 benchmark for visually rich documents, Pro scores an average of 85.8, an 8.8-point increase over Embed 4. It surpasses competitors including Voyage 4 Large (83.7), Gemini Embedding 2 (83.2), and OpenAI text-embedding-3-large (75.5). In financial domains, Pro leads on FinanceBench (80.1) and FinQA (90.0), outperforming the next best non-Cohere model by an average of 3.3 points.

The Fast tier is optimized for latency-sensitive workloads, offering 2.4x higher throughput than Pro. It achieves a ViDoRe V3 score of 84.5, beating smaller competitors like Jina v5 Small (74.5) and Qwen3-VL-Embedding-2B (64.2) by substantial margins.

Multilingual and Multimodal Capabilities

Embed 5 supports over 100 languages with a 128K token context window. In European languages, Pro averages 77 points, with notable gains in Persian (+13), Telugu (+12), and Hindi (+12) compared to Embed 4. For multimodal tasks, Pro handles fused text-image inputs with an average score of 82.3 across five datasets, significantly outperforming Fast (81.2) and Gemini Embedding 2 (61.3).

Deployment and Optimization

The models are available via Cohere API, Model Vault, Microsoft Foundry, and Amazon SageMaker. Cohere introduces RCP-nDCG@10, a new evaluation metric that better reflects corpus-wide performance by accounting for query-specific relevance criteria. To optimize storage, Embed 5 supports Matryoshka embeddings and quantization formats (float, int8, binary). Using 256-dimensional binary vectors can reduce storage requirements by up to 256x compared to 2,048-dimensional float32 vectors, dropping raw vector storage for 100 million chunks from ~819 GB to ~3.2 GB.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.