AI Briefing
KO
Pick

Nemotron 3 Embed, the No. 1 embedding model on RTEB

·2026.07.17 01:01

Key point

NVIDIA's Nemotron 3 Embed 8B model achieved No. 1 (78.5%) on the RTEB benchmark, and the company also released a lightweight 1B variant and a Blackwell-optimized version.

1 / 2

Details

NVIDIA has released three versions of its Nemotron 3 Embed embedding model for agentic search and RAG systems.

Model lineup:

  • Nemotron-3-Embed-8B-BF16: No. 1 on the RTEB benchmark (78.5%), achieving 75.5% on MMTEB
  • Nemotron-3-Embed-1B-BF16: 27% error rate reduction compared to the previous generation, 72.4% on RTEB (for lightweight deployment)
  • Nemotron-3-Embed-1B-NVFP4: optimized for the Blackwell architecture, 2x throughput and 99% memory savings compared to BF16

Key performance: More accurate retrieval reduces repeated queries and unnecessary reasoning by agents, lowering token costs. In evaluations, the Nemotron 3 Embed 8B model achieved the lowest expected agentic token cost on ViDoRe V3, BRIGHT, and BrowseComp-Plus.

Technical characteristics:

  • 32k context window
  • Support for multilingual and code retrieval
  • Open weights and fine-tuning recipes provided
  • Immediately available on HuggingFace, NVIDIA NIM microservices, and vLLM

Architecture: The 1B model compresses the Ministral-3-3B backbone through staged pruning and distillation, going from 3B → 2B → 1.14B, inheriting knowledge from the 8B teacher model to preserve accuracy.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.