NVIDIA, 6 Million-Example Multilingual Reasoning Dataset
Key point
NVIDIA has released the Nemotron multilingual reasoning dataset, comprising 6 million examples across 5 languages, along with the Nemotron Nano 2 9B model.
Details
NVIDIA has released the Nemotron Post-Training Dataset v2 to support the open ecosystem. This dataset consists of 6 million examples created by translating existing English reasoning data into 5 languages: French, Spanish, German, Italian, and Japanese.
To improve data quality and prevent hallucination, NVIDIA applied verification mechanisms such as sentence-level translation, format enforcement using special brackets (〘〙), and fastText-based language identification.
The accompanying NVIDIA Nemotron Nano 2 9B model uses a Transformer-Mamba hybrid architecture, delivering up to 6x higher token generation speed compared to existing models. It also features a 'Thinking budget' function that can reduce inference costs by up to 60%, making it optimized for deployment in customer service agents or edge devices.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.