Cohere releases Tiny Aya L2-Thinker for in-language reasoning across 60 languages
Key point
The 3.35B parameter model achieves over 93% in-language reasoning rates by using a three-pillar data mixing strategy.
Details
Cohere has released Tiny Aya L2-Thinker, a 3.35B parameter model designed to perform reasoning directly in the user's language rather than defaulting to English. This approach addresses the limitations of existing models, which often hide reasoning traces from non-English speakers and lose cultural nuances. By enabling in-language reasoning, the model allows users to verify and correct the AI's thought process while leveraging native cultural knowledge.
Data Mixing Strategy
The model's performance relies on a specific data mixing strategy composed of three pillars:
- English reasoning data (ER): Approximately 1.7 million step-by-step solutions generated by gpt-oss-120b to build core problem-solving capabilities.
- Multilingual reasoning data (MR): English reasoning examples automatically translated into 44 languages, which increased the in-language reasoning rate from 12.8% to 86.1%.
- Multilingual non-reasoning data (NR): Question-answer pairs without reasoning traces to help generalize to new languages.
Combining all three pillars resulted in a 57.4% accuracy rate and a 96.4% in-language reasoning rate across benchmarks.
Performance and Efficiency
Tiny Aya L2-Thinker demonstrates strong performance across 60 languages and six benchmarks, including MGSM, PolyMath, and GlobalPIQA. Compared to its English-only counterpart, Tiny Aya En-Thinker, the model maintains similar accuracy on five out of six benchmarks, with a slight improvement in open-ended writing tasks. The primary exception is PolyMath, where competitive-level mathematics accuracy drops, likely due to the absence of a reinforcement learning stage.
The model excels in efficiency, particularly for low-resource languages. It utilizes a multilingual-optimized tokenizer to keep thinking tokens below 5,000 on average. In contrast, competing models like Magistral and M-Thinker suffer from significant performance degradation or increased reasoning length in low-resource settings. Tiny Aya L2-Thinker avoids the "doomlooping" phenomenon, where models repeat steps without reaching a solution, maintaining high reasoning quality even as resource availability decreases.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.