Multilingual-focused Aya Expanse models released
Key point
Cohere For AI unveiled the Aya Expanse 8B and 32B models, which maximize multilingual performance.
Details
Cohere For AI released the Aya Expanse family (8B, 32B) as open weights, developed to close the performance gap in multilingual models.
Aya Expanse 32B achieved a new SOTA (State-of-the-Art) on the Arena-Hard-Auto multilingual evaluation, outperforming Gemma 2 27B, Mistral 8x22B, and Llama 3.1 70B. Aya Expanse 8B also showed high win rates of 60.4%~70.6% compared to models in the same class, such as Gemma 2 9B and Llama 3.1 8B.
The key technical innovations are as follows:
- Data Arbitrage: Instead of relying on a single teacher model, it strategically samples from a pool of multiple models to generate high-quality multilingual synthetic data, preventing Model Collapse.
- Multilingual Preference Training: Preference training optimized for multilingual environments was applied.
- Model Merging & Safety Tuning: Both performance and stability were secured simultaneously through model merging and safety tuning.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.