AMD Releases Instella-MoE
Key point
AMD has released a fully open 16B MoE model trained on ROCm.
Details
AMD has released Instella-MoE, a sparse-activation MoE language model with 16B total parameters and 2.8B active parameters per token. The model was trained from scratch using only AMD Instinct MI300X and MI325X GPUs and the ROCm software stack.
Instella-MoE aims to be a fully open model, releasing not only weights but also the following training assets:
- Training code and configurations
- Data mixtures and stage-by-stage training recipes
- Pre-training and post-training checkpoints
- Intermediate checkpoints and model configuration files
The training utilized AMD's open-source framework Primus, while reinforcement learning-based post-training used Miles. The architecture incorporates Gated MLA and FarSkip-Collective.
The entire training pipeline consists of 6 stages, progressing from first- and second-stage pre-training, mid-training, long-context extension, supervised fine-tuning, to reinforcement learning. A total of 7.1T tokens were used for pre-training. Stage-by-stage checkpoints allow for the reproduction and analysis of the impact of mid-training and post-training on model capabilities.
Comparative results of the released materials indicate that Instella-MoE achieves a performance band higher than OLMo-3-7B, which has more than twice the active parameters, with 2.8B active parameters. The key takeaway is the addition of both MoE and AMD GPU-based training cases to the fully open model ecosystem.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.