Kakao Releases 'Kanana-1.5-15.7B-A3B', an MoE Model Achieving Dense 8B-Level Performance with 3B Active Parameters
Key point
Kakao has open-sourced 'Kanana-1.5-15.7B-A3B', a model utilizing an MoE structure to achieve performance comparable to Dense 8B models with only 3B active parameters.
Details
The Kakao Kanana LLM organization developed and open-sourced the 'Kanana-1.5-15.7B-A3B' model on Kakao Hugging Face, introducing a Mixture of Experts (MoE) structure to maximize inference efficiency. This model adopts a Fine-grained structure by replicating and splitting the MLP layers of the existing Dense model (Kanana-Nano-1.5-3B) to form 64 ultra-small Experts, activating only 8 Experts per input. Training involved Upcycling from the Dense model followed by two stages of Pre-training (Stage 1: 700B tokens of high-quality general data, Stage 2: 300B tokens of domain-specific data). Benchmark results showed that with only 3B active parameters (approximately 37% of the total 15.7B), the model achieved performance equal to or better than Dense 8B models in certain categories, with further performance improvements in the Post-training stage through the application of Staged RL and On-policy Distillation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.