AI Briefing
KO

Kakao's Kanana Essence, trained on 3T tokens, surpasses llama-3.1-8b benchmark average

·2024.11.14 00:00

Key point

Kakao's Kanana Essence, trained on 3T tokens, achieved higher benchmark averages than llama-3.1-8b, while Kanana Nano maximized efficiency through pruning and distillation techniques.

1 / 8

Details

Kakao adopted a Two-staged pre-training strategy to overcome the Scaling law limitations of existing LLMs. Through this strategy, Kanana Essence was trained on a total of 3T tokens (Stage 1: 2.7T, Stage 2: 0.3T), achieving a benchmark average score of 57.52 with significantly fewer tokens than those used to train llama-3.1-8b, surpassing llama-3.1-8b (50.54). Notably, it showed significant performance improvements on Korean benchmarks (KMMLU, HAE-RAE). However, the score on the English benchmark (MMLU) was slightly lower than that of llama-3.1-8b. The ultra-lightweight model Kanana Nano applied Pruning & Distillation techniques instead of From Scratch training, securing superior performance (average 52.07) with only 1/10 of the data compared to From Scratch training.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.