AI Briefing
KO
Pick

Kakao Enhances Kanana-o with LM-SPT for Efficient, Style-Controllable Voice Generation

·2026.08.04 00:00

Key point

Kakao integrated LM-SPT into Kanana-o, improving voice efficiency and style control, achieving 94.50 on the Korean InstructTTSEval benchmark, close to Gemini-2.5-flash-preview-tts's 95.38.

1 / 4

Details

Kakao integrated LM-SPT, a low-frame-rate speech tokenizer, into Kanana-o to enhance voice generation efficiency and style control. LM-SPT operates at 12.5 frames per second, simplifying the architecture and reducing computational load. The system uses Multi-objective online reinforcement learning to better follow natural language style instructions. On the Korean InstructTTSEval benchmark, Kanana-o achieved a score of 94.50, demonstrating performance comparable to, though slightly below, Gemini-2.5-flash-preview-tts (95.38). Additionally, the model demonstrated cross-language transfer, applying style controls learned in Korean to English speech generation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.