Yandex Music Replaces 15+ Recommender Components with Single Transformer Model 'Sona' in A/B Test
Key point
The single-model approach achieved +4.53% Active Users and +6.30% Total Listening Time in a 7-day A/B test on smart speakers.
Details
Yandex Music has developed Sona, a single transformer model that replaces the production recommender's traditional architecture of 15+ candidate generators, pre-ranking, and ranking models. While the model has not yet shipped to full traffic, an A/B test on smart speakers demonstrated significant performance gains over the production control.
Architecture and History Compression
To manage the computational cost of full attention over long sequences, Sona reads up to 8,192 events using a technique called History Compression, which roughly halves inference costs. The history is split into two blocks:
- Older events: 6,144 events
- Recent events: 2,048 events
These blocks exchange information via cross-attention and a single full-history self-attention layer. A subsequent 7-layer stack processes only the recent 2,048 events. This design retains most of the quality of full attention while keeping older events visible to the decoder and Ranking Module, which both read the same encoder output to ensure the encoder runs only once per request.
A/B Test Results
In a 7-day A/B test involving 15% of users in each arm, Sona outperformed the production stack with statistically significant results (p < 0.01):
- +4.53% in Active Users
- +6.30% in Total Listening Time
Candidates are generated via beam search as Semantic IDs and scored immediately. While engagement metrics improved, the team noted that catalog coverage is lower than with the production stack and is investigating the cause. A long-term A/B test is currently underway.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.