Kakao Releases 'Melon Playlist Dataset' Based on Melon DJ Playlists
Key point
Includes metadata for over 148,000 playlists and 649,000 tracks, along with Mel spectrograms.
Details
Kakao's recommendation team has released the Melon Playlist Dataset, based on DJ playlists from the music streaming service Melon. This dataset was first introduced at the 3rd Kakao Arena competition and can be used for developing collaborative filtering and content-based filtering models, as well as for Playlist Auto Tagging research.
The dataset consists of JSON files for training, validation, and testing, track metadata, a genre mapping table, and Mel-Spectrogram files extracted from audio sources. Additionally, a paper utilizing this dataset was accepted to ICASSP 2021, acknowledging its academic value.
The dataset includes a total of 148,826 playlists and music data for 649,091 tracks after deduplication. It contains 30,652 tags, 107,824 artists, and 269,362 albums after deduplication. The playlists were curated by Melon DJs who selected tracks and assigned tags and titles.
The Mel-Spectrograms were generated from audio data with a 16kHz sampling rate, targeting segments between 20 and 50 seconds. Frames were created using a window size of 512 and a hop length of 256, and 48 mel filters were applied to each frame to convert them into matrices of size 48 x 1876. These matrices are matched to track IDs, saved in npy format, and provided as 40 split compressed files.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.