Yandex releases Yambda dataset for recommendation systems

Yandex has released Yambda, a new dataset containing approximately 5 billion anonymized user interactions collected from its music streaming platform, Yandex Music.

12punto

Yandex has announced Yambda, a new dataset containing approximately 5 billion anonymized user interactions from its music streaming platform, Yandex Music. This dataset stands out as the world's largest open-source event data for recommendation systems.

Yambda includes interactions such as listening, liking, and disliking collected over a 10-month period. The data is provided alongside timestamps, audio embeddings, and organic discovery information, enabling recommendation algorithms to be tested under real-world conditions. Details on 1 million anonymized users and 9.3 million music tracks offer a critical resource for those developing recommendation models in fields such as e-commerce, social media, and short-form video platforms.

Yandex has made Yambda available via Hugging Face in three different sizes (50M, 500M, and 5B events). The dataset is provided in Apache Parquet format and is compatible with systems such as Spark, Hadoop, and Pandas.

Nikolai Savushkin, Head of Recommendation Systems at Yandex, states that Yambda brings together both academia and the industry and will accelerate innovation in recommendation systems.