Leaked source code reveals that the AI music generation service 'Suno' was scraping millions of songs from YouTube and Deezer.



Source code allegedly leaked from the AI music generation service 'Suno' has revealed the specific methods Suno used to collect millions of songs and lyrics from YouTube Music, Deezer, Genius, and other sources to build training data for its AI model. The code included the names of the services targeted, the scale of the collected data, and evidence that it used an external proxy service to obtain audio from YouTube.

Hack Reveals Suno AI Music Generator Scraped YouTube, Deezer, and Genius

https://www.404media.co/hack-reveals-suno-ai-music-generator-scraped-youtube-deezer-and-genius/



Suno AI music generator reportedly hacked, source code apparently data scraping | brief | SC Media
https://www.scworld.com/brief/suno-ai-music-generator-hacked-source-code-reveals-alleged-data-scraping

Suno snatched millions of songs from YouTube, Genius, and Deezer | The Verge
https://www.theverge.com/ai-artificial-intelligence/966072/suno-ai-music-training-scraping-youtube-hack

Suno is an AI service that can generate songs, including vocals and accompaniment, simply by inputting textual instructions or lyrics. However, it has not revealed much detail about the dataset used for training, and has been sued by the Recording Industry Association of America and others for using a large amount of copyrighted music without permission.

Music giants such as Sony, Warner, and Universal are suing music generation AI services 'Suno' and 'Udio' for copyright infringement - GIGAZINE



This information came to light when a hacker calling himself 'ellie.191,' who successfully gained unauthorized access to Suno, provided the data he obtained to the technology media outlet 404 Media. According to the hacker, the intrusion was carried out by using a worm called 'Shai-Hulud' to launch a supply chain attack on one employee and steal their credentials for GitHub and other cloud services.

The leaked documents included source code believed to be from 2023 and 2024, describing processes for retrieving data from YouTube Music, Deezer, Genius, Pond5, Jamendo, Freesound, and the International Music Score Library Project, among others.



Based on the code, it appears that Suno automatically retrieves audio and lyrics from various services, excludes non-music data, and then compiles them into a training dataset. The file 'youtube_music' recorded that it had imported 2,013,545 music clips from YouTube Music as of its last update.

In the data collection from YouTube, there is evidence that Bright Data, which provides proxies and infrastructure for web scraping, was used. By using a proxy, it is possible to send a large volume of access while switching the IP address of the connecting source, making it easier to continue automated data collection while avoiding access restrictions.

Furthermore, the code also included a process to search for a cappella versions of songs on YouTube. It's possible that they were trying to collect not only completed songs but also vocal-only recordings without accompaniment, and use the vocal data for training.

According to other leaked files, the dataset created by Suno included 113,879 hours of music from YouTube Music, 12,287 hours from Deezer, and 17,615 hours of Genius-related data. In addition, it is reported that 62,117 hours were collected from Pond5, 3,726 hours from Jamendo, and 19,514 hours from the International Music Score Library Project.



Another YouTube Music-related dataset called 'ytm_tagged' reached 152,162 hours. However, each dataset may contain overlaps, so simply summing them up may not represent the actual amount of training.

Suno was attempting to collect podcasts on a large scale, not just music streaming services. According to the code, it was planning to use PodcastIndex to identify approximately 420,000 shows that had published five or more episodes longer than 30 minutes, and download a total of approximately 1 million hours of audio.

The Recording Industry Association of America (RIAA) is in court alleging that Suno circumvented YouTube's copy protection measures and performed 'stream ripping,' saving streamed audio as files. The leaked code supports the claim that Suno was at least directly obtaining a large amount of audio from YouTube.

In response, Suno explained that its AI model was trained using music files and associated metadata accessible on the open internet. In past court documents, the company admitted to having collected a nearly comprehensive collection of reasonably good quality music files that were publicly available, while respecting passwords and paid access barriers, and argued that using copyrighted works to train the AI constituted fair use.

It has also been reported that, in addition to source code, users' email addresses, phone numbers, and payment information related to Stripe were obtained during the unauthorized access. Suno explained that it became aware of the limited security incident in November 2025 and contained it in a short period of time, stating that 'the leaked information mainly consisted of old code that is no longer in use, and highly confidential personal information was not compromised, and in the first place, Suno does not have access to all of its customers' credit card numbers.'

Furthermore, some customers have testified to 404 Media that they did not receive any infringement notice from Suno. Suno explained that, given the limited nature of the customer information that was discovered, they determined that individual notices were not required under applicable law.

in AI,   Web Service, Posted by log1i_yk