Dark Mode Light Mode

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Follow Us
Follow Us

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

AI Training Under Fire After Investigation Reveals Millions Of Copyrighted Songs 

Image Source: , Billie Eilish, Music Live 

The debate over AI and music copyright has intensified after an investigation revealed that more than 21 million copyrighted recordings are circulating among datasets used to train artificial intelligence systems. 

Conducted by journalist Alex Reisner for The Atlantic, the investigation uncovered four major music datasets containing recordings from both globally recognised artists and independent musicians. Among those identified are Flume, Tame Impala, Sia, Billie Eilish, Taylor Swift, Nirvana, Bad Bunny, and thousands more, highlighting the scale at which copyrighted music has become intertwined with AI development.

According to the report, two of the datasets contain more than 100,000 recordings each, while the remaining two are substantially larger, housing between nine and 12 million tracks apiece. The datasets have reportedly been downloaded thousands of times, although it remains unclear which AI companies have incorporated them into commercial training models due to the industry’s limited transparency around training data.

The investigation did identify Google and Stability AI as having used the Free Music Archive dataset during AI training. However, whether the larger copyrighted collections have been utilised by specific developers remains unknown.

To improve transparency, The Atlantic has launched an AI Watchdog search tool, allowing artists to check whether their work appears within the four datasets. Early searches have revealed hundreds of recordings attributed to electronic music artists, including Eric Prydz (54 tracks), Honey Dijon (126), Björk (411), Moby (213), Fatboy Slim (175), The Chemical Brothers (153), Daft Punk (151), and Charlotte de Witte (89), among many others.

The findings have prompted immediate backlash from musicians. Singer-songwriter SZA publicly criticised the discovery after finding hundreds of her own recordings listed in the datasets, expressing concern that even unreleased material may have been included without permission.

The investigation arrives as legal and ethical debates surrounding generative AI continue to intensify across the music industry. Rights holders, publishers, and collecting societies have increasingly questioned whether copyrighted works can be used to train AI models without explicit licensing agreements, while technology companies argue that existing copyright law leaves room for interpretation.

The findings are expected to add further pressure on AI developers and lawmakers as artists, publishers, and rights organisations continue to push for greater transparency and licensing requirements around the use of copyrighted music in AI training. 

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Previous Post

Ibiza Drug Network Case Sees 9.5-Year Prison Sentence Upheld

Next Post

Apple and Xbox Raise Prices as AI Boom Drives Global Chip Shortage