The Atlantic Publishes Searchable Database of Music Used to Train AI
Journalist Alex Reisner has located and made searchable four music datasets used to train AI models, two of them containing over 9 and 12 million tracks respectively.
Twelve million tracks in a single dataset. That's what journalist Alex Reisner from The Atlantic discovered while investigating audio data being used to train artificial intelligence models. This week, The Verge covered the project: Reisner has identified four music datasets and made them available to the public through a fully searchable interface.
The two largest datasets contain approximately 12 and 9 million tracks respectively. The other two are smaller but still represent considerable volumes of protected, or potentially protected, material used without most artists, labels, or listeners being aware of it.
What these datasets contain and why they matter
The particularity of Reisner's work is not just locating these datasets, a laborious task in itself, but structuring them so that any musician, rights manager, or researcher can search whether their work is included. Until now, opacity surrounding the training data of generative audio models was nearly complete: companies are not required to publish inventories of their sources and rarely do so voluntarily.
The initiative carries additional weight in the current legal context. Several countries have open litigation over the unauthorized use of protected works to train AI. In the United States, artists and labels have filed lawsuits against various technology companies alleging copyright infringement. In Europe, the debate over practical implementation of the AI Act's data mining exemption remains largely unresolved. Knowing what's in a training dataset is the first step toward any legal claim or licensing negotiation.
Who this tool is useful for
Reisner's database has practical value for at least three groups:
- Artists and songwriters who want to know if their catalog has been used without explicit consent and evaluate whether they have grounds to file claims.
- Legal teams and rights managers who need baseline documentation for litigation or negotiations with AI companies.
- Researchers and journalists interested in auditing the actual composition of data underlying generative audio models, a still-emerging discipline.
The underlying problem: opacity as standard practice
What makes this journalistic work significant is precisely what it reveals about the state of the industry: that months of investigation were needed to locate and make accessible data that should, in principle, be auditable says a great deal about how transparency has been managed in the sector so far.
This isn't Reisner's first time tackling this terrain. Back in 2023, he published investigations into text datasets used to train LLMs, with results that generated significant coverage and some pressure on affected companies. The pattern repeats itself: training data is collected at scale, the industry advances, and accountability comes later, if it comes at all, and from outside.
The tool doesn't resolve the structural problem, but at least it makes it visible. That artists and users can search their name in a database and find concrete answers represents a shift in scale from the previous situation. Whether that translates into regulatory changes or new licensing models will largely depend on what those with the capacity to litigate or legislate do with this information.
From our perspective, we view it positively that data journalism provides this kind of public transparency infrastructure. The conversation about what data feeds models has been too abstract for too long. Having a searchable tool is at least a concrete starting point.
Sources
Read next
Synthesia moves from corporate video to avatar roleplay
Synthesia launches AI Roleplay Sessions: employees rehearse tough conversations with avatars that score and give feedback. What changes and for whom.
AI mania is degrading decision-making at large companies
Nik Suresh collects anecdotes from his consulting work: AI strategies signed off by executives who never used the tools, plus internal token consumption leaderboards. Simon Willison recommends it.
Spotify removes 75 million spam and 'AI slop' tracks
Spotify says it pulled 75 million fraudulent tracks in a year and rolls out a spam filter, AI use disclosure and a tougher rule against voice impersonation.