大西洋月刊が AI 学習に使用された音楽の検索可能データベースを作成
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Verge AI
大西洋月刊の記者アレックス・ライズナー氏が、AI モデルの学習に利用されている4つの音楽データセットを特定し、計2100万曲以上を含む検索可能なデータベースとして一般公開した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
数百万の楽曲が、本来は公開すべきではないにもかかわらず、データセット内で自由に利用可能となっています。
数百万の楽曲が、本来は公開すべきではないにもかかわらず、データセット内で自由に利用可能となっています。
Terrence O'Brien による記事
2026年6月21日 GMT+9 午前3時46分


画像:Cath Virginia / The Verge
Terrence O'Brien は The Verge の週末編集者です。彼はテクノロジー業界を18年以上にわたり取材しており、シンセサイザーについても詳しい知識を持っています。
*Atlantic* 紙の記者 Alex Reisner 氏は最近、AI モデル(人工知能モデル)の学習に使用されている音楽の 4 つのデータセット を発見し、これらを一般向けに 完全に検索可能 にしました。そのうち2つのセットは非常に巨大で、それぞれ1200万曲と900万曲に及びます。残りの2つは比較的小さいものの、それでも各セットとも10万曲を超える楽曲を含んでおり、学習データとして重要な規模を占めています。
ライズナーによると、これらのセットは数千回ダウンロードされており、誰が使用したかを正確に知ることは不可能ですが、Google と Stability はともに研究論文においてそれらを使用したことを確認しています。Free Music Archive データセットなどの一部のソースは、個人利用のためにストリーミングが無料ですが、商用アプリケーションにはライセンスが必要です。
これらのデータセットは理論上インターネット上で自由に入手可能ですが、トレーニングデータとして使用するには、単に ZIP ファイルをダウンロードして AI モデルに読み込ませるだけでは不十分です。ライズナーの説明によると:
**私が発見した 3 つのデータセットは、YouTube または Spotify の楽曲へのリンク一覧として配布されています。AI 開発者は、自動化されたツールを使用して実際のオーディオファイルをダウンロードしますが、これらのツールの一部には、クリエイターが収益を得たり登録者を増やしたりするためのログイン、広告、および仕組みを回避できるものもあります。このようなツールは、各プラットフォームの利用規約に違反しています。
データセットに登場する名前は、ポップスターの Lady Gaga や Fred Again.. から、Radiohead、Aphex Twin、Wu-Tang Clan、Bruce Springsteen、そして実験音楽作曲家の Hainbach まで多岐にわたります。ご自身で、世界の AI モデル(人工知能モデル)の学習に使用されている楽曲や書籍、その他のメディアを検索するには、*Atlantic* の AI ウォッチドッグ サイトへお越しください。
この記事のトピックや著者をフォローして、パーソナライズされたホームページフィードで類似の記事をもっとご覧になったり、メールでの更新を受け取ったりしてください。
- Terrence O'Brien
-
-
-
The Verge Daily
最も重要なニュースを毎日お届けする無料ダイジェスト。
メールアドレス(必須)
原文を表示
Millions of tracks are freely available in datasets, even if they’re not supposed to be.
Millions of tracks are freely available in datasets, even if they’re not supposed to be.
by Terrence O'Brien
Jun 21, 2026, 3:46 AM GMT+9


Image: Cath Virginia / The Verge
Terrence O'Brien
is the Verge’s weekend editor. He’s covered the tech industry for over 18 years and knows a thing or two about synths.
*Atlantic* reporter Alex Reisner recently uncovered four datasets of music being used to train AI models and made them fully searchable for the public. Two of the sets are absolutely enormous at 12 million and 9 million tracks. The other two are much smaller, but still represent a significant amount of training data at over 100,000 songs each.
According to Reisner, the sets have been downloaded thousands of times and, while it’s impossible to know exactly who has used them, Google and Stability have both confirmed they have in research papers. Some of the sources, like the Free Music Archive dataset, are free to stream for personal use but require licensing for commercial applications.
While the datasets are freely available on the internet in theory, using them as training data is not as simple as downloading a ZIP file and feeding it to an AI model. As Reisner explains:
Three of the datasets I found are distributed as a list of links to songs on YouTube or Spotify. AI developers download the actual audio using tools that automate the job, some of which allow developers to bypass logins, advertisements, and mechanisms that might earn money or subscribers for creators. Such tools violate the terms of service of these platforms.
Names that pop up in the dataset range from pop stars like Lady Gaga and Fred Again.., to Radiohead, Aphex Twin, Wu-Tang Clan, Bruce Springsteen, and experimental composer Hainbach. You can hop over to the *Atlantic’s* AI Watchdog site and search through the songs, books, and other media being used to train the world’s AI models yourself.
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
- Terrence O'Brien
-
-
-
-
The Verge Daily
A free daily digest of the news that matters most.
Email (required)
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み