AI企業が旧書を学習データとして購入
AI企業が生成されたテキストによる汚染を避けるため、2022年以前の印刷書籍を大量に購入しトレーニングデータとして活用する動きが加速している。
キーポイント
AIスロップ回避のための物理的書籍の価値向上
Webスクレイピングで入手可能なデータの多くが生成AIによる汚染を含んでいるため、2022年以前の印刷書籍は「構造的にクリーン」なデータとして注目されている。
モデル崩壊(Model Collapse)のリスクと対策
AI生成データを基に学習を繰り返すと性能が劣化する「モデル崩壊」を防ぐため、人間によって編集・査読された書籍データへの依存度が高まっている。
ISBNdbによる大規模な書籍取得プラットフォーム化
ISBNdbはAI企業向けに1,000冊から100万冊単位の印刷書籍購入・スキャン支援サービスを提供し、データ収集の効率化を推進している。
著作権訴訟と倫理的な対立の激化
AnthropicやGoogleが書籍のスキャンを巡り著作権侵害で訴えられており、著者側も意図的にモデルを破壊する「ポイズニング」を行う動きが出ている。
AI企業による書籍購入の匿名性と倫理的配慮
ISBNdbは、AI企業が印刷物を破棄する悪印象(optics problem)を避けるため、秘密保持契約(NDA)の下で身元を隠して大量の古書を取得している。
専門書商による販売急増と在庫の消滅懸念
4月以降、外国語や希少な書籍の販売が歴史的な水準に急増しており、AI学習用として大量購入された結果、入手困難な書籍が失われる可能性が指摘されている。
AI企業の購入パターンの特徴
購入者は特定の分野に限定されず、価格に関係なく高価な書籍も大量に購入しており、その不規則性と資金力からAI企業による購入と推測されている。
重要な引用
"The world's best AI training data is sitting on a shelf."
"Print books from the pre-LLM era are structurally guaranteed to be free of this contamination."
"Books represent curated, peer-reviewed, domain-specific human knowledge..."
'AI company destroys two million books' is not a headline that generates sympathy.
I don't like the end-use, and I don't like that uncommon books are being pulped.
"Basically, almost every library in the world has lost their budget... There's a total disregard for the price of the book. That's kind of a tell for AI because they have just so much money."
影響分析・編集コメントを表示
影響分析
この動向は、生成AIの学習データ供給源がインターネット上のスクレイピングから、物理的かつ高品質な書籍へとシフトする重要な転換点を示しています。これにより、モデルの性能維持や倫理的なデータ利用の観点から、伝統的な出版業界とAI企業間の新たな協力・対立構造が生まれる可能性があります。
編集コメント
AI業界が「データの質」に直面する中、物理的な書籍というアナログなリソースがデジタルの未来を支える重要なインフラとして再定義されています。これは単なるデータ収集の手段を超え、生成AIの健全性を保つための新たな戦略的アプローチと言えます。
AI 企業がモデルの精度向上のためにさらなる学習データを模索する中、ある企業は古書こそが理想的なデータ源だと提案している。その理由は、これらの書籍には AI 企業が生成した「ゴミのようなデータ」が含まれていないことが確実だからだ。
"世界最高の AI 学習データは、すでに棚に眠っている"。AI 企業向けに大量の書籍取得サービスを提供し、自社のデータベースを"世界最大級"と謳う ISBNdb は、そのウェブサイトでこう述べている。「書籍とは、キュレーションされ、査読を経て、特定の分野に特化した人間の知恵です。ウェブクロールでは再現できない形で構造化されています。密度が高く、編集され、権威ある情報なのです」
ある記事では、ISBNdb が「2022 年以前に出版された印刷書籍は、AI 生成テキストが含まれていないため、AI の学習データとして理想的である」と説明しています。同記事が正しく指摘している通り、現在 AI 企業がインターネットから収集できるデータの多くには AI 生成テキストが混在しており、これが「モデル崩壊」を引き起こす恐れがあります。これは、AI 生成データを学習対象とした結果、エラーを起こしやすくなる劣化したモデルが生まれるプロセスを指します。
また記事では、自身の著作が学習目的でスクレイピングされることに反対する著者たちが、現在では AI モデルを意図的に汚染させる手法を用いることで、生成された AI モデルの機能を撹乱・破壊する文章を作成しやすくなった点にも触れています。
「LLM 登場以前の印刷書籍は、構造的にこの汚染から免れていることが保証されています。これ自体が大きな利点です……」
「この日付(2022 年以前)までに出版された物理的な書籍は、現代的な汚染ツールによる影響を完全に受けていません。」
ISBN は「国際標準図書番号」の略称で、書籍の裏面に記載される数値形式の商業識別子兼バーコードです。長年 ISBNdb は、書店や図書館、流通業者が在庫管理を行い、書籍を検索・販売するのを支援してきましたが、生成 AI の急成長により、AI 企業にとって極めて貴重な存在となっています。
同社は書籍メタデータへのアクセス販売に加え、現在では AI ラボに対して、1 回あたり 1,000 冊から最大 100 万冊に及ぶ大量の印刷書籍を調達する仲介役も担っています。ISBNdb のデータベースを活用すれば、AI 企業は重複を防ぎながら、体系的に印刷書籍を取得・スキャンし、トレーニングデータへと変換することが容易になります。
AI 企業がトレーニング用として印刷書籍を買い占めようとする動きは、1 月に作家たちによる Anthropic への著作権訴訟で注目を集めました。この訴訟で開示された内部文書には、同社が数百万冊の印刷書籍を取得・スキャンし、その過程で破棄する計画の詳細が記されていました。ワシントン・ポスト紙の報道によると、Anthropic は「ベター・ワールド・ブックス」という企業から書籍を購入しており、これは図書館や小売店、個人が書籍を販売できる複数の市場の一つです。また、Google も同様に、著作権で保護された書籍を用いて Gemini をトレーニングしたとして出版社から訴訟を起こされています。
ISBNdb は、AI 企業の身元を秘匿できることを謳っています。
「すべての取引には厳格な秘密保持契約(NDA)を締結します」と、ISBNdb のサイトは明記しています。「すべてのプロジェクトは法的拘束力のある秘密保持契約から始まります。あなたの身元、戦略、および購入対象が他者に開示されることはありません。」
ISBNdb は、AI 企業がスキャン工程で印刷された書籍を破棄している現場に遭遇したくないと考えている可能性があると指摘しています。
「イメージの問題は現実的な課題です」と同サイトは続けます。「『AI 企業が 200 万冊の書籍を廃棄』という見出しが、共感を呼ぶことはありません。」
これらの市場で外国語書籍の販売に特化するあるプロの古書商は、私にこう話しました。4 月以降、彼自身や他の古書商たちが歴史的な販売数の急増に気づき始めたといいます。この古書商は、プラットフォーム上で引き続き取引を続けるために匿名性を保つよう希望しています。
「個人的には、これらすべてに対して複雑な感情を抱いています」と、AI 企業向けにトレーニングデータとして数百冊の書籍を売却したと推測する古書商は語りました。「財務的には利益となり、売れにくい古い在庫を整理できる点でもメリットがあります。私は海外からの在庫や外国語書籍を活用して、こうした販売に適応してきました。一方で、最終的な用途には納得がいきませんし、希少な書籍がパルプ化されてしまうことにも不満を抱いています。」
この古書商によると、彼の在庫には希少品や外国語書籍、流通量の少ない本が多数含まれています。これらがトレーニングデータへの加工過程で破棄されれば、今後入手することがさらに困難になるでしょう。
販売者は、通常であれば好調な週でも週に約20冊しか売れないと言います。しかし4月以降は、毎週数百冊が定期的に売れるようになりました。購入元がAI企業であるという明確な証拠はないものの、販売者自身はその可能性を強く疑っています。
まず、彼が扱う書籍は専門書が多く、通常は学校や図書館によって購入されるものです。しかし、予算削減の影響でこうした組織からの需要は減少傾向にあります。また、大量購入は特定のトピックへの関心を反映するものですが、最近の非常に大規模な注文では、ISBN(国際標準図書番号)を持っているという一点以外に共通点がありませんでした。
この販売者はISBNを持たない希少書も扱っていますが、それらは今回の大量購入の対象には含まれていません。なお、これらの売上がISBNdbを介して行われたことや、顧客がAnthropic社や他のAI企業であるという確実な証拠は現時点では確認できていません。
「数量だけでなく、注文の奇妙さにも驚かされます」と、古書商は語りました。「以前も図書館からの注文はありましたが、通常は特定の分野や、せいぜい少し広い範囲に限定されていました。しかし現在、世界のほぼすべての図書館が予算を失っています。米国の大学図書館もほとんど購入しなくなりました。オーストラリアの図書館も同様です。注文される書籍の種類には、全く理屈がありません。さらに、価格に対する無関心も顕著です。この経路で販売された書籍の中には、明らかに高価すぎるものもありました。これは AI の仕業を示す一つの証拠でしょう。AI には資金が潤沢にあるからです。」
「私だけでしょうか?それとも昨年後半以降、自動購入(AutoBuy)の注文数が増えているのでしょうか?」と、ある古書商は 2 月、書籍販売市場である Alibris のフォーラムに投稿しました。Alibris の自動購入機能を使えば、顧客が欲しい本を登録しておき、その本が販売開始された際に自動的に購入できるようになります。「何が起きているのかコメントをください。AI がすべての本を読み込もうとしているのでしょうか?選定基準についての洞察はありますか?状態や形式(ハードカバーかソフトカバーか)、価格などに大きなばらつきがあるように見えます。」
「新規の大口購入者が現れ、一般書籍を大量に買い占めているため、多くの販売者から注文が殺到しています」と、Alibris のクライアントサービス担当ディレクターであるマイク・フェルドマン氏は回答しました。
ある古書販売業者は、同様の大量注文が「Biblio」と呼ばれる別の市場を通じて行われていると明かしました。顧客は購入したい書籍の ISBN を記載したスプレッドシートを Biblio に提供し、あとは同社が処理を進めます。
6 月にはオランダのある出版物が、複数の古書販売業者への取材を行いました。彼らは AI 企業による大量購入だと推測される、同様の大口注文があったと報告しています。
ISBNdb や Biblio、Alibris といった書籍市場やデータベースが購入者の身元を非公開にしているため、これらの書籍が AI 企業のトレーニングデータ収集のために購入され、あるいは使用後に破棄されているという事実を断定するのは困難です。大量の書籍はまず配送センターへ送られ、そこで Alibris が例のように書籍の品質チェックを行った上で、最終的な顧客へと発送されます。
著作権訴訟で明らかになった Anthropic の内部文書には、同社が数百万冊のスキャン計画を立てていたことが記されていますが、なぜその過程で書籍を破棄する必要があったのかは明確ではありません。Google Books の作成にも携わり、Anthropic からプロジェクト統括を任された Tom Harvey 氏の証言録取では、スキャン業務を委託された企業のひとつに Datamation が含まれていることが示されました。Datamation は「高容量の破壊的・非破壊的書籍スキャン」サービスの両方を提供しています。破壊的スキャンでは本の背表紙を切断してページを取り出し、機械に読み込ませるため、非破壊的な方式よりも高速かつ低コストで処理できます。
著作権訴訟において、著者側を訴えたアンソロピック社に対し、連邦裁判所のウィリアム・アルサップ判事は、デジタルコピーの作成が違法ではないと判断しました。その根拠は、元の書籍が廃棄された点にあります。
「本件では、購入した印刷版のすべてのコピーが、保存スペースの節約と、デジタル化による検索機能の向上を目的として行われました」とアルサップ判事は判決文で述べています。「印刷された原本は破棄され、新しいデジタルコピーがその役割を引き継ぎました。また、新たなデジタルコピーが社外で閲覧・共有・販売されたという証拠は一切ありません」。
アルサップ判事はこれを「明白に変容的な行為」とし、著作権法第107条に基づくフェアユース(公正利用)に該当すると結論づけました。
この法的根拠は、ISBNdbのウェブサイトでもAI企業向けに紹介されています。同サイトには、「二次市場から大量の紙書籍を購入しても、創作者が本来受け取るはずだった収入を奪うことにはならない」と記載されています。「これらはすでに創作者に対する金銭的な義務を果たし終えた書籍だからです」。
また、廃棄された書籍のリサイクルが重要である理由については、「責任ある物理的な調達とは、単なる書物焼却ではない。それは書籍のライフサイクルを完結させる行為だ」と説明しています。「木から知識へ、そして再び木へと還る循環こそが、その本質です」。
ISBNdbとアンソロピック社はいずれもコメント依頼に対して回答していません。
原文を表示
imageAs AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.
“The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”
In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.
“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”
ISBN stands for International Standard Book Number, the numerical commercial book identifier and barcode on the back of most books. For years, ISBNdb helped book sellers, libraries, and distributors manage their inventory and find and sell books, but the generative AI boom has made it valuable to AI companies. In addition to selling access to book metadata, ISBNdb now helps AI labs source bulk printed book purchases of between 1,000 to 1 million books per order. ISBNdb’s data makes it easier for AI companies to methodically acquire, scan, and turn printed books into training data while avoiding duplication.
AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process. The Washington Post article found that Anthropic was buying books from one company called Better World Books, one of several marketplaces where libraries, retailers, and individuals can sell their books. Google was recently sued by book publishers for similarly training Google Gemini on copyrighted books.
ISBNdb advertises that it can keep the identity of AI companies secret.
“Strict NDA [non-disclosure agreement] on every engagement,” ISBNdb’s site says. “Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed.”
ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process.
“The optics problem is real,” ISBNdb’s site says. “‘AI company destroys two million books’ is not a headline that generates sympathy.”
One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. This bookseller asked to remain anonymous so he can continue to do business on these platforms.
“I personally have mixed feelings about all of this,” the bookseller, who suspects he’s sold hundreds of books to AI companies for training data, told me. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”
This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
The seller told me that, normally, on a good week, he’d sell about 20 books. Since April, he has regularly sold hundreds of books a week. While the seller didn’t have clear evidence that the purchases were being made by AI companies, the purchases made him suspect that they were. First of all, he said, the kind of books he sells are specialized and are usually bought by schools and libraries. Purchases from these organizations have been trending downward because of reduced funding, he said. Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases. I have not seen any evidence that this bookseller’s recent sales were facilitated by ISBNdb or that the client was an Anthropic or another AI company.
“It's not just the quantity, but the weirdness of the orders,” the bookseller told me. “I've had library orders before, and usually they're mostly confined to a single subject or maybe a slightly broader range of subjects. But basically, almost every library in the world has lost their budget. I know all the U.S. college libraries don't buy much anymore. The Australian libraries don't buy much anymore. The type of books [...] there's no rhyme or reason to it. Also, there's a total disregard for the price of the book. I've had some books that sold through this way that were [...] greatly overpriced. That's kind of a tell for AI because they have just so much money.”
“Is it just me, or has there been an uptick in the number of AutoBuy orders since the tail end of last year?” one bookseller wrote on the forums for Alibris, another marketplace for selling books, in February. The AutoBuy function allows a customer to flag books they want to automatically purchase once they become available for sale on Alibris. “Any comment on what is happening? Is an AI going to read every single book? Any insight into how the selections are made? They seem to vary quite a bit in condition, format (hardcover and softcover), price and so on.”
“We have a couple of new bulk buyers that are scooping up trade books so lots of sellers are getting lots of orders,” Mike Feldman, director of client services at Alibris, responded.
One bookseller told me that similarly large orders of books were coming through another marketplace called Biblio. Customers can provide Biblio with a spreadsheet of ISBNs they want to purchase and the company takes it from there.
In June, a publication in the Netherlands talked to several rare booksellers who reported similar large bulk purchases they assumed were coming from AI companies.
It’s hard to say for a fact that the books are being bought for training data and possibly being destroyed by AI companies because ISBNdb and book marketplaces like Biblio and Alibris keep the identity of the buyer hidden. Large bulk purchases of books are first sent to distribution centers where, for example, Alibris checks the quality of the books before sending them off to the client.
Internal Anthropic documents about its plan to scan millions of books, revealed in the copyright lawsuit, don’t make clear why the company wanted to destroy the books in the process. A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both “high volume destructive and non-destructive book scanning” services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.
Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropic’s creation of digital copies of the books was legal specifically because the books were destroyed.
“Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,” Alsup wrote in his ruling. “The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
This, Alsup said, was “clearly transformative” and therefore qualified as fair use under Section 107 of the Copyright Act.
ISBNdb’s site advertises this legal argument to AI companies as well.
“Purchasing paper books at scale from the secondary market does not deprive any creator of income they would otherwise have received,” ISBNdb’s site says. “These are books that have already fully discharged their financial obligation to their creators.”
“Responsible physical sourcing is not book burning,” ISBNdb’s site in a section about why it’s crucial to recycle the destroyed books. “It is the completion of a book’s lifecycle: from tree to knowledge to tree again.”
ISBNdb and Anthropic did not respond to a request for comment.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み