Anthropic、Claude のテキストに付与するウォーターマークの詳細を公開
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TechCrunch AI
Anthropic は EU AI 法への対応として Claude の生成テキストにシンセティック ID を採用する方針を明らかにし、編集による完全な除去は困難だが軽微な改変では検知される仕組みと、第三者向け検出 API の提供を発表した。
AI深層分析を開く2026年8月16日 04:16
AI深層分析
キーポイント
EU AI 法への対応としての導入
同社は EU AI 法の透明性コード遵守のため、AI 生成コンテンツの識別を可能にするシステムの実装を決定し、これに対するユーザーの懸念や解約の動きについても言及している。
シンセティック ID(SynthID-Text)技術の採用
Google DeepMind が 2024 年に概説した手法を採用し、読者には判別できないが鍵を持つ者は検出可能なパターンを生成する仕組みを導入すると発表した。
編集による回避の限界と検出 API
軽微な編集では水印は完全には除去されないと示唆しつつ、完全書き換えの場合は除去可能であり、同社は検出用の API をリリースする計画である。
既存の AI 検出手法との明確な区別
文章の構造的な「手がかり」を探す Pangram などの従来型検出とは異なり、水印は根本的に異なるアプローチであり、出力品質への影響はないと強調している。
軽微な編集や校正されたテキストの扱い
Claudeによる軽微な編集や校正が施されたテキストでは、人間の執筆部分がほとんどを占めるため、ウォーターマークが検出される可能性は極めて低い。
重要な引用
"Watermarking does not impact the quality of Claude's output. To a reader, a watermarked response is indistinguishable from an unwatermarked one."
"light editing probably won't remove the watermark completely," while "a complete rewrite where every word is replaced will."
"In the latter case, of course, it's arguable whether the text can any longer be described as AI-generated"
"Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code. But by definition, it will have a negligible effect on the actual code produced."
編集コメントを表示
編集コメント
生成 AI の普及に伴い、コンテンツの真偽を証明する技術的・法的枠組みが急速に整備されつつある。今回の Anthropic の発表は、単なる技術実装を超え、AI ガバナンスの実践例として注目されるべき動きである。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Anthropic は金曜日、チャットボット「Claude」が生成するテキストへの透かし(ウォーターマーク)の仕組みについて、基本的な疑問に答えるためのブログ記事を公開しました。具体的には、「透かしの仕組みは実際にどうなるのか」「編集で隠せるのか」「コード生成にはどのような影響があるのか」といった点です。
同社が今週初めに、EU AI 法(欧州連合人工知能法)の透明性規定への対応として、AI が生成したコンテンツを識別可能なシステムを導入すると発表した際、Claude の利用者たちはこの方針について議論を始めていました。この透かし導入は、AI 企業が生成されたコンテンツを特定できる仕組みを使用することを義務付ける EU AI 法の透明性コード(Transparency Code)に準拠するためです。
例えば Reddit では、ある投稿者がこれを「無実の Claude ユーザーに対する陰謀だ」と表現し、別のユーザーは「これに反対する唯一の理由は、人々を欺くためだろう」と主張しました。また Business Insider の報道によると、X(旧 Twitter)では「数十名」のユーザーが、Anthropic の新しい AI 透かし導入を理由に Claude のサブスクリプションを解約したと述べています。
Anthropic の新ブログ記事では、まずウォーターマーキングの概念について概説されています。具体的には、「低リスクな選択」——例えば天候を説明する際に「曇り」と「灰色」のどちらを使うかを選ぶ場合など——において、Claude はその回答にパターンを生み出すと説明しています。このパターンは読者には検出できませんが、それを符号化する鍵を持つ人であれば検出可能です。
同社は「ウォーターマーキングは Claude の出力品質に影響しません」と述べています。「読者にとって、ウォーターマーク付きの回答とそうでない回答を見分けることは不可能です」。
より具体的には、Anthropic は 2024 年に Google DeepMind チームが発表した SynthID-Text アプローチを採用し、さらにウォーターマーク検出 API の公開も計画しています。また、同社は AI 生成コンテンツの検出を目的とする Pangram などの企業が提供する手法とは明確に区別して説明しました。これらの企業は、「彼の発言は [X] ではなく、[Y] だ」といった文章構成のような「兆候」を検知することで AI の使用を明らかにしようとしますが、Anthropic は「こうしたパターンを読み取ることは、根本的にウォーターマークの確認とは異なります」と強調しています。
では、誰かがテキストを書き換えてウォーターマークを隠すことは可能でしょうか? Anthropic によると、それは技術的には可能ですが、「軽微な編集ではウォーターマークが完全に消えることはありません。ただし、すべての単語を置き換える完全な書き換えであれば、ウォーターマークは除去されます」。
「後者のケースでは、もちろん、そのテキストをもはや AI 生成と見なせるかどうかは議論の余地がある」と同社は述べています。
また、Claude が校閲や編集のみを行ったテキストにウォーターマークが検出可能かについては、「テキストの長さや Claude の編集の度合いによる」と回答しました。軽微な編集であれば、「ほぼすべての単語」は人間著者が書いたものとなり、「ウォーターマークが埋め込む余地はほとんど(あるいは全く)ない」としています。
一方、コードには他のテキストよりもウォーターマークの影響が少ないと予想されます。モデルは動作するコードを作成する必要があり、同等に有効な選択肢から自由に選ぶことができないためです。
「ただし、コード内の特定の単語や用語の選択が任意となる領域では、ウォーターマークを利用できます。例えばコード内のコメントなどが該当します」と Anthropic は説明しています。「しかし定義上、実際に生成されるコードへの影響は極めて小さいものとなります」
Anthropic はまた、Claude だけがウォーターマーク付きテキストを生成する AI チャットボットではないとも明言しました。「他の主要なモデル開発者も同様の行動規範に署名しており、それぞれ独自のウォーターマークを実装していく予定だからです」
*当記事内のリンクを通じて購入された場合、私たちは少額のコミッションを受け取る可能性があります。これは当社の編集の独立性には影響しません。*
アンソニーへの連絡や、彼からのアウトリーチの検証は、anthony.ha@techcrunch.com までメールでご連絡ください。
原文を表示
Anthropic published a blog post Friday seeking to answer some basic questions about how it will watermark the text generated by its chatbot Claude. Such as: How will the watermarking actually work? Can it be hidden with editing? And how does this affect code?
Claude users have been debating the move since the company revealed earlier this week that it would be doing this watermarking to comply with the EU AI Act’s Transparency Code, which requires AI companies to use systems that make it possible to identify AI-generated content.
On Reddit, for example, one poster characterized this as a conspiracy against innocent Claude users, while another claimed, “The only reason you wouldn’t want this is to lie to people.” And Business Insider reports that “dozens” of users on X have claimed to cancel their Claude subscriptions as a result.
Anthropic’s new post starts with a general overview of the watermarking concept, explaining that when making “low-stakes choices” — like choosing between the words “overcast” and “grey” to describe the weather — Claude can create a pattern in its responses that is “undetectable to the reader, but *is* detectable to anyone who has a key that encodes it.”
“Watermarking does not impact the quality of Claude’s output,” the company said. “To a reader, a watermarked response is indistinguishable from an unwatermarked one.”
More specifically, Anthropic said it will be using the SynthID-Text approach that the Google DeepMind team outlined in 2024, and that it plans to release a watermark detection API. It also noted that watermarking is distinct from the AI detection approaches offered by companies like Pangram that look for “tells” in the writing (like the construction “his isn’t [X], it’s [Y]”) to reveal AI usage: “Picking up on these patterns is fundamentally different from checking for a watermark.”
Could someone just rewrite the text to hide the watermark? Anthropic said it’s possible, but “light editing probably won’t remove the watermark completely,” while “a complete rewrite where every word is replaced will.”
“In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated,” the company said.
As for whether the watermark will be detectable in text that was only proofread or edited by Claude, Anthropic said that will depend on “the length of the text and how heavily Claude has edited it.” If it’s only been lightly edited, “nearly all the words” will have been written by the human author and “there’s very little (if anything) for the watermark to attach to.”
Code, meanwhile, should have less of a watermark than other text, because the model will need to create working code and won’t have the freedom to choose between a variety of equally valid options.
“Having said that, in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code,” Anthropic said. “But by definition, it will have a negligible effect on the actual code produced.”
Anthropic also said that Claude won’t be the only AI chatbot to generate watermarked text, as “other major model developers have signed the same Code of Practice and will be implementing their own watermarks.”
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Anthony Ha is TechCrunch’s weekend editor. Previously, he worked as a tech reporter at Adweek, a senior editor at VentureBeat, a local government reporter at the Hollister Free Lance, and vice president of content at a VC firm. He lives in New York City.
You can contact or verify outreach from Anthony by emailing anthony.ha@techcrunch.com.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み