Claude の不可視透かし回避法が即座に開発者に発見される
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
WIRED AI
Anthropic が EU の AI 法遵守のため Claude に不視認透かしを導入した発表からわずか数時間で、開発者らが回避ツールを公開し技術的対抗が即座に始まったことを示す。
AI深層分析を開く2026年8月20日 02:51
AI深層分析
キーポイント
即時の技術的対応
Anthropic の透かし導入発表から4時間以内に開発者の Guillaume Meyer が回避コードを公開し、GitHub や X で急速に拡散した。
コミュニティの活発な参加
Meyer のコードは2万回以上のブックマークと100人以上の貢献者を獲得し、多くの開発者が自身のプロジェクトへ組み込み始めた。
EU 規制が背景にある導入
Claude への透かし埋め込みは European Union の AI Act(AI 法)への準拠を目的として先週発表されたものである。
水回避への動機と規制の現状
開発者らはAI生成コンテンツのラベル付けに反対するか、技術的挑戦を好む理由から水回避を試みている。新規則ではプロバイダーに検知可能なラベル付けが義務付けられているが、独立したツールの利用には法的制限がない。
水マークの欠点とリスク
著者は水マークが誤検出やAI使用度の区別不能といった重大なリスクを伴うため不適切だと指摘する。この技術はClaudeの出力に影響を与える可能性があるが、Anthropicは品質低下はないと主張している。
重要な引用
"Anthropic is embedding watermarks in its Claude texts … the issue is practically history just one day later."
developer Guillaume Meyer had published his override.
His code to remove watermarks from Claude-generated text has since gone viral on GitHub
"I just think watermarking in itself is a really bad solution, because it has major drawbacks and risks."
編集コメントを表示
編集コメント
規制遵守を目的とした技術的変更が、開発者コミュニティによって即座に無効化される様子は、AI ガバナンスの現実的な課題を示している。企業は規制対応と技術的有効性のバランスを再考する必要があるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Anthropic が生成したコンテンツに機械可読の不可視透かしを全世界で埋め込むと公式発表してからわずか 4 時間後、開発者のギヨーム・マイヤーはそれを無効化するコードを発表しました。
Claude 生成テキストから透かしを除去する彼のコードは GitHub でバイラルになり、X(旧 Twitter)では 2 万回以上ブックマークされ、100 人以上のコントリビューターが集まりました。さらに多くの開発者がこの技術を自らのプロジェクトに取り入れています。「Anthropic は Claude のテキストに透かしを埋め込んでいますが、その問題は発表からわずか 1 日でほぼ過去のものになりました」とある AI スペシャリストは書き込みました。彼の記事には、鎖を断ち切り、崩れかけた EU と Anthropic の旗の上に立つマイヤーの姿が添えられています。
Anthropic は先週、欧州連合(EU)の AI 法への準拠のため Claude に透かし機能を導入すると発表しました。これを受けてマイヤー氏らは一足早く、透かし技術がどのように機能しているのか調査を始めています。
マイヤー氏は WIRED に対し、一部のユーザーが AI 生成コンテンツにラベル付けすること自体に反対しているため、透かし回避を試みていると語った。一方で彼自身を含む他の人々は、単に技術的な挑戦を楽しんでいるだけだと述べている。フリーランスのコンテンツライターやソーシャルメディアクリエイターからも、マイヤー氏に対してコードを使った支援を依頼する声が寄せられているという。
今月初めに施行された新規則では、Anthropic や OpenAI などのモデルプロバイダーに対し、生成された音声・画像・動画・テキストにラベル付けし、機械が AI 生成であることを検出できるようにすることを義務付けている。これに従わない場合、年間売上の最大 3% に相当する罰金が科される可能性がある。規則では、回避ツールの販売を禁止しているものの、独立したツール自体の利用には法的な制限はない。
「透明性には反対しませんし、コンテンツの帰属も支持しています」とマイヤー氏は語ります。「しかし、ウォーターマーキング自体が非常に悪い解決策だと考えています。なぜなら、重大な欠点とリスクを伴うからです。」彼は誤検知のリスクや、AI の使用量が軽微か重篤かを区別できない可能性を懸念しています。特にネイティブのフランス語話者であるマイヤー氏は、Claude や Grammarly などの他の AI ツール を頻繁に使用して文章を編集しています。Anthropic 社自身も、この技術がテキストが Claude によって改変された確率を示すだけだと認めているにもかかわらず、これを証拠として用いることは、採用担当者が候補者を不当に拒否したり、検出器が反応しただけで研究者が AI を使用したと過剰な非難を受けたりする事態を招く恐れがあるとマイヤー氏は指摘しています。
Anthropic は、Claude の単語やフレーズの選択に特定のパターンを埋め込むことで、テキストを目に見えない形でウォーターマークしています。このパターンは人間には検出できませんが、適切な手法を知った機械であれば識別可能です。この技術により Claude の出力に影響が出るため、一部のユーザーは回答の質が低下するのではないかと懸念していますが、Anthropic はその心配はないと主張しています。
この手法「SynthID」は Google が開発したもので、同社は 2023 年から AI 生成コンテンツにこの技術を用いてウォーターマークを施し続けています。コンピューターサイエンティストのスコット・アロンソン氏は OpenAI 在籍時に類似の方法を提案しましたが、その際「顧客が製品に対して抵抗感を抱くのを懸念したため、同社は実際に導入しなかった」と述べています。
マイヤー氏の除去手法は、ウォーターマークを付与しない大規模言語モデルを用いて複数の書き換えを生成し、同義語の置換や内容の再構成を行います。ただし、この手法は他社の非ウォーターマーク付き大規模言語モデルに依存しており、必ずしも安全な策とは限りません。なぜなら、OpenAI、Microsoft、Meta などの主要プロバイダーを含む 190 の組織が EU の AI 生成コンテンツ透明性コード実践に署名しているからです。今後、これらの機関のどれほどが実際にウォーターマークを実装するかは不透明ですが、8 月以降にリリースされる新モデルには必須となり、既存モデルへの統合も 12 月までに完了させる必要があります。
Anthropic がウォーターマークを検出するソフトウェアを公開するまでこのツールの有効性が確実視されるわけではありませんが、Claude のウォーターマーク化の基盤となっている基本的な SynthID-text のアプローチを理解していれば、マイヤー氏の手法が機能することは十分推測できると、シリコンバレーに拠点を置く主権 AI スタートアップ「Haimaker」の共同創設者兼最高技術責任者のウェイン・パン氏は述べています。同氏は、Claude が軽微な編集であってもコンテンツにウォーターマークを付与することや、ユーザーには見えない形で埋め込まれること自体に反対しており、マイヤー氏のオープンソースツールを自社のプラットフォームに取り入れました。
他の開発者たちも独自の除去ツールを開発しています。ソフトウェアエンジニアのエリック・ヒューズ氏は、Claude を活用してわずか 15 分で ツール を作成しました。このツールは、目に見えない文字や類似の文字を削除し、段落内の文の順序を入れ替え、複数の単語を同義語に置き換える機能を備えています。
オックスフォード大学の客員研究員であるレオン・クロン氏は、Claude の回答を要約し、英語とは意味体系が全く異なるアラビア語などの方言に翻訳した上で、再度元の言語に戻すことで、ウォーターマークを除去できると指摘しています。Anthropic 社自身も 認めています が、大幅な編集や要約、翻訳が施されたコンテンツにはウォーターマークが含まれない場合があると。
WIRED へのコメントで Anthropic の広報担当者は、「EU AI 法(欧州連合人工知能法)への対応として、Claude の出力にマーキングを追加しています。他の研究機関も同様の措置を講じています。AI 生成テキストの特定は困難ですが、この仕組みにより人々が識別するためのより良いツールを提供できます。Claude Code を含む対応モデルからのテキストには目に見えないウォーターマークが埋め込まれますが、これは回答の意味や品質、可読性を変えるものではありません。また、ユーザー自身がこうした処理をより多く行えるよう、テキスト検出 API の提供も計画しています。」と述べています。
Anthropic はテキストに対するウォーターマーク検出の実装方法を検討中で、近い将来にそのツールを公開する予定だ。これにより開発者は、自らの手法が本当に抜け目ないものかどうかを確認できるようになる。同社はまた、ウォーターマークシステムの改善にも取り組み続けている。「彼らは誠意を持って取り組んでいることを示したかったのだろう」とパンは語るが、「どんな状況でも耐えうる完璧なウォーターマークなど存在しないと思う」
原文を表示
Within four hours of Anthropic confirming that Claude models would globally embed invisible, machine-readable watermarks into any AI-generated content, developer Guillaume Meyer had published his override.
His code to remove watermarks from Claude-generated text has since gone viral on GitHub, has been bookmarked more than 20,000 times on X, and has drawn more than 100 contributors, with many more incorporating the technology into their own projects. “Anthropic is embedding watermarks in its Claude texts … the issue is practically history just one day later,” wrote one AI specialist, accompanied by an image of Meyer breaking out of chains and standing on crumpled EU and Anthropic flags.
Meyer and others started investigating how watermarking works after Anthropic announced last week that Claude would adopt it in order to comply with the European Union’s AI Act.
Some are trying to evade the watermarking because they disagree with the idea that all AI-generated content should be labeled as such, Meyer told WIRED, while others, including himself, say they simply relish the technical challenge. Freelance content writers and social media creators have also contacted Meyer asking for assistance using the code, he says.
The new rules, which came in earlier this month, stipulate that model providers like Anthropic and OpenAI must label synthetic audio, image, video, or text so that this material can be detected by a machine as AI-generated—or face fines of up to 3 percent of annual turnover. While the rules say providers cannot market circumvention tools, there is no legal restriction on independent tools.
“I'm not against transparency, and I'm all for content attribution,” says Meyer. “I just think watermarking in itself is a really bad solution, because it has major drawbacks and risks.” He is concerned about the risk of false positives and that the watermarking might not distinguish between light or heavy AI use, especially since, as a native French speaker, he often uses Claude and other AI tools like Grammarly to edit his writing. Using the watermark as evidence–when even Anthropic admits it can only generate a probability that the text has been touched by Claude–could lead to employers unfairly rejecting candidates or overblown accusations of researchers using artificial intelligence just because the detector flags it, he says.
Anthropic watermarks text invisibly by leaving a pattern in Claude’s choice of words and phrases that is indiscernible to a human reader but would be detectable by a machine that knows how to look for it. Because this influences Claude’s output, some users are concerned this will degrade the quality of Claude’s responses, though Anthtropic insists this won’t be the case. The technique, called SynthID, was developed by Google, which has been using it to watermark its AI-generated content since 2023. Computer scientist Scott Aaronson proposed a similar method when working at OpenAI but says the firm never deployed it because the company was worried that watermarks would put customers off its product.
Meyer’s removal method uses a non-watermarking large language model to generate multiple rewrites, swapping in synonyms and slightly reorganizing content. Of course, this relies on using other large language models which do not insert watermarks—possibly not a safe bet since 190 organizations—providers OpenAI, Microsoft, and Meta among them—have signed the EU’s transparency code of practice. It remains to be seen how many of these laboratories are going to implement their watermarks, which must be included in all new models released from August and must be integrated into existing models by December.
While there’s no certainty this tool works until Anthropic releases the software it uses to detect a watermark, understanding the basic SynthID-text approach underpinning Claude’s watermarking makes them fairly sure the method works, says Wayne Pan, chief technology and cofounder at Silicon Valley–based sovereign AI startup Haimaker. He incorporated Meyer’s open-source tool into his platform because he similarly disliked the idea of Claude watermarking content even when it’s only been lightly edited and disagreed with the watermark being invisible to the user.
Other coders have developed their own removal tools: Software engineer Erik Hughes took 15 minutes to knock up a tool with Claude that removes invisible and look-alike characters, reorders sentences within paragraphs, and swaps several words for synonyms. Leon Chlon, a Visiting Fellow at the University of Oxford, says the watermarks can be removed by condensing Claude’s response, translating it into a dialect like Arabic, which has very different semantics compared to English, and then translating it back. Anthropic itself acknowledged that heavily edited, paraphrased, or translated content might not carry a watermark.
In a statement to WIRED, a spokesperson for Anthropic said: "We're adding marking to Claude's output to comply with the EU AI Act, and other labs are taking similar steps. It’s hard to identify AI-generated text, and this gives people better tools for identification. Text from supported Claude models, including output from Claude Code, will carry an invisible watermark, and it doesn't change the meaning, quality, or readability of Claude's responses. We also plan to ship a text-detection API so users can do more of this themselves.”
Anthropic says it’s working out how to implement watermark detection for text and plans to release a tool to do so soon—at which point developers will finally be able to see whether their methods are foolproof. It’s also continuing to work on improving the watermarking system. “I think they wanted to show that they're in good faith doing it,” says Pan, “but I don't think you can ever have a watermark that will withstand everything.”
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み