Anthropic のテキスト透かし、AI 企業の文章への関心欠如を証明
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
404 Media
Anthropic は Claude の生成テキストに埋め込まれるウォーターマークについて、隠し文字やメタデータの追加ではなく、アルゴリズムがランダム性の源を変更することで単語選択を微妙に変える手法を採用すると発表した。
AI深層分析を開く2026年8月19日 01:12
AI深層分析
キーポイント
技術的実装の具体化
Anthropic はテキストに隠し文字やメタデータを追加せず、生成時のランダム性の源を変更して単語選択を微妙に変えることでウォーターマークを実現すると発表した。
読者への不可視性
同社によると、この手法により水付きテキストと非水付きテキストの違いは読者には判別できず、意味の伝達にも影響しないと説明している。
専門家の批判的視点
記事執筆者は、人間が文脈に応じて単語を選択する繊細さを無視し、アルゴリズムが単語を完全に互換性のあるものとして扱っている点に懸念を示している。
検出メカニズムの原理
この手法は低 stakes の選択(例:'grey'と'overcast')を多数積み重ねることでパターンを形成し、鍵を持つ者だけがそのパターンを検出可能にする。
水文字化による言語の価値低下
Anthropicの水文字化システムは言葉の選択をランダムなゲームとみなし、同義語の使い分けを無意味にすることで文章の質を損なう。
重要な引用
"Nothing is added to the text and there are no hidden characters," Anthropic wrote in that company blog post.
Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude's responses.
Anthropic’s algorithm sees these words as totally interchangeable and thus its watermarking algorithm has decided that it can "nudge" the word choice one way or the other for purposes of watermarking.
"Anthropic declares words fungible, language random, choice meaningless"
編集コメントを表示
編集コメント
生成 AI の透明性と検出可能性を高めるための技術的アプローチが、言語の微妙なニュアンスや人間の表現の自由とどう折り合いをつけるかが今後の課題となる。この手法は検出精度を高める一方で、AI による文章作成の本質的な価値に対する議論を再燃させる可能性がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

先月、Anthropic は Claude の将来バージョンにおいて、AI 生成であることを示すウォーターマークを含むテキストを生成すると発表しました。当時、Anthropic はその仕組みについて詳しく説明しておらず、ポッドキャストでは「どのようにテキストに埋め込むのか」「不可視文字を使うのか」「メタデータを利用するのか」など、様々な憶測が飛び交いました。
しかし今週末のブログ記事によって、その手法は AI の書き方そのものを変えるものであることが明らかになりました。
同社のブログ記事には、「テキストに何も追加せず、隠し文字も含まれない。ウォーターマーク付きと非付きのテキストを一般読者が区別することはできない」と記されています。仕組みについては、Anthropic が独自のアルゴリズムのみが知る方法で、AI 生成テキストの語彙選択を微妙に変えることで実現すると説明されています。Anthropic はウォーターマーク検出アルゴリズムを持っており、「単語を選択する際のランダム性の源」を変更することで、その文章が AI によって生成されたかどうかを検知するツールを開発できるのです。
この研究アプローチはデータサイエンスの観点から興味深いものですが、一般向けに説明した内容からは、同社がいかに執筆という技術や、人間が意図を伝えるために使い分ける言葉の微妙な違いについて軽視しているかが浮き彫りになっています。
Anthropic は、2 つの文の違いについて考えるよう求めています。「『今日の天気は寒かった』という文の次の単語が『甘い』になる可能性は極めて低いですが、『曇り』や『灰色』になる可能性は十分にあります。ほとんどの場合、モデルが最終的にどちらの言葉を選ぶかは読者にとって大きな問題ではありません。どちらを選んでも文の意味はほぼ同じだからです。こうした選択は、ランダムな数値によって決定されます」と Anthropic は記述しています。「ウォータマーキングはこのように、生成されたテキスト全体で何度も起こる低リスクの選択を利用し、Claude の回答にパターンを残します。このパターンは読者には検出できませんが、そのパターンを符号化する鍵を持つ人であれば検出可能です。ウォータマーキングを使用する場合も選択はランダムに行われますが、ランダム性の源泉が異なるのです。」
何かを書いた経験がある人なら、Anthropic が説明する通り、「今日の天気は寒くて灰色だった」と「今日の天気は寒くて曇りだった」の違いが常に低リスクであるとは限らないことを理解できるはずです。「灰色」と「曇り」は異なる言葉であり、文脈によっては人間がどちらかを選ぶ理由はいくらでもあります。しかしこの例では、Anthropic のアルゴリズムはこの 2 つの言葉を完全に互換性のあるものとして扱い、ウォータマーキングのために一方へ「押しやる」ことが可能だと判断しています。
Anthropic は続けています。「任意の乱数生成器を使って次の単語を選ぶのではなく、ウォーターマーキングでは鍵と直前の数語を用いて、モデルが選ぶべき単語を決定します。つまり、Claude が選ぶ単語自体は依然としてランダムですが、その単語列を検査することで、もし Claude が鍵を使用していた場合にどのような選択をするかという一貫性があるかどうかを確認できます。もし一致すれば、そのテキストが Claude によって生成された確率を算出できるのです。」
Anthropic は「ウォーターマーキングは Claude の出力品質に影響しません。読者にとって、ウォーターマーキング付きの回答とそうでない回答を見分けることは不可能です」と主張し、「内部テストでは、ウォーターマーキングがテキストの内容、創造性、読みやすさに何ら影響を与えていないことを確認しました」とも述べています。
しかし、Anthropic のウォーターマーキングシステムに対して人々は激しい怒りを示しています。その理由は十分理解できます。同義語は文脈によっては置き換え可能ですが、常にそうとは限りません。これは『Daring Fireball』のジョン・グラバーやジャーナリズム学者のジェフ・ジャ维斯が指摘した通りです。特にジャビスは、Anthropic のこの選択が「文章を軽視している」と主張しています。「Anthropic は単語を交換可能なものとし、言語をランダムなものと見なし、人間の選択を無意味なものだと宣言しているのです」と彼は記しています。
この投稿を書くために席に着いたとき、私はアンソロピックが不公平な操作をしていることに怒っていました。同社は自社の生成モデルの出力を意図的に操作し、その結果生じたテキストが、おそらく本来出力されるはずだった他の AI 生成テキストと質的に全く同じであるかのように主張しているからです。
しかし、執筆を進めるうちに気づきました。私の問題の本質は「テキストの透かし」そのものではなく、「AI によって生成されたテキスト」という存在全体にあるのです。Claude のゴミのような AI テキストがどういった形式で出力されるかは、私にとって本質的な問題ではありません。重要なのは、巨大テック企業のデータサイエンティストたちが「言葉選びなど大した意味を持たない」と考えたり、文の意味を損なうことなく統計的にどこでも同義語を置き換えられると信じているという事実です。
ブログ記事全体を通じて、アンソロピックは執筆行為を確率的なギャンブルに例えています。同社自身の言葉によれば、その生成プロセスは「任意の乱数発生器」の結果であることがあり、システムがパターン認識に基づき「次に来る単語の選択肢のうちどれを選んでも大差ない」と判断した状況では、「ランダム性」が優先されます。
これは LLM というゴミには当てはまるかもしれませんが、人間が文章を書く経験においては真実ではありません。だからこそ、人間の執筆と AI による執筆は、ほとんど常に異なる質感を呈するのです。
このウォーターマーキング手法と、それに関するアンソロピックのブログ記事は、すでに明らかであるべきことを浮き彫りにしています。つまり、印刷された膨大な数の書籍をスキャンして破壊し、盗まれたコンテンツで大規模言語モデル(LLM)を学習させたことで有名な同社が、実は文章の技術や努力に全く関心がないということです。彼らは言葉を互換性があり、重要ではないものとして捉えています。
アンソロピックは、この変更を欧州連合(EU)の新規 AI 規制の一環として実施すると述べていますが、これは善意に基づくものではあるものの問題を含んでいます。確かに、AI 生成コンテンツを検出する追加手段を持つことは有用ですが、アンソロピックがこの決定を発表した際の無頓着さは、LLM を使用して文章を作成することに伴うより広範な問題を浮き彫りにしています。同社が指摘するように、これらは確率的なツールであり、人間のように「書く」のではなく、すでに起こり、取り込まれた過去の出来事であるトレーニングデータを模倣するに過ぎません。
これとは対照的に、アンソロピックはコードに対しては「正確な出力が必要」と考えています。一方、文章については、異なる言葉がしばしば「同等に良い」と示唆しています。何度も繰り返して、アンソロピックやこの種のウォーターマーキングに取り組む研究者たちは、テキストを人間に気づかれず、かつ「品質」に影響を与えることなく、このように「微調整(nudged)」できると主張し続けています。
しかし、AI が生成した出力の「品質」を評価している人々は、データサイエンティストか、あるいは AI ツールに執筆を任せている人々であって、読書や執筆そのものを大切に思う人々ではないという点には注意が必要です。Anthropic が引用している科学論文は、Google の研究者が Google のウォーターマーキングツール「SynthID」について行ったものであり、Anthropic のウォーターマーキングはこの SynthID に基づいています。
SynthID の研究では、品質評価のためにランダムに一部の Gemini 出力にウォーターマークを付与し、Gemini ユーザーに対してその回答に「いいね(親指アップ)」か「いやだ(親指ダウン)」かを尋ねました。「クエリのランダムな割合をウォーターマーク付きモデルへルーティングし、同等の数を非ウォーターマーク版へ振り分けました。Gemini のユーザーインターフェースでは、ユーザーは親指アップ(良い回答)と親指ダウン(悪い回答)を通じてモデルへのフィードバックを提供できます。私たちは約 2,000 万件のウォーターマーク付きおよび非ウォーターマーク付き回答を分析し、親指アップ率と親指ダウン率(両方とも、受け取った総フィードバック数に対する割合として計算)を算出しました。その結果、両モデル間の親指アップ率は 0.01% の差しかありませんでした。」
この記事をクリックした人なら誰でも、チャットボットに質問した人物に対して回答を「いいね」や「よくない」と評価させることが、「文章の質」を測る適切な方法ではないことは明白でしょう。Google が行ったもう一つの人間による評価では、水マーク付きと水マークなしのテキストを並べて比較し、その質を評価してもらいました。研究の付録に掲載されている例を見ると、人々はどちらかを明確に好む傾向はなかったようです。
imageこれらの文章は、同じ内容を説明する異なる表現方法だと主張できるかもしれません。しかし、明らかに「同一」ではありませんし、なぜ片方がこの表現で、もう片方が別の表現なのか、読者にとってその理由は不明瞭です。なぜ LLM はある例では「呼吸不全」と書き、別の例では「呼吸停止」と書いたのでしょうか。両方の答えは、「任意の乱数生成器」と、AI 企業が制御する独自性のブラックボックス化されたアルゴリズムによる重み付けシステムにあります。水マーク付きバージョンでは、もともとパターンマッチングに過ぎない機械に対して、追加で「誘導」したり操作を加えたりしているのです。
つまり、ここで意識的な思考や意思決定が行われているわけではないので、AI によって生成されたテキストにウォーターマークが施されていること自体は、通常の AI テキストほど悪質ではないかもしれません。しかし、これらの機械を開発する企業がそれをこれほどまでに露骨な形で提示していることは、彼らが実際に「文章を書く」ことにどれほど無関心であるかを示しています。
では逆に、私がなぜある単語を選んだのかと問われれば、その理由を正確に説明できるかどうかはわかりません。しかし、私が目指していたもの、私の文体、対象読者、その日の気分、心が高鳴っていたかどうか、どこで何をしていて、朝のうちに何を過ごし、その後何をしていたかといった背景については説明できるでしょう。
もしかしたらそれは、小学 3 年生の担任教師がいつも使っていた言葉かもしれませんし、先週読んだ記事に出てきた言葉や、友人との間のジョーク、最近ハマっている言葉、あるいは多用しがちな言葉なのかもしれません。私がなぜそのように書いたのか、あるいはそもそもなぜ何かをするのかは、赤ちゃんとして言語を獲得した時から現在に至るまでの、私自身の経験の複雑な混ざり合いの結果です。それをすべて説明できるかどうかもわかりませんが、それが私独自の文章スタイルを生み出しているのです。
私は速く作業したり、不注意にメッセージを送ったりする際にも、このことは当てはまります。その時、思考が脳から指先へ、そしてキーボードへと流れ込みます。自分が何を言っているのか意味があるのかさえわからなくても、それは人間の脳から生まれたものであり、ランダムな数値生成器ではないため、おそらく可読性は保たれているのです。
だからこそ、AI が生成した短い文章は、繰り返し指摘してきた通り、魂の欠けたものや画一的なものに感じられるのです。また、より人間らしく、あるいは盗用されていないように見せるために、あまり使われない単語を選ぶことで AI による文章を「人間化」するツールも多数存在します(つまり、「AI が好きならさらに AI を使う」という皮肉な構造です)。しかし、「スピンナー」や「ヒューマナイザー」と呼ばれるこれらのツールが出力するテキストは、AI 自体の文章と同じくらい不気味で奇妙なものになりがちです。あるいは、ニュース記事の引用のように「正確な出力が必要」とされる場面でこれらを使用すると、事実誤認、名誉毀損、あるいは単なるゴミのような出力になることがよくあります。
原文を表示
imageEarlier this month, Anthropic announced that future versions of Claude will generate text that includes watermarks showing it was AI-generated. At the time, Anthropic did not explain how this would work, leaving us to speculate on the podcast: Would it somehow encode this into the text? Include invisible characters? Do something with the metadata? We now know, thanks to a blog post over the weekend, that Anthropic will do this by changing how its AI writes altogether.
“Nothing is added to the text and there are no hidden characters,” Anthropic wrote in that company blog post. “The difference between watermarked and un-watermarked text will not be distinguishable to readers.” The way it will work, the post explained, is that Anthropic will subtly alter the word choices in AI-generated text in a way that is only known to Anthropic and its algorithms. Anthropic will know the watermarking algorithm, which will change “the source of the randomness used to pick among words” and thus can write a tool to detect whether something has been AI-generated.
This research and approach is interesting in a data science kind of way, but Anthropic’s layperson explanation for how this will work shows how little the company thinks about the craft of writing or the subtle differences between words a human author might want to use to convey their thoughts.
Anthropic asks us to consider the difference between two sentences: “Take the sentence ‘The weather today was cold and…’. The next word is very unlikely to be ‘sugary.’ But it is quite likely to be ‘overcast’ or ‘grey.’ Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses—the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number,” Anthropic writes. “Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses. That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it. When watermarking is used, choices are still made at random, but the source of the randomness is different.”
Anyone who has written anything would, I hope, understand that the difference between the sentences “The weather today was cold and grey” and “The weather today was cold and overcast” are sometimes “low stakes,” as Anthropic describes, but not always. “Grey,” and “overcast” are different words, and there are any number of reasons why a human author might pick one over the other in a given context. In this example, however, Anthropic’s algorithm sees these words as totally interchangeable and thus its watermarking algorithm has decided that it can “nudge” the word choice one way or the other for the purposes of watermarking.
Anthropic continues: “Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key. If it is, one can assign a probability that the text was generated by Claude.”
Anthropic claims “Watermarking does not impact the quality of Claude’s output. To a reader, a watermarked response is indistinguishable from an unwatermarked one,” and that “in internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text.”
People are quite mad about Anthropic’s watermarking system, and understandably so. Synonyms are sometimes interchangeable, but not always, as is pointed out in this excellent essay by John Gruber of Daring Fireball, and by journalism academic Jeff Jarvis, in which he claims Anthropic “devalues writing.” In making this choice, “Anthropic declares words fungible, language random, choice meaningless,” Jarvis writes.
When I sat down to write this post, I was mad because it seems like Anthropic is putting its thumb on the scale, messing with the outputs of its machine and saying that the resulting text is qualitatively just the same as the other AI text it was probably going to output. But as I began writing this, I realized that my problem is not necessarily with text watermarking but with AI-generated text altogether. It does not matter to me, necessarily, whether the output of Claude’s garbage AI text is one way or is a slightly different way. But it does matter to me that AI data scientists at huge tech companies think that word choice doesn’t matter, or that it is possible to statistically use synonyms wherever without fucking with the meaning of a sentence.
Throughout the blog post, Anthropic describes the act of writing as being akin to a probabilistic game of chance. In Anthropic’s own words, its writing is sometimes the result of an “arbitrary random number generator,” and “random” whenever its systems encounter a situation where its tool believes, based on pattern recognition, that the choice between several possible next words isn’t all that important. That may be true for LLM garbage, but is not true for the human experience of writing, which is why human writing almost always feels different than AI writing.
This watermarking approach, and Anthropic’s blog post about it, highlights something that should already be clear about a company that famously scanned and destroyed huge numbers of printed books and has trained its LLMs on stolen content: Anthropic does not care about the craft or effort of writing, and sees words as fungible and unimportant. Anthropic says it is making this change as part of the European Union’s new AI regulations, which are well-intentioned but problematic. While it can definitely be useful to have additional ways of detecting AI-generated content, the carelessness with which Anthropic has announced this decision highlights the broader problem with using LLMs to write: They are, as Anthropic notes, probabilistic tools that do not “write” in the way that humans do, rather, they mimic their training data which is, by definition, things that have already happened and been ingested.
Contrast this with how Anthropic sees code, something where it says an “exact output is required.” In writing, meanwhile, Anthropic suggests different words are often “equally good.” Over and over again, Anthropic and the researchers who work on this type of watermarking claim that text can be “nudged” in this way without being noticeable to humans or without impacting “quality.”
But it is worth noting that the people judging the “quality” of the AI-generated outputs are either data scientists or people asking AI tools to do their writing for them, not, say, people who care about reading or writing. The scientific paper that Anthropic cites was done by Google researchers on a Google watermarking tool called “SynthID,” which Anthropic’s watermarking is based on.
In the SynthID study, quality was assessed by randomly putting watermarking on some Gemini outputs, then asking Gemini users to either thumbs-up or thumbs-down the response: “A random fraction of queries were routed to a watermarked model and an equivalent number to the unwatermarked counterpart. The Gemini user interface allows users to provide feedback on model responses via a thumbs-up (good response) and a thumbs-down (bad response). We analysed approximately 20 million watermarked and unwatermarked responses and computed the thumbs-up and thumbs-down rates (both as a fraction of the total number of thumbs-up and thumbs-down feedback received). We found that the thumbs-up rate for the two models differed by 0.01%.”
I hope it is clear to anyone who has clicked on this article that asking someone who asked a chatbot something to thumbs up or thumbs down a response is not a very good way of assessing the “quality” of “writing.” The other human assessment that Google did was to ask people to assess side-by-side watermarked and unwatermarked text for quality. Here are examples given in an appendix of the study; apparently people did not really have a preference one way or the other:
imageOne could argue that these passages are two different ways of explaining something, yes. But they are definitively not the “same,” and it is unclear to any reader why one version is one way and the other version is another way. Why did the LLM write “respiratory failure” in one example and “cessation of breathing” in the other? The answer for both is an “arbitrary random number generator” and proprietary black box algorithmic weighting systems controlled by the AI company. In the watermarked version there’s been an additional “nudging” or messing with the machine that’s already just a pattern matcher.
The point is, there is no conscious thought or decision-making process happening here, so perhaps watermarked AI text is not all that much more offensive than regular AI text. But to see it laid out in such stark terms by the companies building these machines shows how little they actually care about writing. If you asked me, on the other hand, why I used one word instead of another, I might not be able to tell you exactly why, but I could probably explain to you what I was going for, the style of writing I do, my intended audience, my mood that day, whether my heart was racing or not, where I was, what I was doing, what I did earlier that morning and what I did later that day. Maybe it was a word my third grade teacher used all the time or which I read in an article last week or is an inside joke with my friends or which I have recently become obsessed with or tend to overuse. Why I wrote what I wrote or why I did anything at all is the result of my some mix of human experiences dating back to when I first acquired language as a baby and continuing on to this very moment that I may or may not be able to explain, but which result in a certain style of writing that is mine.
This is the case even when I’m working fast or carelessly dashing off text messages, when the thoughts just kind of flow from my brain to my fingers to my keyboard where I don’t know if what I’m saying is making sense at all but is probably legible because it’s coming from a human brain and not a random number generator.
This is why short passages of AI-generated text feel soulless and generic, as we have written about repeatedly. And there are many AI tools that use AI to make AI writing seem less generic (yo dawg, we heard you like AI so we put AI in your AI) by using synonyms that are supposed to make a passage sound more human — or less plagiarized — by picking words that are less commonly used. The text outputted by these tools, which are called “spinners” or “humanizers” are often just as uncanny and weird as AI writing itself. Or, when applied to things where, to use Anthropic’s own language, “an exact output is required” such as quotes in a news article, the output is often factually inaccurate, libelous, or just plain garbage.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み