Google の AI が「Google」や他の単語のスペルも間違える理由
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TechCrunch AI
TechCrunch は、Google の生成 AI モデルが自社の社名や一般的な単語のスペルを誤る現象について分析し、その技術的・データ上の原因を解説している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Google の名前に P はいくつあるか?Google によると、2 つある。
また Google の AI オーバービューは、「poop」という単語には「r」がちょうど 1 つあり、「journalism」という単語には「d」が 2 つあると述べているが、実際には j-o-u-r-n-a-d-i-s-m と誤って綴っている。Google は少なくとも米国の大統領の姓に P が 1 つあることは特定したが、それを t-r-p-u-m と誤って綴った。
Google の AI 重視型検索の大規模改修がうまくいかないことは、予言者でなくても予測できたはずだ。私たちは以前にも同じことを経験している。初めて Google が検索に AI オーバービューを追加した際、この機能は The Onion や Reddit の風刺記事 を引用し、人々に岩を食べたりピザに接着剤を塗ったりするようアドバイスしたりした。
今回は、Google が 29 年続く主力製品である検索の中心に生成 AI を据えるというコミットメントをさらに強化している最中であり、それがつまずくのは驚きではない。
「単語内の文字数カウントは LLM(大規模言語モデル)において既知の課題であり、私たちはこの特定の課題の解決に取り組んでいる」と Google は TechCrunch への電子メール声明で述べた。
これらの基本的なスペルミスは、どこか懐かしいように思えるかもしれません。チャットボットやその他のテキスト生成器を動かす人工知能の一種である大規模言語モデル(LLM)は、スペルを理解するように作られていません。企業が新しい AI モデルを発表するたびに、「strawberry」という単語に「r」が何個あるか聞いてみるべきだというのは、長年にわたる冗談になっています。これらの AI モデルは数秒でアプリをコーディングしたり、数十年間数学者たちを悩ませてきた問題を解決したりできるのに、スペルに関しては幼稚園児並みのレベルなのです。
しかし、Google の AI 概要に関する問題点は、おかしなスペルミスだけにとどまりません。Google は先週、"disregard"という単語を検索すると、あたかもその単語の辞書定義が示されるかのような結果が表示されるという問題を修正済みです。ただし、表示された定義は「了解しました。新しいプロンプトや質問があればいつでも教えてください!」というものでした。しかし、これらのスペルミスが笑いを誘い続けるのは、それらを根絶することが極めて困難だからです。
研究者たちは 以前に説明した 通り、これらの綴りの難問について質問すると、AI は文章を単語や文字で構成された言語の単位として認識していません。多くの大規模言語モデル(LLM: Large Language Model)はトランスフォーマーモデルに基づいて構築されており、このモデルはテキストをトークンに分解します。トークンは、モデルによって異なりますが、完全な単語、音節、あるいは文字そのものになり得ます。AI は人間のように「読む」のではなく、テキストを数値表現に変換し、それを文脈化することで論理的な回答を導き出そうとします。
image画像クレジット: TechCrunch
「LLM はこのトランスフォーマーアーキテクチャに基づいていますが、これは文字通りテキストを読んでいるわけではありません。プロンプトを入力すると、それはエンコーディングに変換されます」と、アルバータ大学の AI 研究者兼准教授であるマシュー・グズディアル氏は TechCrunch に語りました。「『the』という単語を見ると、『the』の意味に関する一つのエンコーディングを持っていますが、'T' や 'H'、'E' については知りません」。
Google の AI オーバービューのような LLM を支えるトークンベースのアーキテクチャは本質的に制限があり、研究者たちは綴りの問題を解決できることに対して楽観的ではありませんでした。
「言語モデルにとって『単語』とは具体的に何を指すべきかという問いを避けるのは難しいし、仮に人間のエキスパートが完璧なトークン語彙で合意したとしても、モデルはさらに細かく『チャンク(分割)』する必要があると考えるだろう」と、ノースイースタン大学で大規模言語モデルの解釈可能性を研究している博士課程学生である Sheridan Feucht は TechCrunch に語った。「この種の曖昧さゆえに、完璧なトークナイザーなど存在しないというのが私の推測だ」。
これは必ずしも研究者たちの頭を悩ませる緊急の問題ではない。なぜなら大規模言語モデル(LLM)の有用性は、スペル能力にあるわけではないからだ。しかし、これらの明白な失敗は、AI が時に見た目には理解を超えた全知全能の力のように思えるとしても、決して完璧ではないことを私たちに思い出させてくれる。その正確性を二重に確認することなく、AI の出力を盲目的に信頼してはならない。
*当記事内のリンクを通じて購入した場合、私たちは少額のコミッションを獲得する可能性があります。これは私たちの編集の独立性には影響しません。*
アマンダへの連絡や、彼女からの outreach の確認は、amanda@techcrunch.com へメールを送るか、Signal で暗号化メッセージを @amanda.100 宛てに送ることで可能です。
原文を表示
How many Ps are in Google? According to Google, there are two.
There’s also is also “exactly 1 ‘r’ in the word ‘poop’,” Google’s AI Overview says, as well as two ‘d’s in the word journalism, yet spelled it: j-o-u-r-n-a-d-i-s-m. Google did at least identify that there is one P in the last name of the U.S. president, but spelled it as t-r-p-u-m.
You didn’t need to be a prophet to predict that Google’s AI-forward Search overhaul was going to go over poorly. We’ve done this before. The first time Google added AI Overviews to Search, the feature ended up citing satirical posts from The Onion and Reddit, advising people to eat rocks and put glue on their pizza.
This time around, as Google doubles down on its commitment to make generative AI the centerpiece of its 29-year-old flagship product, it’s not surprising to see it stumble.
“Counting within words has been a known challenge for LLMs, and we’re working to fix this particular issue,” Google told TechCrunch in an emailed statement.
These basic spelling errors may seem familiar. LLMs, the kind of artificial intelligence that powers chatbots and other text-generators, are not built to understand spelling. It’s been a running joke for years that whenever a company unveils a new AI model, you should ask it how many ‘r’s are in the word strawberry. These AI models — which can code an app in seconds, or solve problems that have stumped mathematicians for decades — are about as good as a kindergartener at spelling.
Google’s AI overview woes reach beyond silly spelling mistakes though. Google already patched an issue from last week in which searching the word “disregard” would yield what looked like a dictionary definition of the word, only the definition was shown as, “Understood. Let me know whenever you have a new prompt or question!” But these spelling errors have remained amusing because they’re so difficult to quash.
As researchers have previously explained when we’ve asked about these spelling conundrums, AI doesn’t perceive sentences as units of language made up of words and letters. Many LLMs are built on transformers models, which break down text into tokens, which can be full words, syllables, or letters, depending on the model. Instead of “reading” like a human would, the AI converts the text into numerical representations of itself, which are then contextualized to help the AI come up with a logical response.

“LLMs are based on this transformer architecture, which notably is not actually reading text. What happens when you input a prompt is that it’s translated into an encoding,” Matthew Guzdial, an AI researcher and assistant professor at the University of Alberta, told TechCrunch. “When it sees the word ‘the,’ it has this one encoding of what ‘the’ means, but it does not know about ‘T,’ ‘H,’ ‘E.’”
The token-based architecture that powers LLMs like Google’s AI overview is inherently limiting, and researchers haven’t been optimistic that they can solve the spelling problem.
“It’s kind of hard to get around the question of what exactly a ‘word’ should be for a language model, and even if we got human experts to agree on a perfect token vocabulary, models would probably still find it useful to ‘chunk’ things even further,” Sheridan Feucht, a PhD student studying large language model interpretability at Northeastern University, told TechCrunch. “My guess would be that there’s no such thing as a perfect tokenizer due to this kind of fuzziness.”
This isn’t necessarily an urgent problem on researchers’ minds, since the utility of LLMs doesn’t come in their capacity to spell. But these blatant failures help us remember that AI is not perfect, even if it may sometimes seem like an all-knowing power beyond our comprehension. We cannot blindly trust AI outputs without double-checking their accuracy.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Amanda Silberling is a senior writer at TechCrunch covering the intersection of technology and culture. She has also written for publications like Polygon, MTV, the Kenyon Review, NPR, and Business Insider. She is the co-host of Wow If True, a podcast about internet culture, with science fiction author Isabel J. Kim. Prior to joining TechCrunch, she worked as a grassroots organizer, museum educator, and film festival coordinator. She holds a B.A. in English from the University of Pennsylvania and served as a Princeton in Asia Fellow in Laos.
You can contact or verify outreach from Amanda by emailing amanda@techcrunch.com or via encrypted message at @amanda.100 on Signal.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み