AI の知能を正しく理解しているか、サンタフェ研究所のメロニー・ミッチェルが警鐘
本文の状態
日本語全文を表示中
詳細モードで約54分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
サンタフェ研究所のメレニー・ミッチェルは、AI が人間のような推論を行っているのか単なる模倣に過ぎないかを判断する新たな測定方法と、6 つの評価原則を提案している。
AI深層分析を開く2026年8月22日 22:12
AI深層分析
キーポイント
機械認知の測定不足と異星人知性
メレニー・ミッチェルは、現在の技術では機械の認知能力を適切に評価する方法が欠如しており、AI は非人間型の認知機構を持つ「異星人知性」であると主張する。
心理学的手法の転用と6 つの原則
ミッチェルは、乳幼児や動物の認知を研究する心理学者が用いる手法を AI に適用し、機械認知をより良く評価するための 6 つの原則を提示した。
数学的推論と歴史的事例の教訓
AI による数学分野での最近の画期的な成果や、20 世紀初頭の計算馬が示す評価の罠について言及し、知能の評価における慎重さを求めている。
内部解釈の難しさ
これらのシステム内で何が起きているかを解釈することの困難さが議論され、ブラックボックス化された AI の挙動をどう理解するかが課題として挙げられた。
AI開発における人間の知能理解の欠如
私たちは人工知能に対して過度に興奮しているが、赤ちゃんや子供がどのようにして知的になるかという人間側のメカニズムについてはほとんど理解していない。
重要な引用
AI is a form of 'alien intelligence' that operates through non-human cognitive mechanisms.
The distinction isn't just philosophical, this determines what we can trust AI to do, how closely we need to supervise it, and ultimately what its real-world impact will turn out to be.
we're so excited about the artificial mind when we have very little comprehension of the human mind
we're trying to skip a step
編集コメントを表示
編集コメント
この議論は、AI の能力評価が単なる性能比較から、認知プロセスの理解へとパラダイムシフトする必要があることを示唆している。サンタフェ研究所のメレニー・ミッチェルによる提言は、業界全体が抱える「ブラックボックス」問題への重要な対案となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
LLM が質問に答えるとき、それは人間のように推論しているのか、それとも推論に見えるテキストを生成しているだけなのか。この区別は単なる哲学的な議論ではなく、AI に何を信頼できるか、どの程度厳密に監視する必要があるか、そして最終的に現実世界でどのような影響を与えるかを決定づけるものです。
サンタフェ研究所のメレニー・ミッチェル氏は、機械的認知を測定するための適切な方法が不足しており、AI は非人間的な認知メカニズムを通じて動作する「異星人のような知性」の一形態であると主張しています。この『The Joy of Why』のエピソードで、ミッチェル氏はスティーブン・ストロガツに対し、心理学者が他の種類の「異星人のような知性」、つまり赤ん坊や動物の認知を研究するために用いる手法を AI への探求に応用できることを説明し、機械的認知をよりよく評価するための6つの原則を提示しました。二人の対話は、これらのシステム内部で何が起きているかを解釈する難しさから、数学分野における最近の AI 支援による画期的な成果、そして1900年代初頭の計算ができる馬が知性を評価する方法に対する教訓となる物語である理由まで、多岐にわたります。
Apple Podcasts、Spotify、TuneIn、またはお気に入りのポッドキャストアプリで聴くことができます。また、Quanta Magazine からストリーミングすることも可能です。
transcript
[音楽が流れる]
スティーブ・ストロガッツ:私はスティーブ・ストロガッツです。
ジャンナ・レヴィン:そして私はジャンナ・レヴィンです。
ストロガッツ:これは『The Joy of Why(なぜという喜び)』です。
レヴィン:Quanta Magazine が提供するこのポッドキャストでは、現代の数学と科学における最大の未解決問題を探求します。
ストロガッツ:こんにちは。これもまた、AI に関する番組ですね。
レヴィン:言っておきますが、人々は AI という話題から離れられず、私ももはや断定的な意見を述べるのをためらうようになりました。状況があまりにも急速に変化しているからです。
ストロガッツ:その通りです。非常に速いペースで進んでいます。私たちが今話していることが、来週には古びてしまう可能性もあります。
レヴィン:そうですね。
ストロガッツ:さて、現在の日付は 2026 年 7 月 23 日です。
レヴィン:2025 年 7 月 23 日の頃とは、私には明らかに違う雰囲気を感じます。
STROGATZ: なるほど、そのタイムラインの話は非常に重要です。今日のゲストであるメアリー・ミッチェル氏は、サンタフェ研究所で認知科学者かつコンピュータサイエンティストとして活躍されています。彼女とは以前にもお話をしたことがあり、ちょうど5年前のことでした。ChatGPT が登場する前の話です。
LEVIN: 当時も彼女は AI に興味を持っていたのですか?
STROGATZ: もちろんです。
LEVIN: なるほど、認知科学だけにとどまらず、ということですね。
STROGATZ: その通りです。メアリーは長い間 AI について考え続けてきました。彼女自身からその経緯を伺うことになるでしょう。ただ、今日私たちが特に議論すべき面白い点は、メアリーの視点にあります。それは、発達心理学などの分野の立場から AI の問題を捉え直すという考え方です。例えば、赤ん坊や幼い子供が、なぜそれほどまでに短期間で知的な存在へと成長するのか?そのプロセスについて考えるのです。
LEVIN: それは非常に興味深いですね。私たちは人工的な知能に対して熱狂している一方で、人間の心そのものについてはほとんど理解できていないのですから。
STROGATZ: その通りです。
LEVIN: つまり、重要なステップを飛び越えようとしているわけですね。
STROGATZ: そうです。しかもそれは人間に限った話ではありません。動物の心についても同様です。比較心理学という分野では、鳥や犬、イルカなどの知能について研究しています。私たち成人人間の知能以外の、さまざまな知能について学ぶべきことはまだ山ほどあるのです。
レビン: はい、その点です。私たちは、生まれてきた赤ちゃんや成長過程にある子供がどのようにして知能のレベルを獲得するのかというメカニズムをまだ理解していないのに、人工知能を生み出すための仕組みを単に理解すればいいと考えているような考え方に疑問を感じます。この二つの視点を組み合わせるのは非常に興味深いと思います。私もその点について楽しみにしています。
ストロガッツ: 素晴らしいですね。では、メレニー・ミッチェルさんにお話を伺いましょう。こちらです。
[*音楽が流れる*]
ストロガッツ: こんにちは、メレニーさん。
メレニー・ミッチェル: やあ、スティーブ。
ストロガッツ: またお会いできてとても嬉しいです。今回は楽しい時間になるでしょうね。数年前にこの番組が『The Joy of X』と呼ばれていた頃にお話ししましたが、あなたが私たちの最初の「再来チャンピオン」かもしれませんよ。
ミッチェル: なんておっしゃるんですか、光栄です。
ストロガッツ: ぜひそう思ってください。私があなたをお招きしたのは、人工知能の分野でここ数年で大きく様変わりしたと感じられることが多すぎるからです。おそらく2021年にお話ししましたが、ChatGPTという津波が世界を襲ったのは2022年の11月頃だったと思います。合っていますか?
ミッチェル: はい、その通りです。
ストロガッツ: 今や誰もが、人工知能は至る所に存在していることを知っています。私たちはそれを話題にし、人々はそれについて心配したり、興奮したりしています。確かに非常に広く利用されていますが、まずお伺いしたいのは、ここ数年で最も驚いたことは何ですか?
ミッチェル: 驚いたことがたくさんあります。ただ、これらのモデルを膨大な量の人間が生成した言語や画像などでトレーニングするだけで、今の水準に到達できるとは思ってもみませんでした。AI の分野で何が起こったかには本当に驚いています。また、AI コミュニティ全体や社会一般で見られる、極端に分かれた反応にも少し驚きました。
ストロガッツ: 分かれているのは、例えば「AI に悲観的な人」と「AI に楽観的な人」を区別するような場合でしょうか?そのことをおっしゃっていますか?
ミッチェル: そういう側面もありますし、もう一つの側面として、「AI は人間より賢い」と信じる人と、「人間の知能に遠く及ばない」と考える人の対立があります。それに関連して「愛する派」と「嫌う派」の別れもあるでしょう。これらは独立した次元ですが、もしかすると相関があるのかもしれません。
ストロガッツ: そうですね。「愛する派」と「嫌う派」は、環境への影響と特定の企業の経済的繁栄といった対立軸にもつながりますし、一方で雇用喪失の問題もあります。このテーマには実に多くの側面がありますね。
ミッチェル: はい、本当に数えきれないほどあります。
STROGATZ: 今日、あなたと特に議論したいのは、複雑系や認知科学、そして人工知能(AI)です。あなたは多くの肩書きをお持ちですが、私はあなたがこれまでに取り組んできた研究に大変興味を持っています。それは、発達心理学の視点、つまり人間の子供のような異質な知性についてどう考えるか、あるいは比較心理学の視点、つまりペットの犬や賢い鳥、イルカなどの動物が持つ異質な知性の観点から AI を捉え直すというものです。AI という「異質な知性」をこのように捉える視点は非常に興味深いです。
MITCHELL: はい。多くの人が AI を「異質な知性」と表現しています。インターネット上の言語や書籍、あらゆるデータを学習対象としているにもかかわらず、その仕組みは人間とは大きく異なるからです。これらのシステムがどのように働き、どのように学び、どのように推論し、どのような行動をとるのかという点は、人間のそれとは根本的に異なります。
このテーマは、特にスタンフォード大学のマイク・フランク氏のような発達心理学者たちによって取り上げられました。彼は、AI 研究者たちは赤ちゃんや幼児の研究からヒントを得るべきだとする論文を執筆しています。また、他の人々はこれを動物の知性へと拡張し、「では動物の知性はどうだろうか」と問いかけています。おそらく、認知科学(Cog Sci)の分野の人々が強く主張しているのは、AI の研究がより本物の科学らしくなるための実験手法を採用すべきだという点でしょう。
ストロガッツ: この視点、とても共感します。おそらく聴衆にはあまり馴染みがないかもしれません。私自身もそうでした。認知科学を専攻したことも、発達心理学の授業を受けたこともないからです。これらの分野の人々は長年この問題について考えてきました。どれくらいでしょうか?ミッチェルさん、お答えください。
ミッチェル: 少なくとも100年です。
ストロガッツ: そうですか、100年も。驚きですね。さて、こちらに向かう間もずっと「AI はブラックボックスだ」と話していました。ニューロンの重みを読み取ることは容易ではなく、仮に読み取れたとしても、それが何を意味するのか分からないからです。しかし考えてみてください。私たちの知能もまた、多くの点でブラックボックスではないでしょうか。
ミッチェル: 確かにその通りです。ブラックボックスを解明する方法はいくつかあります。一つは神経科学で、実際にプローブをニューロンに挿入したり、fMRI(機能的磁気共鳴画像法)や他のイメージング技術を用いたりします。もう一つが心理学で、人間や動物の行動そのものを観察し、そこから背後にあるメカニズムを推測するアプローチです。
この二つの伝統は長い間、互いに独立した分野として存在してきました。しかし認知科学という分野が誕生し、両者を統合しようとしたのです。当初、認知科学には人工知能(AI)も含まれていました。ところが、その統合はうまくいかなかったようです。
ストロガッツ: 社会学的に定着しなかったという意味ですか?
ミッチェル:当初は、AI を人間と同じようにプログラムすればよいと考えられていました。人間の心理学と AI に人間の心理を組み込もうとする試みには密接なつながりがありました。しかし、そのアプローチは成功しませんでした。私たちが目にしたのは、神経ネットワークやデータからの学習であり、プログラムによる制御ではありません。
ストロガッツ:なるほど。
ミッチェル:また、神経ネットワーク自体は当初、神経科学にインスパイアされていましたが、現在の動作原理はその元々の着想から大きく乖離しています。つまり、機械学習の分野は統計学へと大きく舵を切っており、これは認知科学が働く仕組みとはかなり異なる方向性です。
STROGATZ: さて、この点でベンチマークについて少しお話ししたいです。現在、社会全体で議論の中心となっているのはまさにこれだからです。
数学の世界でも大きな話題を呼んだ出来事がありました。最新のフロンティアモデルが、まるで創造性を見せるかのようなことを成し遂げたのです。それは、ハンガリー出身の偉大な数学者ポール・エルデシュが残した難問の一つ、「単位距離問題」を解決したというものです。彼は多くの問題を後世に残しましたが、その中で最近、AI によって非常に賢明な方法で解かれるに至りました。この解決法は、これまで試されたことのない形で数学の二つの分野を組み合わせたものでした。
私がこれを取り上げたのは、前回お話しした際の話と対比させるためです。その時は、古いタイプの AI がアタリゲームなどをプレイするのを学習している話でした。あなたは、その AI はプレイが上手い一方で、パドルを数ピクセル動かすだけで、全く新しい状況に対応できず、最初からやり直しが必要だとおっしゃいました。オリジナルのゲームにわずかな変化を加えただけでも、それを理解してプレイすることができないのです。
当時のお話で私の中で強く印象に残ったのは、「これらの機械は、学習した領域以外にはその卓越性を転移させることができない」という点でした。あれから五年が経ちました。さて、今ではどうお考えですか?それは今でも真実なのでしょうか?
ミッチェル:はい、あの特定のモデルは大規模言語モデルではありませんでした。アタリゲームをプレイするための専用モデルだったのです。一方、現在の大規模言語モデルはあらゆるデータを対象に学習されています。ある意味では、新しい知識を転移させる必要すらありません。すでに学習済みだからです。
しかし、AI や機械学習の分野の人々は「分布内」と「分布外」という概念について語ります。これはつまり、「私たちがモデルに求めているタスクが、学習データで見たことのある類似した事柄なのか、それとも全く異なる事柄なのか」を問うものです。
それが何なのかを知ることは難しいと思います。何が学習されたのかは不明です。これらの問題を解決するモデルは、間違いなく大量の数学的データを学習しています。インターネット上には膨大な量の数学情報がありますし、教科書も学習データに含まれています。さらに、YouTube にあるスティーヴン・ストロガッツ氏の動画すべても学習済みです。
これらのモデルは、ある分野の知識を別の分野と組み合わせて応用するのが非常に得意です。
ただ、特に数学のような分野において、「すべてのものを学習した」モデルに対して「転移」という概念をどう語るべきか、私にはわかりません。
ストロガッツ:なるほど。
ミッチェル:「あらゆるデータで学習した」という表現には、ある程度の意味があると思います。もし「人間に関わるすべてのこと」を学習したと言うなら、それは明らかに事実ではありません。しかし、「数学やコードに関するすべて」を学習したと言うのであれば、どうでしょうか。数学の知識全体が、テキスト形式か動画形式としてどこかに存在しているのでしょうか。
ストロガッツ:なるほど、私に聞かれるとは。最近、数学界隈では何が起きたのかを理解しようとする中で、大きな議論が巻き起こっています。私たちはかつて、「これらの機械は検索が非常に得意だ」「大規模な空間を探索するプログラムは優れている」と考えていました。彼らが膨大な知識を持っているのは、おっしゃる通り、インターネット全体や国会図書館の資料、そして読めるものはすべて読み込んでいるからです。
つまり、知識を持ち、高速で検索・計算し、記憶を失わないといった能力が、彼らの強みとして機能します。しかし、これまで見過ごされていた異なる分野間のつながりを発見し、それを活用して長年の難問を解決する——もし人間がそれを行えば、私たちはそれを美的な高みとみなすでしょう。
数学者たちは、位相幾何学のアイデアが幾何学の問題解決に役立ったり、代数学の考え方が別の分野で活用されたりするのを好む傾向があります。しかし、これはある意味では容易いことかもしれません。すでに知られている知識を網羅的に把握し、ありとあらゆる可能性のある関連性を探索できるのであれば、稀には幸運が訪れることもあるでしょう。今回のケースも、まさにそのようなプロセスを経た結果のように思えます。
ミッチェル: そうですね、その通りだと思います。なぜそうなったのか、誰にも正確に知ることはできません。モデルの内部構造を詳しく分析できない理由がいくつもあるからです。しかし、予期せぬ二つの要素を組み合わせて、実際に機能するものを作り出す行為は、創造的だと言えます。
この状況は、ある意味で1970年代に行われた数学発見プログラムを思い出させます。ドグラス・レンアットという人物が関わっていたと記憶しています。その名前は「EURISKO」でした。これは主に数学的な新アイデアの探索を目指したもので、異なる要素を意図的に組み合わせ、それらを結合させることを試みるシステムでした。そして、数百、数千、あるいはさらに多くの候補生成を試行する仕組みを持っていました。
大半は単なるゴミでしたが、たまに面白いものが出力されることもありました。人間が中に入って確認し、「これは面白いのか?」と判断する必要がありました。機械自身でそれを理解することはできなかったのです。
では、今回のケースではどの程度のことが起きているのでしょうか?私はわかりません。ただ、ここでは明らかにスケールが桁違いに大きいという点に違いがあると思います。この問題を解決する過程で、どれほどの推論の痕跡(トークン)を生成したのか、どれほど多くの誤った道を進んだのか、そしてどうやって正しい道を見つけたのか——これらは、AI の科学の一部であり、現在あまり追求されていない分野だと私は考えています。
STROGATZ: はい、そこについて詳しくお聞きしましょう。それが私があなたに話したいことの核心です。「AI の科学」という言葉は素晴らしい表現ですね。ぜひ、メーリーさんが執筆された「AI の認知能力を評価するための6 つの原則」に関する記事も読んでみてください。ただし、まずはその6 つの原則とは何か、簡単に説明していただけませんか?
MITCHELL: もちろんです。まず第一に、自分自身の人間中心の認知的バイアスに気づくことです。私たちは、流暢な英語で話しかけてくるものに対して、無意識に人間の姿や性質を投影しがちです。そのため、これらのモデルには実際にはない人間のような特性があると誤って考えてしまう傾向があります。
2 つ目の原則は、科学者にとって常識的なことです。仮説に対して懐疑的であり、対照実験を設計することです。これはまさに「科学入門」の基礎ですが、実際にどれほど実践されているかは疑問が残ります。人々は自分の仮説に愛着を抱きがちだからです。
3 つ目の原則は、頑健性や一般化能力を検証するために、刺激やベンチマーク項目に新たなバリエーションを加えることです。
4 つ目は、これらのシステムが必ずしもブラックボックスである必要はないという点です。多様な方法で内部をプローブ(探査)することができ、なぜ特定の結果が得られるのかについて好奇心を持つ研究者が増える必要があります。
5 つ目の原則は、「パフォーマンス」と「能力」の区別を考えることです。これは「見せかけとして示せること」と「実際にできること」の違いを指します。論文ではその具体例を示しています。
6 つ目は、失敗の種類を分析し、ネガティブな結果さえも受け入れる姿勢を持つことです。ネガティブな結果を含む論文は引き出しにしまい込まれて忘れられがちですが、実は非常に示唆に富んだものになり得ます。
STROGATZ: 6 つ目の原則については、私たち全員が直接的な経験を持っていますよね?ハルシネーション(幻覚)を目の当たりにすると、これらのシステムでいったい何が起きているのかと疑問を抱かざるを得ません。実際、エラーから学ぶことは非常に多いです。
MITCHELL: はい。人々はポジティブな結果を称賛し、ネガティブな結果については説明をつけて片付けようとしがちですが、どこで失敗したかを精査することで、真に何が起きているのかを理解することが重要です。
STROGATZ: 記事の中で取り上げられている例の一つは、AI に関するものではなく、生物学や心理学から得られる教訓です。つまり、微妙な現象が起きている可能性があり、それが本当に何なのかを見極めるためには、警戒心と懐疑的な視点を持つ必要があるという話です。そこで、古くからの「賢いハンス」の物語を詳しくお聞かせいただけますか?
MITCHELL: 賢いハンスは、1900 年代初頭のドイツにいた馬です。この馬は算数の質問に答えることができました。例えば、「14 に 12 を足すと?」と尋ねると、彼はその回数だけ蹄を地面に叩きつけて答えを示したのです。まるで天才の馬のように見えました。当時の多くの科学者を含む人々は、彼が数学を行ったり、数を数えたり、人間と同じように単純な問題について推論したりできる動物だと強く信じていました。
人々は非常に興奮しましたが、オスカー・プフングストという心理学者が登場し、「では、ここで統制実験をしてみましょう」と提案しました。心理学における統制実験の概念は、当時としては比較的新しい考え方だったと思います。そして、質問する人が見えない状況でどうなるかを確認したのです。
STROGATZ: なるほど。
MITCHELL: その結果、ハンスは失敗しました。実は彼は、質問をする人の顔にある微妙な合図を読み取っていたのです。さらに驚くべきことに、質問をする人自身が答えをすでに知っていなかった場合も、ハンスは失敗することが判明しました。
実はその人が行っているのは、馬の蹄を叩く音への反応に過ぎません。答えにたどり着いた瞬間、彼が読み取っている無意識のシグナルを送っているのです。つまりこの馬は天才ですが、人々が思っていたような分野での天才ではありません。むしろ、人間の顔にある社会的なシグナルを読み取る点において天才なのです。
STROGATZ: この寓話から、AI が行う一見天才的な行為に感銘を受けたとき、私たちが得る教訓は何でしょうか?制御実験を行うべきだということですか?
MITCHELL: はい。ある AI システムが科学論文の図表に関する推論に非常に優れていることが示されました。実際、これは実例です。彼らは図表について質問にも答えられるのです。しかし、その後の対照実験では、図表を見せずに質問だけを与えました。これは奇妙に思えますよね?図表を見ずにどうやって図表についての質問に答えられるのかと。ところが驚くべきことに、AI はこのタスクを遂行できてしまいました。なぜなら、質問の言葉と正解の間には、何らかの不適切な関連性が存在していたからです。
STROGATZ: 振り返れば、これはベンチマークを試みた誰かの実験設計が不十分だったケースと言えるでしょう。
ミッチェル: 振り返ってみると、心理学や他の分野でもよくあることですが、実験デザインが不適切だったケースが多々あります。実験デザインは非常に難しく、交絡変数など様々な可能性が潜んでいます。だからこそ、科学における「再現性」の概念がこれほど重要視されるようになったのです。あるグループが実験を行い結果を得たとしても、その結果を盲目的に信じるべきではありません。その結果は、意図していなかった実験デザインの別の側面によるものだった可能性があります。そのため、独立したグループが研究を再現することが極めて重要です。しかし、AI 分野ではこの作業があまり行われていません。
ストロゴッツ: いいえ、なぜでしょうか?再現性が地味で、二番手として参加することになり、インセンティブがないからでしょうか。それは科学のあらゆる分野に共通する真実ですね。
ミッチェル: はい、科学のどの分野でも当てはまると思います。しかし同時に、AI 研究の多くが、実験手法に焦点を当てた背景を持たない、コンピュータサイエンスや関連分野出身者によって行われていることも要因の一つです。私はコンピュータサイエンティストですが、実験手法に関する講義を受けたことはありません。所属する学部でそのようなコースが開講されたこともありませんでした。コンピュータサイエンスの本質の一部として認識されていなかったのです。これが現在の AI 議論において欠けている点の一つだと考えています。AI が様々なことができることを示すこれらの実験や研究の結果を、どうやって信頼すればよいのでしょうか?
[*音楽が流れる*]
レビン: 興味深いですね。量子力学でよく議論される「観測者の干渉」という現象が、認知科学の文脈でも起こっているように思えます。つまり、実験者自身が実験やその結果に干渉しているという役割です。これは非常に面白い点です。もちろん、「クレバー・ハンス」はあまりにも有名ですが、私も同意します。この馬は社会的な手がかりを読み取る能力において、とても賢い存在です。
しかし、もしこれが AI においても起こっているとしたらどうでしょう。単に実験者の役割が干渉しているだけでなく、実験者の心理そのものが干渉しているという可能性です。これほど興味深いことはありません。
ストロガッツ: その通りです。理論科学や数学を専門とする多くの私たちには、メアリー・フリーリーも率直に認めているように、この次元については教育を受けていません。私は実験デザインに関する講義を受けたことがありませんし、あなたのような物理学者でも実験物理学の基礎課程は受けたかもしれませんが……
レビン: はい。私の実際の研究では、それはあまり重視されていません。実際、実験的なアプローチではありません。ですから、私が優れた実験を設計するアーキテクトになるのは難しいでしょうね。
ストロガッツ氏:その通りです。現在、AI 企業は頻繁にベンチマークを用いて、自社のシステムが人工一般知能(AGI)や超人的な知能の実現に向けた旅路のどこまで到達しているかを評価しています。あるいは単に競合他社との競争で優位に立つためにも利用されていますね。私たちは、これらの新しい機械学習システムや他の AI が持つ能力について知りたいと考えています。
レビン氏:私は、実は人間自身の知能をどう評価すべきか、あるいは人が思考している過程を本当に理解できているのかさえも、まだ確信が持てないと思っています。私たち自身についても、よくわかっていないのです。自分自身の内面を正確に報告することもできません。「今この文を組み立てる際、脳内でこうして処理が進んでいる」といったことを、私があなたに説明できるわけではありません。私はそれを聴き、そのプロセスを追体験したとしても、そう簡単に言語化はできないでしょう。ただ自然に言葉が溢れ出るだけです。私自身も内部のメカニズムを完全に把握しているわけではありません。AI についても同様のことが言えるのです。
多くの人が、私たちの番組で他の認知科学者やコンピュータサイエンティストと対談した際、「なぜ AI に直接問いかけないのか」と言うのですが、実は AI も正確に自己反省することは難しいのです。
ストロガッツ: この「ブラックボックス」という謎について。私たちは AI に対してこの用語を頻繁に使いますが、もちろん私たちの知性そのものもブラックボックスです。他人から見ただけでなく、自分自身にとってもそうです。あなたが強調されていた通りです。
しかし、それでは不思議に思えてきます。マジシャンに役割があるのではないか、と。マジシャンや手品師は、人間の心理的・身体的な限界を巧みに見せてくれます。いかに簡単に騙されるか、どのような認知的な誤りを犯しやすいかを示すのです。AI の欠陥や常識のなさを見せる「マジシャン」のような人々もいますね。彼らは AI に対して、まるで魔法使いが仕掛けるようなゲームを演じているわけです。それが、真剣な科学的な意味でどれほど示唆に富むものになるのか、私は興味を抱きます。
さて、メーラニーは AI の認知や理解の深さについて、さらに多くのことを語ります。また、数学を含む科学全体がどのように変容しうるかについても触れます。休憩後、その続きをお聞きください。
[*音楽が流れる*]
ストロガッツ: 『The Joy of Why』へようこそ。今日はサンタフェ研究所のコンピュータサイエンティスト、メーラニー・ミッチェル氏にお越しいただきました。
STROGATZ: 長い間、大学で教鞭を執ってきました。学生たちと接していると、彼らは正解を導き出せる一方で、その理解の深さを掘り下げていくと、「なぜ正解できたのか」に疑問が生じます。実は、間違った理由で正解しているケースもあるのです。これは、教師としていかに有益な指導ができるかという点において非常に重要です。
この点は、「能力(competence)」と「パフォーマンス(performance)」の区別にもつながります。AI の文脈では、この概念をどう捉えるべきでしょうか?
MITCHELL: 「能力」と「パフォーマンス」の違いは、心理学や言語学における古くからの議論です。ある認知機能に対する「能力」は備わっているのに、特定の課題を実行できない場合があるという考え方です。例えば、問題を解く能力はあるのに、感情的な理由で動けなくなってしまうようなケースがこれに当たります。
逆のパターンも存在します。「能力がないのにパフォーマンスだけ見せる」というケースです。オフィスアワーの学生が、教科書の特定の課題とその解答を丸暗記していたとしましょう。しかし、その背後にある一般的な原理を理解していないため、少し問題文を変えられただけで対応できなくなります。これが「能力なきパフォーマンス」の実例です。
ストロガッツ: じゃあ、AI が理解しているかどうかを検証する方法を考えようとした場合、どのような証拠が「理解」の証明になるのでしょうか。もしあなたが AI の支持者で、「これらの新しいシステムは、規模を拡大したから、あるいは世界モデルや社会モデルといった優れた新アーキテクチャのおかげで、すでに理解する段階に達した」と主張するとしましょう。単なる計算能力を超えて、実際に理解しているのだと。その「理解」の証拠とは具体的に何を指すのでしょうか?
ミッチェル: おっと、これは「理解」という言葉の厳密な定義について議論し始めるのは避けたいのですが、この言葉には実に多くの異なる意味が含まれています。
ストロガッツ: なるほど。
ミッチェル: 私たちがサンタフェ研究所で行った講演で、哲学者が「理解」を25種類に分類して解説していました。
ストロガッツ: ああ、その質問をしてしまったことで、自分がどんな泥沼にはまり込むか知らなかったですね。
ミッチェル: つまり、P 型の理解や G 型のような理解があり、非常に多様な「理解」のタイプが存在するのです。そして、私が確信しているのは、「本当の理解」という単一の概念が本当に存在するのかどうかは疑問だということです。最近、私と共同研究者たちが取り組んでいるテーマの一つは、理解を構成するさまざまな次元を分析することです。例えば、ある言語モデルやチャットボットに物語を作成させるとしましょう。「何かについての短い物語を作って」と指示すれば、彼らは美しい、一貫性のある短い物語を生成します。しかし、その物語について質問し始めると、奇妙な形で失敗することがよくあります。
ストロガッツ: ふむ。
MITCHELL: 生成した側であっても、同じことが言えると思います。多くのタスクにおいて、ある一つの次元では理解していても、別の次元では理解していないケースがあります。ある意味で、深い理解とは、こうした多様な次元すべてにわたって理解できている状態のことかもしれません。
STROGATZ: なるほど。それは有望な方向性のように聞こえますね。もう少し「タスク」という言葉についてお話ししましょう。この表現はあなたの著作でも見かけたものです。「タスクの専制(tyranny of tasks)」とは、いったい何を指すのでしょうか。
MITCHELL: 私はまず、哲学者のシャノン・ヴァルロからその言葉を聞きました。AI における世界は「タスク」によって分割されているという考え方です。つまり、AI システムが何ができるかを考える際、人々はこう言います。「ああ、要約はできるね。記事の要約能力をテストしよう」とか、「図に関する質問に答える能力を試そう」、あるいは「他のベンチマークを使おう」などです。
STROGATZ: 確かに最近は、国際数学オリンピック(IMO)のような非常に難しい高校レベルの問題や、研究レベルの問題でも頻繁にベンチマークされています。さらに現在では、未解決のオープン問題も対象となっています。これらはすべて、数学分野における3 つの異なるレベルのベンチマークと言えるでしょう。
ミッチェル: そうです、AI の能力はこうしたベンチマークで定義されています。例えば法科大学院の弁護士国家試験のような基準があり、AI はそこで非常に高い成績を収めます。そのため、「弁護士は警戒すべきだ。AI システムが人間と同等のレベルに達したことで、あなたの仕事は脅かされている」と言われることがあります。
しかし、私たちは特定の質問やタスクに対する AI の出来栄えを見ることで、その能力を定義しています。でも、実際の「仕事」は、独立したタスクが次々と続くこととは全く異なります。AI がいくつかのタスクをこなせるからといって、それらに関連する人間の仕事全体を代行できると考えるのは、一種の誤謬(ごびゅう)ではないでしょうか。
具体例を挙げましょう。ジェフリー・ヒントン氏には有名な発言があります。「AI システムは放射線画像の診断や解釈において驚異的な能力を持っている。もう放射線科医になるために学校に通う必要はない。5 年以内に AI がすべての仕事を奪うだろう」という趣旨です。
これは 2016 年の話で、すでに 10 年前になります。現在、むしろ放射線科医が不足しています。ヒントン氏の発言が原因かどうかは定かではありませんが、AI システムがベンチマーク上で医師を上回る成績を収めたとしても、それが現実世界での「仕事」をこなせることと同じ意味を持つわけではありません。現実はもっと開かれた課題に満ちており、明確に定義されたタスクの羅列だけから成り立っているわけではないのです。
STROGATZ: それでも、例えば放射線科の例のように、AI がそのタスクを本当に得意になった場合、人間である放射線医に何が残るのか気になります。私たちはまだその分野に関わるべきなのでしょうか?私の専門である数学の世界でも同様です。定理証明は非常に得意になっても、新しい概念を生み出すこと、あるいは「理論構築」と呼ばれるような領域ではまだ十分ではないかもしれません。ここには「問題解決」と「理論構築」の明確な区別があります。つまり、私たちはそれぞれの得意分野を見つけて、AI ができない部分を担うようになるのでしょうか?放射線科の場合なら、AI はスキャン読解は得意だが、より開かれた問いへの対応は苦手、といった具合です。それが私の疑問の核心です。
MITCHELL: はい。
STROGATZ: そうなるのでしょうか?
MITCHELL: 可能性はあるでしょう。これらの新しいツールによって、あなたの仕事のようなものも大きく変わるかもしれません。数学者にとって極めて有用なツールになるはずです。つまり、ご自身の業務内容が変わる可能性は十分にあります。かつてパーソナルコンピュータが登場したときのようにです。ただし、ジョージ・ラコフとラファエル・ヌニェスが共著した『数学と言葉の起源』という素晴らしい本があります。そこでは、数学におけるアイデアがメタファーを通じてどのように生まれるかが論じられています。二人は、人間の身体性(エンボディメント)が理解や数学を深める上で極めて重要な要素だと考えています。
ストロガッツ: そうです。それが私たちの唯一の希望だと思います。なぜなら、現在の機械には優れた身体性(エンボディメント)がないからです。また、数学における多くの素晴らしいアイデアは、現実世界との経験からインスピレーションを得ているという点もごもっともです。私が言いたかったのは応用数学についてで、純粋な数学よりもさらにそうだと感じています。自然や工学、社会などから得られるインスピレーションが非常に多いため、人間が応用数学の分野で有用である可能性は、純粋な数学よりもずっと高いと考えます。
ただし、純粋な数学の方が応用数学より先に人類にとって不要になる(あるいは価値を失う)のではないかとも思います。あるいはどちらもそうならないかもしれません。永遠に続いていくこともあるでしょう。あなたはどうお考えですか? 数学は往々にして「黄金基準」のようなものとして捉えられていますね。AI 企業にとっても数学は非常に有用で、自社のシステムがどれだけ優れているかを証明できます。なぜなら、問題を解決できたかどうかを検証できるからです。
ミッチェル: それは大きな疑問です。もし私の予測通り、純粋な数学が人類にとって何らかの形で不要になったとしたら、他の分野にはどのような影響があるのでしょうか? 機械がすべての分野を支配し始める道筋にあるのでしょうか? それとも、1997 年(あるいはその頃)にディープブルーがカスパロフを破ったようなケースでしょうか。当時、チェスで人類最強の棋士に勝利したからといって、それが他の分野でも同じように通用するとは限らなかったのです。
ストロガッツ: わからないですね。あなたはどう思いますか? 私には科学の方が、その点において数学よりもはるかに開かれたもののように感じられます。
ミッチェル: はい、その通りだと思います。エルデシュの問題をすべて解決したからといって、一般の人々が職を失うことを恐れる必要はないと考えます。
ストロガッツ: なるほど。そうすると、純粋な頭脳派の世界、つまり科学者や数学者の世界においてさえも、生物学的分野には測定すべきことがあまりにも多く、まだ収集していないデータが膨大にあり、新たな観測方法も次々と生まれています。私には数学と比較して、これらは尽きることがないように思えます。
ミッチェル: 同意します。おそらく数学に近い物理学においても、明確な定式化がされていない未解決の問いや、証明を構築できないような問題があまりにも多く存在していると思います。
ストロガッツ: しかし、数学への希望は、現実世界から引き続きインスピレーションを得続けることにあると感じています。フォン・ノイマンも同様のことを述べていました。「数学が芸術のための芸術となりすぎ、源流である自然や現実から遠く離れすぎると、それは無機質で枯れたものになる」と彼は言っています。
つまり、数学が自然からのインスピレーションをより多く取り入れるようになれば、純粋な数学にとって非常に良い時代になり得ると思います。20 世紀から 21 世紀にかけてはそれが少なかったですが、もし再びその原点に戻れば、おそらく人間が数学から得られる喜びをあと数百年にわたって享受できるでしょう。
ミッチェル: AI には「簡単なことは難しく、難しいことは簡単だ」という格言があります。
ストロガッツ: その通りです。
ミッチェル: 純粋な数学は、人間にとって知性と輝きの最も高貴な現れと見なされています。それはまさに「難しいこと」ですが、実は機械にとっては難しいことが容易で、簡単なことが困難になるという逆転現象が起きているのです。
ストロガッツ: はい、そして「ソフト」という言葉もありますね。科学の分野では「ハードサイエンス(硬い科学)」と「ソフトサイエンス(柔らかい科学)」を区別しますが、経済学や心理学、人類学といったソフトサイエンスこそが、実は最も難しい領域なのです。
ミッチェル: その通りです。
ストロガッツ: では、5 年後にまたお会いできることを楽しみにしています。
ミッチェル: 『ガンマの喜び』か、あるいは何か別のタイトルで。
ストロガッツ: はい、その頃には『オメガの喜び』になっているかもしれません。その頃までには、AI システムについて何を理解できるようになっていればいいと思いますか?あるいは今日では不可能な、どのようなテストを行えるようになることを望んでいますか?
ミッチェル: そうですね。私が AI 科学においてうまくいってほしいと強く願っているのは、「メカニスティック・インタープリタビリティ(機械的解釈可能性)」と呼ばれる分野です。これは神経科学のアナロジーとなるもので、システムの活性化や重み、そしてご存知の通りシステム内部の複雑な構造そのものを直接観察し、それらがより高次なレベルでどのような役割を果たしているのかを理解しようとするアプローチです。
最近、この分野は比較的小さなサブフィールドになっており、fMRI などのツールに似たようなものを開発しようとする人たちがいます。しかし、まだ誰もこれを正しく行う方法を完全に解明したわけではありません。ただ、いずれ達成できることを願っています。もしそれが実現すれば、AI の限界や能力、できないこと、犯しやすいミスの種類、そしてその修正方法などについて、真に理解するための道が開けるでしょう。
STROGATZ: その点に触れたのは興味深いですね。私があなたを知ったのも、広い意味でその文脈でした。あなたが以前「細胞自己組織化(CA)のための遺伝的アルゴリズム(GA)」という専門用語で呼ばれていた研究に取り組んでいた頃を思い出します。ジム・クラッチフィールド氏とともに、特定のクラスの問題、特に難解なコンピュータサイエンス問題を解決できるアルゴリズムを進化させるという課題に挑んでいました。その際、選択プロセスを通じて不断に改善されるより良いアルゴリズムを選別するために、進化型アルゴリズムを活用していたのです。
しかし、あなたが非常に優れたシステムを構築した後に、それを「機械的解釈可能性」のアナロジーとして捉え直した部分こそが、私にとって最も創造的だと感じました。あなたは、そのシステムがいかにして賢くなっているのかを探るため、図に示された特定のルールに従って互いに衝突する粒子のような要素として分析を試みました。
ミッチェル:はい、その通りです。私はちょうどそのような接続を明確にはしていませんでしたが、それは興味深いですね。
ストロガッツ:まさにそれこそが「解釈可能性」なのです。
そうです。そして、複雑系分野では、この現象は「創発」という概念として語られることが多いのです。
ミッチェル:私たちはそれを一種の「創発的計算」と捉えていました。AI システムにも同様に、容易には見つけられないが確実に存在する創発的計算があると考えられます。もしそれらをより深く理解できれば、システムが実際にどのように動作し、何をしているのかを解明できるはずです。
STROGATZ: 確かに興味深い姿勢ですね。正直に言って、とても温かく、古風な響きがありますね。「ああ、君たちは限られた知能しか持っていないけれど、科学を続けていけばいいんだ」という希望がそこにある。そう、笑っているでしょうね、私がどこへ向かおうとしているかを察して。少し皮肉なことを言っていますが、この傲慢さ——つまり、私たちが有限の頭脳で科学を続け、これらの AI がどうやって動作しているのかを理解し、それが私たちのゲームであり続けるのだと信じる態度です。科学の世界ではいつもそうでしたから。
しかし、私の暗い側面は、こうしたガジェットがどんどん巨大化していく中で、私たちができる日が数え切られていると考えています。彼らについて科学を続け、理解し続けることができるのか?誰が保証できるのでしょうか?それに対するあなたの反応は?他にやるべきことはありません。試すしかないのです。
MITCHELL: 興味深い質問ですね。では、そもそもなぜ私たちは科学を行うのでしょうか?問題を解決したいからというのが一つの原因です。しかし同時に、物事を理解したいという衝動に駆られて行うのです。
STROGATZ: はい。
MITCHELL: これは小さな子供たちにも見られることです。彼らは理解しようとする衝動に駆られています。多くの場合、最初に出てくる言葉の一つが「なぜ」です。彼らはそれを絶えず問いかけます。つまり、これは人間に備わった本能的な欲求であり、それに抗うのは難しいのです。だからこそ、あなたと私が科学の世界に入ったのも、それが私たちにとって重要だからなのです。
さて、私はある学会のシンポジウムで「科学における AI の役割」について議論している際、少し絶望感を覚えました。パネルには著名な研究者たちが集まり、AI が気象予測や遺伝学、宇宙論などあらゆる分野を革命化するだろうと熱弁していました。しかし、私は最後にこう尋ねました。「では、これは人類の世界理解に貢献するのでしょうか?」彼らの答えは「なぜそれを気にする必要があるのですか?」というものでした。
STROGATZ: そうですね。私にとってこれは今、誰もが考えている岐路です。科学には二面性があるからです。物事を解明することの喜びや、「なぜ」を追求する快感は、私たちの種に深く刻まれたものです。もちろん好奇心を持つことは素晴らしいですが、一方で長年、科学は技術や医療において私たちを支える実用的な手段として機能してきました。
私が抱く疑問、そして多くの方が同じように感じているのは、重要な問題解決において人類がもはや最前線にいなくなった時、私たちは依然として「好奇心の喜び」を楽しみ続けることができるのか、ということです。最後に、これまでの会話をお聞きになっていない方のために伺います。あなたがこの分野に関心を持ったきっかけは何だったのでしょうか?もし今日から始めるとしたら、同じような好奇心を抱くと思いますか?
ミッチェル: はい、素晴らしい質問ですね。私が子供の頃、論理パズルが大好きでした。真実だけを話す騎士たちと、嘘ばかりつく愚か者たちの話です。このジャンルのパズルを多数執筆した数学者レイモンド・スミスリアンの本は、いくつかありますがどれも楽しくてたまらなかったものです。
大学に進むと、ダグラス・ホフスタッターの『ゲーデル、エッシャー、バッハ』を読みました。これは、先ほどのパズルの世界版のようなものでした。彼はゲーデルの定理や数学的論理のパラドックスについて語り、それが認知や思考、創造性などいかに関連しているかを説明していました。その内容に私は完全に感銘を受け、「これが私の人生でやりたいことだ」と確信しました。何をするべきかまではっきりとはわかりませんでしたが、おそらく人工知能(AI)に関連する分野だろうと思いました。そこでダグラスを指導教員として選び、彼のグループに参加して、新しいパズル群であるアナロジー・パズルを通じてのアナロジー研究を行いました。そのすべてに私は夢中になりました。
もし私が今、あの頃の年齢に戻れるとしたら、心配することになるでしょう。実は私には機械学習で博士号を取得している息子がいます。彼も機械学習の研究をしたいと考えていますが、AI がすべての機械学習研究を行い、自らを改善していく未来では、人間が機械学習の研究をする役割はもう残らないのではないかという点に非常に不安を感じています。もし私が今の年齢なら同じように考えるだろうか。それはわかりません。
STROGATZ: 5年後にこの話題を再考する必要があるかもしれません。その頃には答えが出ている可能性もありますね。すべてがあまりにも速く進んでいるので、誰にもわかりません。貴重なお時間を割いていただき、本当にありがとうございます。これは広範で、ある種形のない対話でしたが、まさに開かれた議論でした。これに勝るガイドはいないと私は思います。ご参加いただき、心から感謝いたします。
MITCHELL: ありがとう、スティーブ。とても素晴らしい時間でした。
[*音楽が流れる*]
LEVIN: ふむ。ふむ。学生時代、ニュートンの法則を初めて学んだ時のことを思い出します。そしてケプラーの法則も学びました。これらは天体の循環への応用を通じて、ニュートンの法則の美しさを際立たせるものです。「私はこれが得意ではないから学ぶべきではない」とは思いませんでした。また、「自分がこの分野で最も優秀になるまで、情報を得る経験に喜びや楽しさを感じてはいけない」とも考えませんでした。
もちろん、多くの人はすでに他者が熟知していることを学んでいます。そこでふと考えるのですが、AI が私たちよりも先に知識を得ているとしても、私たちは依然として自分自身で理解を深める必要があるのではないでしょうか。その過程には、先ほどの学習体験に似たような喜びがあるはずです。AI が私々と自然を直接対話する間のフィルターになってしまうのではなく、私たちは依然として知識を獲得し、その経験を楽しむことができるのではないかと思うのです。確信はありませんが、もしかするとすべてが私たちの通り過ぎ去ってしまうのかもしれません。
STROGATZ: そうですね、もう少し掘り下げてみましょう。特に「最良である必要はない」という点へのあなたの強調に共感します。ある意味で、それはとても解放された考え方だと思います。
私が大学に進学した直後に気づいたのは、「最良ではないこと」がどういうことかということです。この最適化の時代において、「ナンバーワン」であることに固執するのはどうでしょうか。より速く、より安くという最適化アルゴリズムは溢れています。しかし、私たちの生活においては、往々にして私たちは「最良」ではありません。テニスではもちろん私が一番上手いわけではありませんが、それでもテニスを愛しています。チェスでも同様です。私はまだチェスが大好きですが、世界一であるわけではありません。そして、最高の父親になろうと努めていますが、果たしてそうなのかはわかりません。
それでも、これらすべてはそれ自体のために価値のある行為なのです。そうではありませんか?それらは私たちに喜びをもたらします。
私はこの点について、非常に哲学的で、ほとんど宗教的な感覚を抱いています。私たちは地球上に生きている時間は限られています。AI に関するこれらの問いかけは、究極的には「人生の意味」という問いへと繋がります。私たちは何を目指しているのでしょうか?もし人生の意味が、「ある分野で世界一になること」や「世界を変える発見をすること」だとしたら、大多数の人々の人生は無意味なものになってしまいます。私はそれが正しい人生の定義だとは信じたくありません。
私の父には当てはまりません。彼が大学に進むことさえ叶わなかったのです。ご存知の通り、彼は大恐慌の時代に育ちました。その時代、進学などという選択肢はありませんでした。彼の人生とは、良き親であり、自分が経営する靴店で靴を買ってくれる人々を大切にする生活でした。小さな町では、彼が客全員の靴のサイズを知っていたほどです。そして彼は亡くなる際にも、立派な名前を残しました。多くの人々が彼をよく覚えてくれています。
レビン: その通りです。
ストロガッツ: では、なぜそれが科学を扱うこの番組で取り上げられるのでしょうか?
レビン: そうですね。私の考えでは、人生の意味を「富の獲得」や「資産の蓄積」と捉える人々にとって、AI は大きな意味を持つでしょう。彼らはこの技術に熱狂するはずです。なぜなら、これは今よりも素早くアクセスでき、活用してさらに多くの富を得られるような、あらゆる機能を活用できる新たなツールとなるからです。
一方で、歌を歌うことや詩を書くこと、小説家になること、あるいは数学に取り組むことに人生の意味を見出す人々もいます。そして私は、これらの分野の人々が少し不安を抱いているのだと思います。自分たちの居場所が今後どうなるのか、その地位をどう守っていくべきか、そしてどう向き合うべきかを再考しているのです。
もし「もしもの話」に夢中になるなら、AI は単なるスーパーコンピュータの一種という世界もまだ存在します。これは以前にもスティーブと議論した通りです。どんなに多くの計算を高速で処理できても、結果が記号の羅列として提示されるだけなら、それは私たちにとって意味のある答えではありません。たとえ何らかの解答が含まれていようとも、誰もそれを価値あるものとは認めないのです。
私たちは依然として、スーパーコンピュータが銀河の画像を生成したり、生体医学的な神経マップの画像を見たりする際に、人間としての重要な役割を果たしています。実際には、科学者たちの仕事を奪ったわけではありません。つまり、これは私たちを追い越して置き換えるものではなく、あくまでツールとして使い続けられる可能性が高いのです。
STROGATZ: その問いが核心ですね。私には二つの plausible なシナリオがあります。一つは、AI が引き続きツールとして機能し、最先端の科学や数学において人間に不可欠な役割が残るというものです。もう一つの選択肢は、実は私の心の中ではこれが現実だと信じていますが、私たちは最先端から取り残されるようになるというシナリオです。そしてそれは、非常に近い将来に起こり得ることでしょう。
では、その場合、私たちの存在意義は何になるのでしょうか?
それでもなお、意味はあると思います。高校時代に数学の発見に触れた時のように、私個人にとっては新たな発見でした。それが世界にとっての発見だったわけではありませんよね。私たちは、おそらくそのような状況を受け入れるしかないのかもしれません。世界に向けた真の発見を成し遂げることはもうできないのです。
その役割は AI が担うようになります。私は本当に、それが非常に近い将来に起こると信じています。もしかしたら間違っているかもしれません。AI がそれを達成できない根本的な理由があるのかもしれません。例えば、彼らには身体がないし、社会的な生活もないからです。
原文を表示
When an LLM answers a question, is it reasoning like humans, or just producing text that looks like reasoning? The distinction isn’t just philosophical, this determines what we can trust AI to do, how closely we need to supervise it, and ultimately what its real-world impact will turn out to be.
Melanie Mitchell (opens a new tab) at the Santa Fe Institute argues that we lack adequate methods for measuring machine cognition, and that AI is a form of “alien intelligence” that operates through non-human cognitive mechanisms. In this episode of *The Joy of Why*, Mitchell tells Steven Strogatz how methods that psychologists use to study cognition in other kinds of “alien intelligence” — babies and animals — can be adapted to probe AI, and she lays out six principles for better assessing machine cognition. Their conversation ranges from the challenge of interpreting what’s happening inside these systems, to recent AI-assisted breakthroughs in mathematics, to why a math-performing horse from the early 1900s offers a cautionary tale for how we assess intelligence.
Listen on Apple Podcasts (opens a new tab), Spotify (opens a new tab), TuneIn (opens a new tab) or your favorite podcasting app, or you can stream it from Quanta.**
Transcript
[*Music plays*]
STEVE STROGATZ:** I’m Steve Strogatz.
JANNA LEVIN: And I’m Janna Levin.
STROGATZ: And this is *The Joy of Why*.
LEVIN: A podcast from *Quanta Magazine* where we explore some of the biggest unanswered questions in math and science today.
STROGATZ: Well, hello, hello. This is unsurprisingly yet another show about AI.
LEVIN: I’m telling you, it’s a topic people can’t seem to get enough about, and I’m becoming reluctant to pontificate anymore. It’s changing too quickly.
STROGATZ: It’s true. It is moving very fast. Anything we say could be obsolete by next week.
LEVIN: Oh yeah.
STROGATZ: As we speak, it’s July 23rd, 2026.
LEVIN: And it feels different to me than it did in July 23rd, 2025, that’s for sure.
STROGATZ: Mmm. That’s actually relevant, this talking about timelines, because our guest today, Melanie Mitchell, who is a cognitive scientist and computer scientist at Santa Fe Institute, is someone that we had on the show previously. She and I spoke about five years ago, and that is before ChatGPT.
LEVIN: Right. And was she interested in AI then?
STROGATZ: Oh, yes.
LEVIN: Okay, so it wasn’t just cognitive science.
STROGATZ: Absolutely. I, I mean, yes, I should say Melanie has been thinking about AI for a long time, and she’ll tell us about that. But the thing that’s gonna be so interesting, I feel, for us to discuss today is, um, Melanie’s point of view, which is to think about the problem of AI from the standpoint of fields like developmental psychology. Like, how does a baby or a young child get to be as intelligent as they soon become?
LEVIN: Oh, I think that’s so interesting ’cause we’re so excited about the artificial mind when we have very little comprehension of the human mind.
STROGATZ: Exactly.
LEVIN: Right, so we’re trying to skip a step.
STROGATZ: Well, that’s right. And not just human mind, but also animal minds, right? So there’s the field of comparative psychology where we look at intelligence in birds or dogs or dolphins, whatever. Um, we have a lot to learn about thinking about intelligences other than our own adult human intelligence.
LEVIN: Yeah, and this idea that we’re going to somehow simply understand a mechanism to generate an artificial intelligence when we, again, don’t understand the mechanism that brings a baby to have its level of intelligence when it’s born or when it’s developing. I mean, I think that’s really interesting to combine those two. So I’m looking forward to this one.
STROGATZ: Well, great. So then let’s dive in with Melanie Mitchell. Here she is.
[*Music plays*]
STROGATZ: Hi there, Melanie.
MELANIE MITCHELL: Hey, Steve.
STROGATZ: Very excited to see you again. This is gonna be fun. We talked a few years ago back when this show was called *The Joy of X*, and I think you may be our first return champion.
MITCHELL: Oh boy, I’m honored.
STROGATZ: Well, you should be. And, I have you back because so much feels like it’s changed in artificial intelligence. We talked, I think it was maybe 2021, and ChatGPT tidal wave hit the world at something like November of 2022. Is that right?
MITCHELL: That’s right.
STROGATZ: So everybody knows that AI is everywhere. We seem to be talking about it. People are worrying about it. Some people are excited about it. It’s certainly very widely used. I suppose I’d like to start by asking, what has surprised you the most about the past few years?
MITCHELL: Oh, wow. So much has surprised me. Just the thought that we could get to where we are now just by training these models on huge amounts of human-generated language and images and so on. I never would’ve dreamed it. So I’ve just been really surprised by what’s happened in AI. Also just the kind of polarized reaction that appeared in the AI community and society at large, I think, has been a little surprising to me, too.
STROGATZ: Polarized in terms of, like, sometimes people will distinguish AI doomers and AI optimists. Is that the kind of thing you’re talking about?
MITCHELL: There’s that dimension, then there’s the dimension of people who believe that AI is smarter than humans and people who think that it’s far, far from being anywhere near human-like intelligence. I guess related to that is sort of the love-it and hate-it. And these are separate dimensions, but maybe they’re correlated.
STROGATZ: Well, and right, and the love-it and hate-it can be also tied to things like the impact on the environment versus, you know, the economic prosperity for certain companies, but then again, what about job loss? There’s so many dimensions to this.
MITCHELL: Oh, there’s so many, yeah.
STROGATZ: But the thing that I really wanna focus on with you today is complex systems, cognitive science, artificial intelligence. You have a lot of different hats but I’m really very curious about the work that you’ve been doing to look at AI through the lens of either developmental psychology, like the way that we try to think about the alien intelligence of human babies, or comparative psychology with the alien intelligence of our pet dogs or smart birds or dolphins or that kind of thing. I mean, it’s a really interesting take on this alien intelligence of AI.
MITCHELL: Yeah. Many people have described AI as an alien kind of intelligence ’cause it’s very different from humans, even though it’s been trained on human language and books and everything on the internet and so on. But the way that these systems work, the way that they learn, the way that they reason, the way they do what they do is just really different from the way humans do it.
And this theme was actually picked up by people in developmental psychology, especially, Mike Frank at Stanford, who wrote this paper about how AI people should take some inspiration from the study of babies and young children, developmental psych. And then other people have extended that to, what about animal intelligence? And I guess one of the things that people in cog sci have been urging is that people in AI actually adopt some experimental methodologies that would make AI more like a science.
STROGATZ: Yeah, I really like this point of view, and I think it may not be so familiar to our listeners. I have to admit it wasn’t that familiar to me. You know, I never studied cognitive science, or never took a course in developmental psychology, and people in those fields have been thinking about these issues for… Well, I don’t know. You tell me.
MITCHELL: Yeah, at least 100 years.
STROGATZ: Yeah, 100 years now. Wow. And I was thinking on the way over we constantly talk about AI as a black box. That we can’t read the weights on the neurons very easily, or even if we can, we don’t know what they tell us. But for that matter, couldn’t you say that our own intelligence is in a lot of ways a black box?
MITCHELL: Absolutely. I mean, we have different ways to penetrate the black box. One is neuroscience, where we actually stick probes into neurons, or we use fMRI or other imaging techniques. There’s also psychology, where you actually look at just the behavior of a person or an animal, and try and infer from that underlying mechanisms.
And those two traditions have, for a long time, been quite separate. But the field of cognitive science tried to integrate them, and originally, the field of cognitive science also included AI. Somehow that integration didn’t work.
STROGATZ: You mean it didn’t catch on sociologically, or what do you mean?
MITCHELL: You know, originally it was thought we’re going to program them the way that humans work. And there was a very close connection between human psychology and people trying to build human psychology into AI. And then that actually didn’t yield success in AI the way that we’ve seen neural networks and learning from data rather than trying to program it in.
STROGATZ: I see.
MITCHELL: And neural networks itself was originally inspired by neuroscience, but the way that neural networks work today has diverged considerably from that original inspiration. So I think the field of machine learning has gone much more in the direction of statistics, which is quite separate from how cognitive science works.
STROGATZ: So at this point, I guess I’d like to talk a bit about benchmarks, because they do seem to be a big part of the discussion broadly in society these days. There was something that got a lot of people chattering in the world of math. One of the latest frontier models did something that looked like a kind of creativity, solved an old, longstanding math problem one of the problems that Paul Erdős, the great, Hungarian mathematician, he left lots of problems for people to think about, and one of them that they call the unit distance problem was recently solved in a very clever way by AI, and it involved putting two parts of math together in a way that hadn’t really been tried before. And so I bring that up because the last time we spoke, we were talking about an old AI that was learning to play some Atari game, or something. And you talked about how it was so good at playing, but then if you move the paddle a couple pixels up or something, it had to relearn all over again. It didn’t know how to play the slightest variation on the original game.
So the thing you said at the time that stuck with me: “The strange thing is that these machines don’t seem to be able to transfer their brilliance to any other domain than the one they’ve been trained on.” So that was five years ago. Now I guess I wonder, what do you think? Is that still true?
MITCHELL: Yeah, I mean, that particular model was not a large language model. It was a specific model to play the Atari game. Whereas now we have large language models that are trained on everything. So in some sense, they don’t have to transfer anything. They’re already trained. But, people in AI or machine learning talk about things that are in distribution and out of distribution, and that means that is this thing that we’re asking the models to do similar to things that it’s seen in its training data, or is wholly different?
And I think it’s hard to know. We don’t know what it’s been trained on. The model that’s solving these problems has certainly been trained on a lot of math because there’s a lot of math out there on the internet. It’s been trained on textbooks. It’s been trained on all of Steve Stogatz’s videos that are on YouTube. And these models are pretty good at taking things from one area and putting them together with another area.
But, you know, I don’t know how to talk about this notion of transfer when something’s been trained on everything, especially in a field like math.
STROGATZ: Huh.
MITCHELL: Where you know, “trained on everything” I think has some meaning in a way. If you say it’s been trained on everything that has to do with being human, clearly that’s not the case. But if you say it’s been trained on everything having to do with math or with code, I don’t know. Is all of mathematical knowledge out there in some kind of textual or video format?
STROGATZ: Well, you’re asking me. I, so the thing that is roiling our community in math lately as we try to make sense of what just happened is we used to think, “Okay, these machines are very good at searching,” or, “These programs are good at searching big spaces.” They have a tremendous amount of knowledge because, as you say, they’ve ingested the whole internet and the Library of Congress, and anything you can read, they’ve read.
So anything where knowledge and the ability to search and to compute very fast and to not forget, all that, that plays into their strength. But the, but to spot a connection between different branches that hadn’t been noticed before and to exploit that to solve a longstanding problem, if a human being did that, we would consider that an aesthetic high point.
You know, mathematicians love it when an idea from topology gets used to solve a problem in geometry, or when an idea from algebra helps. But then again, maybe it’s sort of easy. If you know everything that’s been done and you can look for a lot of possible connections, maybe you’ll occasionally get lucky. So that’s what it sort of seems like happened here.
MITCHELL: Yeah. No, I think that’s right. I don’t… You know, who knows how it happened because we can’t really look at the innards of the- these models very well for many reasons. But it is creative to bring two unexpected things together and have something that’s actually working. I consider that creative. But, it sort of reminds me in a way, there was a math discovery program way back in the ‘70s maybe done by this guy, Douglas Lenat. It was called EURISKO, I think. And basically it was trying to find new ideas in math. And it explicitly tried to bring together things and stick them together, and it would generate hundreds and hundreds and hundreds and hundreds of these things.
Most of them were just junk, but occasionally it would come up with something interesting. A human had to go in and look and say, “Is this interesting?” The machine couldn’t figure it out itself. So how much of that is going on here? I don’t know. I think here the difference is that the machine obviously is at a much bigger scale, and I don’t know how many tokens of reasoning trace that it generated in the course of solving this problem, and how many kind of wrong paths it went down, and how it figured out that it was on the right path. I mean, these are things that I think are part of the science of AI that not enough people are kind of pursuing right now.
STROGATZ: Yeah, let’s get into that now because that’s really where I wanted to go with you. It’s a nice phrase, the science of AI. I’d like to encourage people to look at this article of yours, Melanie, about the six principles to assess cognitive capacity of AI. But just, as a teaser, could you enunciate what are those six and say a little about them?
MITCHELL: Sure. So the first one is to be aware of your own anthropomorphic cognitive biases. So we tend to project human likeness onto things that talk to us in fluent English. So people very much think that these models have human-like qualities when maybe they actually don’t.
The second one’s a very common sense one for scientists. Be skeptical of hypotheses and develop control experiments. That’s just like Science 101, although I’m not sure how often it’s really followed through in science. People tend to like their own hypotheses.
The third is to develop novel variations of your stimuli or your benchmark items in order to test robustness and generalization.
Uh, the fourth one is these systems don’t have to be black boxes. You can probe them in many different ways and we need more people who are very curious about why they’re getting the results that they do get.
Fifth principle is to consider performance versus competence, sort of what you can show that you can do versus what you actually can do, and in the paper I give some examples of that.
The sixth is to analyze failure types and to embrace any negative results. We tend to put papers with negative results in a drawer and forget about them, but actually they can be incredibly enlightening.
STROGATZ: We all have very direct experience with number six, don’t we? When we see the hallucinations, it starts to make you wonder what’s really going on with these systems, and it’s true you learn a lot from the errors.
MITCHELL: Yeah, people celebrate their positive results and they try to explain away their negative results, but it’s important to really understand what’s going on by looking at where it fails.
STROGATZ: So one example that you give in your article, this is not about AI, but this is about the kind of lesson from biology or from psychology that subtle things can be happening that you need to have an alert and skeptical mind to notice what might really be going on. So could you just regale us with the old story of Clever Hans?
MITCHELL: So Clever Hans was a horse who lived in the early 1900s in Germany. And Clever Hans was able to answer arithmetic questions. So you’d say like, “What’s 14 plus 12?” And he would tap his hoof that many times. Looked like a genius horse. And people including many scientists living back then, were very convinced that this was an animal who could do mathematics, who could count, who could reason about simple problems in the way that humans do.
And people were very excited. But then a psychologist, named Oskar Pfungst, came along and said, “Well, let’s do some controlled experiments here,” this notion of controlled experiments you know in psychology being kind of a new idea, I think. And let’s see what happens if he can’t see the person who’s asking the question.
STROGATZ: Okay
MITCHELL: And then he fails. And it turns out what he’s doing is he’s reading subtle cues on the face of the person who’s asking the question. It turns out that if the person who’s asking the question doesn’t know the answer already, he also fails.
’Cause what the person is doing is they’re reacting to his hoof taps, and when he gets to the answer, there’s some unconscious signal they’re sending that he’s reading. So he is a genius horse, just not at the things that people thought he was a genius at. Instead, he’s a genius at reading social signals in human faces.
STROGATZ: And so in this parable then, as far as like when we are impressed by something seemingly genius that AI is doing, what is our lesson? That, that we should be doing controlled experiments, or what?
MITCHELL: Right. So, an AI system was shown to be really good at reasoning about diagrams in scientific papers, let’s say, I think this is, actually a real example, and could answer questions about them. But then the control experiment was give the questions without showing the diagrams. Seems crazy, right? How could you answer questions about a diagram without seeing the diagram? And it turned out that the AI could do this task because somehow there was some kind of spurious association between the words in the questions and the correct answer.
STROGATZ: So that seems like a case of poor experimental design on whoever was doing the benchmark attempt in retrospect.
MITCHELL: In retrospect, and in retrospect this happens all the time in psychology and other fields, I’m sure too, poor experimental design. Experimental design is a very hard thing and there’s all kinds of confounding possibilities. So this is why the notion of replication in science became so important. If one group does an experiment and they get a result, we shouldn’t necessarily believe that result. That result might be due to some other aspect of their experimental design that wasn’t intended. That’s why it’s very important for independent groups to replicate studies. This isn’t something that people in AI do very much.
STROGATZ: No, and why not? Is it that the replication is not very glamorous because you’re coming in second like there’s no incentive. That’s true in all parts of science, right?
MITCHELL: Yeah. I think that’s true in all parts of science. But it’s also because I think most of AI research is done by people whose background is in computer science or a related field that’s not focused on experimental methodology. I’m a computer scientist. I never had to take a course in experimental methodology. No such course was ever offered to me in my department. It wasn’t seen as part of what computer science was all about, and I think that’s one of the things that’s lacking in today’s AI discussion. How can we trust the results of these experiments and studies that are done that show that AI can do all these different things?
[*Music plays*]
LEVIN: Fascinating. So it seems to me that there’s this cognitive science version of the interference of the observer that everyone talks about in quantum mechanics, right? The observer themselves is interfering with the experiment or the outcome of the experiment, and that is such an interesting role. Of course, this Clever Hans is very famous, and I agree that that is a very clever horse for being able to read the social cues.
But how interesting if this is also happening with AI, that it’s, it’s not just the role of the experimenter that’s interfering, it’s actually the role of the psychology of the experimenter that’s interfering.
STROGATZ: Yeah. It’s a whole dimension that many of us in the theoretical sciences and math don’t get trained in, as Melanie freely admits. You know, I never took a course in experimental design. You as a physicist, I assume you had to take some experimental physics, but…
LEVIN: Yeah. It doesn’t really weigh in my actual work. It’s really not experimental. Yeah. So I would not be a very good architect of a good experiment.
STROGATZ: Well, and it seems like it is, something that’s a very live issue because these days the AI companies frequently use benchmarks to show how – well, to assess how – how far along are their systems on this quest for either artificial general intelligence or superhuman intelligence, that sort of thing. Or even just to out-compete the other AI companies. We would like to know what the capacities are of these new machine learning systems and other AIs.
LEVIN: Well, I think it might be that it’s just, I don’t think we really know how to evaluate human intelligence, or to really know what somebody’s doing when they’re thinking. I don’t think we know about ourselves. I don’t think we can self-report very well. I can’t say to you, “Oh, this is how it’s working in here right now as I’m constructing this sentence. I listened to it, and this was the process.” I don’t know, right? It’s just natural. It just comes out. And I’m not that privy to the inner workings, and I feel the AI similarly. A lot of people have said, I’ve had conversations on our show before with other cognitive scientists and computer scientists and they say it’s really hard for the AI to answer questions, ’cause a lot of people say, “Why don’t you just ask it?” And it can’t self-reflect either in an accurate way.
STROGATZ: This whole thought, the mystery of the black box. We use the term black box so often for the AI, but of course, our own intelligence is a black box, not just from mine to you, but even me to myself, as you’re emphasizing. But it makes me wonder if there’s a role for magicians because, you know, magicians or sleight-of-hand people are so good at showing us our own psychophysical limitations. How easily we’re fooled, or the sorts of cognitive errors we tend to make, and there are people who are analogous to the magicians who show the deficits and common sense of the AIs, right? They’re sort of playing games that are almost like magic tricks on the AIs. I wonder how revealing those will be, you know, in a serious scientific way.
Well, Melanie has a lot more to say about the depth of AI cognition and understanding, and also how it might change whole fields of science, including math. We will be hearing more about that after the break.
[*Music plays*]
STROGATZ: Welcome back to *The Joy of Why*. We’re joined today by Santa Fe Institute computer scientist Melanie Mitchell.
STROGATZ: You have been a college professor for much of your life. When you’re working with students they can get the answers right, but as you start to probe what they actually understand, you start to realize that they might be getting the right answers for the wrong reasons. They don’t really know what they’re doing, and that’s important if you wanna be a helpful teacher. This brings up another point: competence versus performance. Can you expand on this idea and, what would it mean in the AI context?
MITCHELL: So competence versus performance is kind of an old distinction from psychology and linguistics. The idea is that you might have the competence for a particular cognitive capacity, but there might be some reasons why you can’t perform the task that I’m giving you. Like they have the competence, they could solve the problems, but they’re just emotionally frozen. There’s some performance block.
But then there’s the other way around, which is performance without competence. So if the student in your office hours, say, had memorized a problem from the textbook and the solution, but they didn’t understand the general principle, so if you gave them a slightly different version of the problem, they couldn’t do it. That’s performance without competence.
STROGATZ: Okay. So if we would say that we’re trying to work out ways of testing whether the AI understands, what would count as evidence? Suppose that, you’re an AI advocate who said that these new systems, because we’ve scaled them up or because we have some nice new architecture with world models or social models or whatever, we’ve now crossed a threshold where they actually understand. It’s not just that they can compute, they understand. What would count as evidence of understanding?
MITCHELL: Oh gosh. I hate to get pedantic about understanding, but there’s so many different meanings of it.
STROGATZ: Ah.
MITCHELL: We had a talk here at Santa Fe Institute from a philosopher who broke down understanding into 25 different types.
STROGATZ: Aha. I didn’t know what I was getting myself into with the question.
MITCHELL: So there’s like P understanding and G understanding and there’s this very long typography of understanding. And I’m not sure there is any sort of single notion of real understanding. One of the recent things I and my collaborators have been working on is looking at different dimensions of understanding. One example is you can get one of these language models or chatbots to generate a story. Just generate a short story about something, and they will. They’ll generate a very beautiful little coherent short story. But then if you start asking them questions about the story, they will often will fail in weird ways.
STROGATZ: Hmm.
MITCHELL: even though they generated it. And I think the same thing is true in a lot of different tasks that they understand along one dimension but not along another dimension. And in some sense deep understanding might be just you understand across many different of these dimensions.
STROGATZ: Aha. That sounds like a promising direction. Let’s talk about tasks a little more, because that’s a phrase or a term that I’ve seen in some of your writing, the phrase, the tyranny of tasks. What’s that about?
MITCHELL: I first heard that, from Shannon Vallor, a philosopher. The idea is that in AI, the world is divided in terms of tasks. So when we think about what AI systems can do, people say, “Oh, they can make summaries. Let’s test their ability to summarize articles.” Or, “Let’s test their ability to answer questions about diagrams” or I don’t know, some other benchmark.
STROGATZ: Well, I mean, these days, they’ve been benchmarked a lot on International Mathematical Olympiad, very hard high school problems, then there were research level problems. Now there’s open problems that are unsolved in math. These are all like three levels of math benchmarks that are out there.
MITCHELL: Right, their capabilities are defined in terms of these benchmarks. You know, one benchmark might be the bar exam for law students, and they do really well on the bar exam. And so we say, “Oh, lawyers, you should be afraid. Your job is threatened because these AI systems are as getting as good as you are.” Uh, But the way that we’re defining that is by looking at how well they do on a specific set of questions or a task. And jobs as a whole are not the same as just one independent task after another. This is, I think it’s almost like a fallacy that if an AI system can do a bunch of tasks, it can do the job of a person that is associated with those tasks.
So just one example of this. So there’s a famous quote from Geoffrey Hinton, where he said something like, “AI systems are incredibly good at diagnosing or interpreting radiology images. Nobody should go to school anymore to be a radiologist. AI is gonna take all the jobs within five years.”
Well, that was 2016. That was 10 years ago. Now we actually have a shortage of radiologists. I don’t know if that’s because he said that, but uh, it turns out that even though AI systems can beat human doctors on these benchmarks, that’s not the same as doing this job out in the real world, which is much more open-ended, which is not just a series of well-defined tasks.
STROGATZ: Still, it does leave you wondering, like in the case of radiology, you could imagine if they are really good at that task, then what’s left for the human radiologist? Should we still be in that part of the game? Like in my own world of math, you know, if they’re very good at proving theorems, but they’re not so great yet at coming up with new concepts, or as we sometimes speak of it, theory building, right? There’s this big distinction between problem-solving and theory building. So is it that we’re sort of gonna find our niche, that we can do the parts that they don’t do? So like in the case of radiology, they have the open-ended part but not the scan reading part? I guess that’s what I’m wondering.
MITCHELL: Yeah.
STROGATZ: Is that how it’s gonna go?
MITCHELL: Maybe. I wouldn’t be at all surprised if jobs like yours change quite a bit because of these new tools. These are going to become incredibly useful tools for mathematicians. So it might change your job. Just like when personal computers came out, but there’s a fantastic book by um, George Lakoff and Rafael Núñez about math and where ideas in math come from, via metaphors. And they feel that human embodiment is a very important part of understanding and mathematics.
STROGATZ: Exactly. I think that’s our only hope ’cause right now they the machines don’t have great embodiment. And you’re right, that a lot of great ideas in math are inspired by experience with the world. And that’s what I was gonna say about applied math, that I feel like that’s even more so than pure math, where we get so much inspiration from nature and from engineering and society and all that, that I think we have a lot more chance of being useful as humans in applied math.
But I do think pure math will expire before applied math does, and maybe neither will. Maybe we’ll just keep going forever. What does it look like to you? I mean, math is often thought of as some kind of gold standard like, the AI companies have a lot of use for math, right? They can demonstrate how good their systems are ’cause they can verify that they’ve solved a problem or not.
MITCHELL: Well, that’s a big question I have, which is, suppose that your prediction comes right and math, pure math expires in some sense for humans. What does that mean for other fields? Does that mean that these machines are on their way to taking over everything? Or is it more like 1997 or whatever it was that Deep Blue beat Kasparov and that actually beating the best human at chess did not necessarily mean that was gonna go anywhere in other fields.
STROGATZ: I don’t know. What do you think? It feels to me like science is much more open-ended than math in that respect.
MITCHELL: Yeah, I believe that. I don’t think that solving all the Erdos problems means that the average person has to fear for their job.
STROGATZ: Okay, now we have many different things on the table at that point. But even just in the world of pure brainiacs, whether it’s scientists or mathematicians, just the fact that biology there are so many things to be measured, we have so much data that we could collect that we haven’t collected, so many new ways of observing. I mean, that seems very inexhaustible to me compared to math.
MITCHELL: I agree. And even in physics, I think, which is maybe closer to math, there’s so much you know, open-ended questions that aren’t well-formulated, that don’t have something like a proof that can be constructed.
STROGATZ: But so, I do feel like the hope for math is to continue to take inspiration from the real world. And von Neumann had said something like that too, that when math becomes too much art for art’s sake, when it drifts too far from the source, for him the source was nature or reality, if it becomes too far removed it becomes sterile, said von Neumann.
So I think this could be a a really good era for pure math if it starts taking more inspiration from nature. That’s been less so in the 20th and 21st century, but I think if we go back to that, we can probably eke out a few more centuries of human pleasure in math.
MITCHELL: I’ll just say there’s this dictum in AI which is that easy things are hard and hard things are easy.
STROGATZ: Right.
MITCHELL: And pure math is seen by humans as like the most exalted exhibition of intelligence and brilliance. It’s the hard thing, and yet we know that hard things are easier for machines and easier things are harder.
STROGATZ: Yep, and there’s the word soft also, right? In science, we talk about the hard sciences and the soft sciences, and the soft sciences of economics and psychology and anthropology, and those are the really hard ones.
MITCHELL: Right.
STROGATZ: Well, so if we meet again in five years.
MITCHELL: The Joy of Gamma, or something.
STROGATZ: Yes, The Joy of Omega by then, right. What do you hope we would understand about AI systems by then? Or what kinds of tests would we want to be able to do that we can’t do today?
MITCHELL: Yeah, I mean What I really hope will go well in the science of AI is this field called mechanistic interpretability, which is the neuroscience analog, where you’re actually looking at the activations and the weights and the, you know, all the messy innards of the system, and understanding at a higher-level sort of what they are doing.
These days, it’s kind of a smallish subfield where people are trying to develop tools that do that, analogous to things like fMRI or whatever. And I don’t think anybody’s really figured out exactly how to do this the right way yet, but I’m hoping that’s something that we can accomplish, and then we would have a genuine way of understanding sort of their limitations, what they can do, what they can’t do, what kinds of mistakes they’re likely to make, and maybe how to fix them.
STROGATZ: Interesting that you put your finger on that because the first time I became aware of you, it was in connection with that in a broad sense. So what I’m thinking of is back when you used to work on something that in the jargon was called GAs for CAs, genetic algorithms for cellular automata, you and Jim Crutchfield were looking at this problem of evolving algorithms that could solve a certain class of problems, hard computer science problems, and you were using this evolutionary algorithm to select better and better algorithms that kept improving through a kind of selection process.
But then the part that you did that I found so creative is once you’ve got a really good system, you looked at it in what felt to me like an analog of mechanistic interpretability. You tried to see what was making that system so smart, analyzing it in terms of particles that were colliding with each other according to certain rules in the diagrams. That’s, I don’t know if I’ve summarized it reasonably well, but it seems like this is a longstanding interest of yours.
MITCHELL: Yeah. that’s true. I hadn’t made that connection exactly, but that’s interesting.
STROGATZ: It is this, though. It’s interpretability.
It is interpretability. And it’s also, I think, in the field of complex systems, people talk about this notion of emergence.
STROGATZ: Yeah.
MITCHELL: And we thought of that as a kind of emergent computation. And I think these AI systems also have emergent computations that are not easy to find, but they’re there, and if we understood them better, we would understand how the system is actually working, doing what it does.
STROGATZ: Yeah, it’s an interesting attitude. It feels honestly to me very sweet and very old school. This hope that… Okay, you’re chuckling ’cause you see where I’m going. It’s a mean thing I’m saying, but this conceit that we with our limited minds can keep doing science, you know, and we’re gonna figure out how these AIs are doing what they’re doing, and that’s what our game will continue to be just like it always has been in science.
And I, the dark side of me, thinks our days are numbered to be able to do that as these gadgets get bigger and bigger. Who says we can keep doing science on them and figuring them out? What’s your reaction to that? We have nothing else to do. We have to try.
MITCHELL: That’s an interesting question. Um, why do we do science in the first place? I mean, you know, we do science ’cause we wanna solve problems. That’s one thing. But we also do science ’cause we’re driven to understand things.
STROGATZ: Yes.
MITCHELL: You see this in little children. They’re driven to understand. Often one of their first words is why. They ask it constantly. So I think that’s a human drive, and it’s hard to fight against that. And that’s why you and I both went into science, it’s important to us.
Now, I was a little despairing when I went to a panel discussion at a conference on the role of AI in science. And there were a bunch of famous people on the panel talking about how AI was going to revolutionize weather prediction, and genetics, and cosmology, and you name it. And I asked them at the end “Well, like, is this going to contribute to human understanding of the world?” And they’re like, “Why should we care about that?”
STROGATZ: Yeah. To me, this is the bifurcation that we’re all thinking about now. ’Cause science has this double-edged aspect, that it gives us pleasure, we like figuring things out, there is the joy of why, and as you say, it’s deep in our species. So yes, we’re curious, but then there’s the other side that for so long science has been this instrumental thing that helps us in technology and medicine.
And I guess the question I have, and I think a lot of us have, is will we continue to take pleasure in the joy of curiosity when we are no longer the best at solving the important problems? But let me ask you one last thing, for people who haven’t heard our earlier conversation, what was your draw to this field, and if you were starting out today, do you think you’d have the same kind of curiosity?
MITCHELL: Yeah, that’s a great question. When I was a child, I loved logic puzzles, like the knights and the knaves. The knights who always told the truth and the knaves who always lied. There’s a fun several books by Raymond Smullyan, a mathematician who wrote a bunch of puzzles in this genre that I absolutely loved.
When I got to college, I read Douglas Hofstadter’s book, Gödel, Escher, Bach, which was the real-world version of these in a way. I mean, he was talking about Gödel’s theorem and paradoxes in mathematical logic and how all this related to cognition and thinking and creativity and so on. And I was just completely blown away and that this is what I wanna do in my life. I didn’t exactly know what it was, but it seemed like it might be artificial intelligence. So I pursued Doug as an advisor and got to join his group, and was studying analogy via a new set of puzzles which were analogy puzzles. And, I was very entranced by all of that.
If I were that age today, I would be worried. In fact, I have a son who is getting a PhD in machine learning, and he wants to do research in machine learning, but he’s actually quite nervous that there will be no more roles for humans doing research in machine learning because AI will be doing all the research in machine learning and improving itself and so on and so forth. And I wonder if I’d think the same thing. I don’t know.
STROGATZ: Maybe we do have to revisit this in five years because we may know by then. Given how fast everything is going, who knows? I really appreciate your spending time with us. This has been wide-ranging, a little bit amorphous conversation, but it’s just wide open and I can’t think of a better guide to it. Thank you very much for joining us.
MITCHELL: Thanks, Steve. It’s been great.
[*Music plays*]
LEVIN: Hmm. Hmm. I, I just remember being a student and learning Newton’s laws for the first time, and then Kepler’s laws, which really make Newton’s laws beautiful, this application to the celestial cycles. I didn’t think, “Oh, I’m not the best at this, therefore I shouldn’t learn it.” Nor did I think, unless I one day become the best at this, I cannot feel pleasure or joy in my experience of acquiring this information.”
Of course, lots of people study things that other people already know and are better at. So I, I sort of wonder if maybe the AI will know things before us, but we will still need to acquire the understanding ourselves, and in that acquisition is a similar experience. Instead of maybe the AI will be a filter between us and interrogating nature directly, but we’ll still be acquiring, I don’t know, the knowledge and having that experience. I’m not sure. Maybe it’s all gonna pass us by.
STROGATZ: I– Well, let’s explore this a little more. I like especially your emphasis on not being the best, and how, in a way, unfraught that is. I, I learned as soon as I went to college what it means to not be the best. You know, this, this fixation with being the number one, especially in an age of optimization. There’s so many optimization algorithms. We talk about faster, cheaper. But in our own lives, very often we’re not the best. I’m certainly not the best tennis player. I love to play tennis. I’m not the best chess player, and I’m still happy to play chess. And try to be the best dad, but I may not be. But still, all these things are worth doing for their own sake, right? They give us pleasure.
I do feel very philosophical and almost religious about this. Like, we get a little time on Earth alive and, you know, these questions about AI do tap into questions about the meaning of life. What are we trying to do? If the meaning of life is that you’re gonna be the best in some domain or you’re gonna make a discovery that’s gonna change the world, then most people will have a meaningless life, and I just don’t wanna believe that’s the correct version of the meaning of life.
It was not for my dad. He didn’t even get to go to college. You know, he grew up in the Depression. That was not an option. His life was being a good parent and taking care of the people that bought shoes at the shoe store that he had. And he knew everyone’s shoe size in our little town, and he left a good name when he died. People remembered him well.
LEVIN: Right.
STROGATZ: So okay. What is that doing on our show here about science?
LEVIN: Well, I think that let’s say the meaning for some people of life has to do with acquisition, acquiring wealth. They’re gonna love this stuff, right? ’Cause there’s gonna be this new tool that simply leverages all kinds of buttons that they now have faster access to and can exploit and acquire more wealth.
There are people who found meaning in singing songs or writing poetry or being novelists or doing math, and, and I think all of those fields are a little more nervous, right? About reevaluating what the place is going to be for them and, and how to secure that place and how to think about it.
If I’m playing games of what may or may not happen, I mean, there is still a world in which AI is like a supercomputer, and we’ve talked about this before, Steve. Just ’cause a supercomputer can crunch all of these numbers, if it presents it to us as a string of symbols, even though it has, in some sense, an answer, it’s not a meaningful answer for us, and none of us value it.
We still, as human beings, have a very important role between us and a supercomputer rendering an image of a galaxy or looking at an image of a biomedical neural map. It hasn’t actually robbed scientists of their work. And so it might be that it really will continue to be a tool and not simply something that overtakes and discards us.
STROGATZ: Well, that’s the question, right? I think there are two plausible scenarios. One is that it continues to be a tool, and we always have some essential role in science and math at the cutting edge. The other option is, and actually in my heart I believe this is the case, that we will not be at the cutting edge, and that will happen very soon. And, so then what is the point?
Then I feel like it’s still meaningful, just like when I was in high school and I discovered things about math. They were discoveries to me. They were not discoveries to the world, you know? I think we may have to all settle for that. We’re not gonna be making genuine discoveries for the world.
The AIs will be doing that. I really do believe that’s gonna happen very soon. I may be wrong. I mean, there may be fundamental reasons why the AIs won’t be able to do that. For instance, they don’t have bodies, they don’t have social life, you kn
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み