AI が数学界に存在論的危機をもたらす
本文の状態
日本語全文を表示中
詳細モードで約41分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Verge AI
OpenAI が数学の難問に対する解決策を発表したことで、AI が高度な抽象数学を処理できるようになった事実に、数学者界隈で存在意義に関する危機感が広がっている。
AI深層分析を開く2026年8月21日 10:11
AI深層分析
キーポイント
OpenAI の数学的突破と業界への衝撃
OpenAI が長年の未解決問題に対する一連の解決策を発表し、それが数学界に爆発的な影響を与えた。
基礎計算と高度抽象化の逆転現象
AI は初等算数では依然として苦手である一方、非常に高度な抽象数学においては急速に能力を向上させている。
学術的価値と研究資金への疑問
AI が既存の難問を解決できる状況下で、新たな問題発見や次世代数学者育成のための学術助成金やプログラムに何の意味があるのかという議論が起きている。
マーケティング戦略への懐疑
数学分野への注目が、AI ラボによる単なる大規模なマーケティングキャンペーンであり、古くからある学問分野の存続には関心が薄い可能性が指摘されている。
AIの数学能力における二極化
AIは数や曜日のような基礎的な計算では依然として極めて不十分だが、学術論文に見られるような推論や分野間の接続においてはプロレベルに達している。
重要な引用
It caused a huge debate in the math community
It's funny that AI systems are all still pretty bad at elementary school arithmetic, but getting increasingly good at very high-end abstract math.
What good are academic grants and university programs... if frontier models simply answer all the outstanding questions?
Basically a bit of an existential crisis within, 'what is mathematics? What are mathematicians doing, and what is the role of mathematicians going forward?'
編集コメントを表示
編集コメント
AI が初等計算よりも高度な抽象数学で優位性を示すという事実は、技術の進化方向が人間の直感と逆転していることを浮き彫りにしている。このニュースは単なる技術報告ではなく、学問の未来像を再考させる社会的な問いかけとして捉える必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本日の『Decoder』では、ロバート・ハート氏(The Verge のロンドン拠点 AI 担当記者)と、AI が数学の分野にもたらしている影響や、多くのトップ数学者が抱えている存在意義に関する危機感についてお話しします。
OpenAI は先日、長年未解決だった数学問題に対する一連の解を発表しました。これは数学界に大きな衝撃を与え、ロバート氏はその議論を深めるため、現代を代表する数学者たちへのインタビューを行いました。
面白いのは、AI システムが小学校レベルの算数ではまだ苦手である一方、高度な抽象的な数学においては急速に能力を高めている点です。これは先進的な数学の分野において、大きな疑問を投げかけています。
もし AI がこれほどの水準で数学を処理できるなら、そのスキルは他の分野へも転用可能なのでしょうか?最先端モデルが既存の問題に対する未解決の問いすべてを即座に答えられるのであれば、新しい問題を特定するために訓練される次世代の数学者を育成する学術助成金や大学プログラムには、いったいどのような意義があるのでしょうか。
もしかすると、数学に関するこれらの注目は、最も重要な学問分野の一つである数学がどうなろうとも知ったことではないとばかりに考える最先端 AI ラボにとって、単なる大規模なマーケティング施策なのかもしれません。
この話題は多岐にわたり、ロバート氏は多くの関係者に話を聞き、さまざまな見解を収集しています。
Verge の AI 担当記者ロバート・ハートが、AI が数学に与えている影響について語ります。
※本インタビューは、長さや明確さのために軽微な編集を加えています。
ロバート・ハートさん、The Verge のロンドン拠点から AI 担当記者としてようこそDecoderへ。
お招きいただきありがとうございます。
お話しできて大変楽しみです。特に AI と数学の関係については、最近詳しく取り上げられた記事がありますよね。AI による数学の危機について、多くの数学者に話を伺ったと伺っています。
非常に複雑な問題で、まだ解明すべき点も多いと感じます。この二つの技術がどう相互作用するかは、まさに存在そのものに関わる危機ともいえるでしょう。これは Decoder の読者にとって格好の話題です。全体として何が起きているのか、広くお聞かせください。
「非常に複雑」という言葉が、現状をうまく表していると思います。本質的には、「数学とは何か」「数学者は何をしているのか」、そして「今後数学者にどのような役割が求められるのか」といった、存在そのものに関わる危機が深まっています。
この動きは、過去 6 ヶ月から 1 年ほどの間に AI の能力に関する認識が大きく転換したことが大きな要因です。AI はごく短期間で、かつては非常に苦手だった分野から、プロのレベルで真に高い性能を発揮する領域へと急激に進化しました。これは、他の分野がここ数年かけて取り組んできた課題を、極めて短い期間に圧縮して経験しているような状況です。
ソフトウェアエンジニアリングの分野に並ぶ、AI による数学危機
私はこれをソフトウェアエンジニアリングの領域と重ねて考えています。私たちはすでに一定期間、AI がソフトウェアエンジニアリングにおいて危機的な状況にあることを経験してきました。しかし、2024 年、いや直近の去年に至るまで、一般的な通説は「AI モデルは数学が特に苦手だ」というものでした。
有名な例として、これらのモデルが「ストロベリー(strawberry)という単語に含まれる R の数を数えられない」という事例があります。単に数を数えることさえも、彼らには困難だったのです。
では、なぜ数学の能力が向上したのでしょうか?それとも、一般の四則演算は依然として苦手だが、高度な数学には強いのか、あるいはその中間なのか。
残念ながら、まだ一部の分野では本当にひどい状態です。私が確認したところ、彼らは今やストロベリーの R の数を数えられるようになりました。誰かが調整を加えたのでしょうね。
私は、この「ストロベリー問題」はハードコードされていると思います。 非常に明確に言っておきますが、私の陰謀論では、すべてのモデルにこの仕様が組み込まれていると信じています。
私も同様に思います。その陰謀論には賛成です。
しかし、やはり数学や算数、さらには曜日の問題などについては依然として苦手です。先日、彼氏が「また水曜日だと言ってくる。でも今日は水曜日じゃないのに」と言っていました。あるいは時間についても同様です。エリサ・ウェルが数ヶ月前にChatGPT は時間を認識できないと報告しましたが、今もなおその状態は変わっていません。
これは数学全体の問題ではありません。
数学におけるAIの現状には、ある種のズレがあります。数学が得意であるためには、数えたり足したり掛けたりする能力が不可欠です。しかし、その多くは推論に依存しています。学術論文を見てみると、数字が出てこないケースも多々あり、それがこの状況をよく表していると思います。つまり、AIはまだ苦手な部分もありますが、一方で別の分野では驚異的な成果を上げています。
なぜそうなるのかといえば、ある時点でこれらのシステムが処理できることの臨界点に達するからです。文章生成やプログラミングの領域でその光景を目にしてきました。異なる分野間のつながりを巧みに作り出し、古い手法を新しい形で応用するなどです。現在訓練されている最新のモデルは、まさにその「理解が深まる」段階に到達したようです。そして今、数学もこなせるようになりました。
ただし、数学を単一の学問領域として語るには注意が必要です。特に外部から見るとそう思われがちですが、例えば生物学を考えてみてください。動物の行動を観察して記述することから、細胞レベルのメカニズムや生化学に至るまで、その範囲は非常に広範です。数学もまた、単一な学問ではありません。AIはある分野では極めて得意ですが、数えるような基礎的な作業については依然として苦手としています。
より抽象的なレベルにおいても、数学者の間ではトポロジー(位相幾何学)がAIにとってまだ苦手な領域の一つではないかという議論があります。正直に言えば、私はそれを検証することはできません。私の専門分野の範囲を超えているからです。とはいえ、現状は複雑で、得手不得手があるというのが実情です。
数学という分野は非常に広大で、多くの学術的な関心領域を含んでいることはご承知の通りです。モデルが驚くほど得意な分野もあれば、人々が「数学」として思い浮かべる数えるといった基礎的な部分ではまだ苦戦している箇所もあります。その間には実に幅広い中間領域が存在します。
この中間領域こそが、未来への不安や存在意義に関する危機感の源泉なのでしょうか?それとも、モデルが極めて得意とする分野にこそ問題があるのでしょうか?
実は両方の側面があります。これが今回の議論における共通テーマとなるでしょう。誰も「中程度の数学者」になることを恐れてはいませんが、この分野が果たす役割の大きさは明白です。研究や最先端の領域において、多くの結果が過剰な期待を生んでいる現状を考えると、「AI は何ができるのか」という問いは重要です。現在では、優れた数学者に匹敵する成果を生み出せる領域がある一方で、単純な計算ができないという側面も依然として存在します。
さらに、雇用構造や資金調達構造が書き換えられるかどうかという問題も絡んでいます。また、「数学的知識とは何か」という曖昧な問いや、これらの技術が担う役割についても議論が必要です。これらすべてが複雑に絡み合っているのです。
私は、あらゆる分野で AI が台頭した現状と概ね一致していると考えています。計算資源や処理能力を追加するだけで解決できる問題、あるいは検証可能性が明確な領域では、モデルは着実に進化を続けています。しかし、その中間の領域、つまり世界の知識が必要とされたり、モデル自体に現実世界に関する真の知性が求められたりする場面では、モデルは苦戦しているように見えます。
数学の分野のうち、少なくともあなたの記事や各研究所が語っている範囲では、それらはほぼ完全に自己完結した理論的な問題であるようです。まさにその領域こそが、モデルが証明を生成したり、これまで誰も解けなかった問題を解決したりできる場所です。そして、それが実際に存在するかどうかを検証し、何度も何度も再実行して確認することも可能です。
ここで、今年 5 月の話に戻りましょう。当時、まだ一般には公開されていない内部モデル「Astra」が、80 年もの間未解決だった「単位距離予想」を反証しました。そして最近では、OpenAI の「Astra」について報じられるようになりました。Astra はまさにスイッチが入った瞬間であり、皆がこれを存在論的な危機だと判断した出来事です。
Astra が達成したのは何なのか、なぜそれが大きなニュースとなるのか?
私は、Astra が初期のモデルにも関わっていた可能性が高いと考えています。OpenAI はそれを無名の内部モデルとしてリストアップしただけで、私が質問しても回答は得られませんでした。おそらくそれは Astra でしょう。
OpenAI は数週間前、ブログ記事と膨大な証明資料(数百ページに及ぶ)を公開しました。その内容は「数学および理論計算科学における 10 の進展」と題されていました。これは、最新のモデル「Astra」が何らかの形で解決したと主張された一連の学問分野のリストでした。
例えば量子ゲーム理論に関する問題がありましたが、これをどう説明すればよいのかさえわかりません。さらに説明が難しいのは、3 次元を超える高次元における球充填の問題です。他にも多くの異なる分野が含まれていました。
この発表は学界に大きな波紋を広げました。まさに爆弾のような出来事でした。これまで個別の画期的な成果はありましたが、OpenAI はそれを一度に 10 も発表したのです。しかもどれも非常に重要な内容でした。研究者たちは、「もし人間がこれらの問題を解いたなら感銘を受けるだろう」と話していました。しかし、人間がこれらすべてを同時に解決したと聞いたら、おそらく誰も信じないでしょう。
重要なのは、これらが数学者たちが実際に重視している問題だということです。過去の多くの画期的成果は「数学者が関心を持たない分野での成果だ」と批判されたことがありました。しかし今回取り上げられた問題は、優れた数学者たちが長年取り組んできた難問であり、これまで誰も解決できなかったものです。
さて、この 10 の進展について詳しく話しましょう。あなたがこれらを報じられましたが、OpenAI もいくつかの文書を提供しました。ただし、すべてが根拠のない主張ではないはずです。
彼らの研究は過去の成果に基づいています。そのため、出典の扱いについて疑問の声も上がっています。実際にはどのような反応があったのでしょうか。「モデルがこれを成し遂げた」という声だったのか、それとも AI 分野でよく見られる「他人の成果を無断借用し、巨人の肩に乗ったことに言及していない」といった批判的な反応だったのか。
私が話を聞いた人々の多くは、全体的に非常に感銘を受けたという反応でした。これらの問題は数学者たちが真剣に取り組んでいる課題です。特に注目を集めたのは、その出典の扱い方と、ブログ記事でどのように発表されたかについてです。OpenAI は当初「10 の問題があり、過去 10 年間で進展はなかった」と述べていましたが、論文を精読すると、そのうちの 1 つには明確に「我々は 2 人の研究者の進歩に基づいて構築した」と記されています。
その後、この点は静かに修正されました。私が話を聞いた数学者の一部は、この点についてあまり感銘を受けませんでした。「自分たちが大幅に利用した成果に対して、自らの言及でも認められているにもかかわらず、適切なクレジットを与えていない」という点が気になったようです。ただし、名前が挙げられた研究者の一人には、全体的な事象に対してやや複雑な感情を抱いていた人もおり、反応は賛否両様でした。
全体的な印象としては、非常に驚くべき成果でした。これらは人々を悩ませるほどの実質的なブレイクスルーであり、もし人間の数学者がこれらの成果を達成していたなら、「もし研究者の一人でもこうした問題に取り組んでいれば、学術キャリアは確約されたはずだ」と語る研究者も実際にいました。
面白いのは、学界における credit(功績)や attribution(帰属)こそがゲームそのものであるという点です。AI 企業は、人間であれば決して許されないほど無頓着な態度で臨んでも、何事もなく通り過ぎてしまうように見えます。
発見の規模や作業量が、その無頓着さを上回ったのでしょうか?もし人間が同じ目標を達成し、かつ同様に無頓着だった場合でも、反応は同じだったでしょうか?
私には「まるで盗作ではないか」という要素に頼りたくなる時もあります。しかし、実際に彼らが発表した論文を読み込めば(私が話を聞いた研究者の一人によると、世界中でこの内容を深く読み込むのはおそらく 50 人程度でしょう)、盗作を試みた形跡は全くありません。正直なところ、これは単にプレスリリースが不十分だったケースだと考えられます。
私は科学報道を 10 年以上続けており、時折こうした問題を取り上げることもありますが、プレスリリースでは実際に達成された発見の内容やその重要性、新規性が過大評価されていることがよくあります。今回もまた、まさにそうした典型的な例だと言えるでしょう。
「この成果は再現可能なのでしょうか?彼らは解けなかった 10 の問題を解決したという証拠を提示しましたが、別の 10 の問題も同様に解決できるという証明はあるのでしょうか?」
これが核心となる問いです。本記事のために取材に応じた十数人の関係者に話を聞いたところ、共通して浮かび上がった疑問は「結局、この 10 の問題を解くまでに何回挑戦したのか」という点でした。答えを知っているのは彼ら自身ですが、それを明かすつもりはないようです。
つまり、肝心なのは「これだけの成果を出すために一体どれほどの試行錯誤が必要だったのか」が不明確な点です。もしこれが最初の挑戦でいきなり 10 の問題を解決し、素晴らしい結果が出たのだとすれば、私は驚嘆するでしょうね。
次にどの分野に注力し、なぜそこを選ぶのかについても明確ではありません。おそらく、自社のモデルが対応できると判断した問題を選定・公表することには、ビジネス上の理由があるのでしょう。各 AI ラボは裏側でシニア数学者のチームを積極的に採用していますから、今後同じような成果を出せるかどうかは誰にも予測できません。
証明に関する点では、数学という学問は実験科学とは少し性質が異なります。証明とは「証明されたもの」であり、論理が通ればそれが正解です。ここで重要なのは、そのプロセスを他の専門家も追跡・検証できるかどうかなのです。各分野は非常に専門化されているため、該当分野の専門家がそれぞれの成果を検証します。私が取材した中で、今回取り上げられた分野で研究を行っている人々は「すべてが極めて妥当だ」と評価しています。
数学の分野には、Lean というプログラミング言語兼計算証明ツールが存在します。これを使えば数学的な証明をコード化し、実行して検証することが可能です。
「証明」という言葉をよく使いますが、これはその論理の厳密性や前提条件を徹底的にテストするプロセスです。この手法は実際に論文として発表されており、その有効性は認められています。確かにプレスリリースには過剰な期待が寄せられているという指摘もありますが、私が話をした人々の間では、彼らが主張する本質的なブレークスルーを疑う声はありません。
この話題にもう少し触れておきましょう。数学的証明は生成されました。これはまさに、数学が論理そのものであるように、反復可能なものです。手順を一つずつたどり、「この証明は成功した」と確認できます。高校で微積分の証明に苦労した経験のある方なら、そのプロセスをよくご存じでしょう。そこには純粋な論理が存在します。
一方、Lean などの形式証明言語で記述されるソフトウェアコードの領域では、純粋な論理をコードとして表現し、実行してその正しさを検証することが可能です。AI が理論的にこうしたタスクに優れていることは理解できます。推論を実行させればコードが生成され、そのコードを実行することで検証可能性を得られるのです。ソフトウェア工学の現場でもすでにこのプロセスは展開されており、「コードが動くか動かないか」「検証可能かどうか」をモデルが推論によって判断できることが確認されています。
しかし、私にとって最も重要な疑問は、あなたが示唆された点です。いったい何回実行すればよいのでしょうか?また、その成果が人間の数学者に指示されて得られたものではなく、AI モデル自体が行ったものであるとどうやって検証できるのでしょうか?現在、研究機関が高額な報酬で数学者を雇っている状況から、この分野が混乱しているのではないかという懸念も示されています。
フィールズ賞(数学界の最高峰)を受賞したジェームズ・メイナード氏の言葉があります。彼は「内省を迫られている」と述べています。私はこの言葉を何度も読み返しています。記事には他にも同様の発言をする数学者たちが登場し、彼らは皮肉にも検証不可能な事柄に対して内省を強いられています。つまり、「モデルはどのようにしてこれを成し遂げたのか」「その手法が数学のあり方を脅すほどにスケーラブルなのか」という問いです。私たちは、モデルが実際にどうやってこの成果を達成したのかについて、いったい何を知っているのでしょうか?
私たちが普段抱えている疑問や、まだ解明できていない事柄について、同様の懸念がいくつかあります。まず挙げられるのはモデルそのものの性質です。これは未公開のモデルであるため、独立して検証したい方にとってはハードルが高いでしょう。 proprietary(独自開発)な技術全般に言えることです。
とはいえ、彼らが使用しているモデルについて何らかの嘘をついているとは考えにくく、ある程度の信頼を置いてもよいのではないかと私は考えています。
もう一つの重要な点は、どのようにしてこのモデルにプロンプトを与え、どのようなガイドラインで操作しているかという点です。これは、その分野に精通した数学者が行っているのでしょうか?私が話を聞いた多くの数学者によれば、これが鍵となる要素のようです。これらは一般向けのモデルであることが多いですが、より広範な状況にも通じる課題です。
彼らの見解では、自分が何をしているかを知っており、具体的な指示を出せる場合、その利用は有効だといいます。あるいは、事実確認の方法を熟知しているからこそ活用できるツールとして使うことも可能です。実際、私は ChatGPT が基本的な計算で誤った数値を提示した際、「その数字は違います」と指摘すると、「ごめんなさい、おっしゃる通りです」と謝罪し、修正したものの、依然として間違った答えを返してくるという経験があります。
それでも、このレベルの精度を維持するためには、こうした検証プロセスが不可欠なのです。
これはより広範な問題への示唆を含んでいます。おっしゃる通り、これは「自動化によってその問題を解決できるのか?あるいは、最終的に人間には理解不能な状態に陥るのではないか?」という、ある種の魂の叫びのような問いかけです。そしてそれは、「数学とは何か」「なぜ数学を行うのか」「なぜこの分野を価値あるものとするのか」というより深い問いへと繋がります。
これに対する答えは人それぞれでしょう。しかし、私が話をした多くの人々が抱いているのは、数学が人間の関心の領域を超えてしまうのではないかという恐れです。もしそうなれば、私たちはもはやそれに関与しなくなるかもしれません。あるいは、一部の興味を持つ人々だけがその先へ進み、残りの人々はこれまで通り通常の生活を送り続けることになるのかもしれません。
チューリッヒの研究者ヨハネス・シュミット氏は、「AI によって数学の問題が『刈り取られて』いく状況に向かっている可能性があります。しかし、人間がプロセスから排除され、検証もできず、理解もできておらず、将来どのようなブレークスルーが待ち受けているかもわからないままでは、分野全体は前進しない」と述べています。この懸念はどれほど現実的なものだと感じられますか?大きな問題でしょうか?
確かに、問題を「刈り取る」要素は存在します。特に、若手数学者の訓練の場となり、彼らが基礎を固めるために用いられるような問題についてはその傾向が顕著です。しかし一方で、「数学とは問題解決である」という考え方は、数学という営みに対する外部者の視点に過ぎないとも私は考えています。
私が話をした数学者たちの多くは、数学の分野において「チェックボックスを埋める」作業こそが最も面白く価値のある部分ではないと考えています。彼らにとって真に価値があるのは、「何かを解決した」「証明できた」という結果そのものではなく、そこから何が起こるかです。
ジェームズ・メイナード氏が言った言葉のように、この分野で最も興味深い発見とは、単に問題を解いたことではありません。そこから何が派生するかこそが重要なのです。時にはその解法が、誰も想像していなかった全く新しい研究分野を開くこともありますし、「これはあらゆる場所で楽しく、刺激的に応用できる新たなツールだ」というケースもあります。
数学の一般化、あるいは人々が提起した定理について考えてみると、長く残るのは「問い」の方であり、「答え」ではありません。常に話題になるのはフェルマーの最終定理のようなものであり、「先人が提案した問題に対する解はこちらです」といった内容ではありません。
ここで懸念されるのは、AI がこれらの質問を次々と解決してしまうことです。通常、その過程で新たな興味深い分野へと枝分かれしたり、新しい問いが生まれたりすることを期待するものです。しかし AI はそれをしません。これが最大の懸念点であり、結果として数学の分野は非常に無機質なものになりかねません。すでに多くの成果が積み上げられた一方で、追求すべきものは何も残らないという状況です。
一般的な見解としては、まだ結論は出ていません。現時点で判断するには早すぎます。数学者が人間であっても、こうした事象の影響を実感するには相当な時間がかかります。私が述べた通り、ここ半年から一年で急激に進展しましたが、数学という分野自体が最も速いペースで動く分野ではありません。そのため、これが将来の懸念事項となるかどうかを断言するにはまだ早すぎます。
しかし、確かに懸念は存在しており、それは大きな懸念です。
マイナード氏の5月の発言には、「数学論文の出版基準がAIでは達成できないレベルにあるなら、特に博士号取得に通常4年かかることを考慮すると、課題は『今すぐAIができない問題』を見つけることではなく、『4年後のAIでもできない問題』を見つけることである」という趣旨があります。
これは、モデルの改良速度と深く関連しています。おっしゃる通り、特に数学分野ではその改善速度が加速しているように見えますが、数学のすべての領域で均一に進行しているわけではありません。
この状況こそが、学界内で自問自答を促す要因となっています。もし学生として今日から研究を始め、AIが苦手とするようなニッチな分野を選んだとしましょう。博士論文や博士課程の研究の半ばまで進んだ頃には、その問題がAIによって既に解決されており、あなたの研究は完了してしまいます。これはあなたにとって現実的な問題です。
この分野はまだこの事態に対応しているのでしょうか、それとも「モデルがこれまで不可能だと思っていたことを実行し始めた」という衝撃から抜け出せていない段階なのでしょうか?
正直に言うと、私はこの状況を「戦慄」と表現しています。インタビューのたびに、私の横にある小さなノートに広範な感想をメモしているのですが、そこには"shell shock(心的外傷後ストレス障害)"と記しました。
これは普遍的な反応ではありませんが、あまりにも急激に進んだため、多くの人が驚かされたという感覚があります。理論的には「いずれ来る」と知っていても、実際に起こったスピードは想像以上でした。
AI 数学関連のスタートアップが次々と登場し、同僚たちが異なるラボや研究領域へ移動していく様子を目にしていました。しかし、その変化の速度があまりにも速く、対応する時間がほとんどありませんでした。必ずしも得意分野で起きていることではなく、「いったい何ができるのか」という問いの方が大きくなっています。
私が学生時代、あるいはわずか半年という短い期間にこのような技術が登場していたらどうなっていたか想像もつきません。あの時なら、可能だったことの定義が根本から覆されたでしょう。夏休みを挟んで戻ってきた学生たちは、まるで全く異なる学問分野に戻ってきたような感覚を抱いているかもしれません。
世界には懐疑的な声もあります。AI に対する信頼できる批判者であるゲイリー・マルカスは、繰り返し指摘しています。AI はある特定の領域で成果を上げると、その能力があらゆる分野に一般化できると見なされる傾向がある、と。
ガリー・マルカスのサブスタック記事にある、この名言を引用しましょう。「10 年前に AI が『ジェーopardy!』で優勝したワトソンをがん治療マシンに変えようとした試みが破綻し、最終的に失敗したことを教訓として得たように、ある領域での成功がすべての領域での成功を保証するわけではありません。」
私はこの言葉を二つの解釈で受け取ることができます。一つは「AI は数学を解決した」という主張ですが、これは事実ではありません。あなたが指摘された通り、数えることさえまだ苦手な段階です。もう一つの解釈は、「AI が数学を解決したなら、必然的に他のすべての分野にも手を出してくるだろう」というものです。物理学や法律、あるいは世界中のあらゆる分野にまで及ぶでしょう。検証可能性があれば、AI はそれを解決できると言えます。なぜなら、そのように実行できるからです。
この議論の両側面は理解できます。確かに、ある領域での成功がすべての領域での成功を保証するわけではありません。しかし、AI の発展軌跡を見ると、次々と新しい領域を吸収し続けています。検証可能性がある限り、それらの領域同士を結びつける可能性の方が高いと言えるでしょう。あなたはどうお考えですか?
マルカスが批判している人々が、実際に彼が言うようなことを言っているのかは確信が持てません。過剰な宣伝(ハイク)については語るべき点が多いですが、いずれにせよ、ある領域での成功がその領域全体での成功を意味するわけではなく、ましてや他の分野での成功を意味するわけではありません。
とはいえ、ここ数年で能力の幅が確実に広がっているという傾向は否定できません。AGI(汎用人工知能)の潮流に乗る必要はありませんが、その変化が何らかの影響を及ぼすことは認めざるを得ないでしょう。
私が記事の中で引用したオックスフォード大学のアンドラス・ユハス教授の話です。彼は私に、「ChatGPT を少し触ってみたが、幾何学的な直感など全く備わっていないと思う」と語りました。それがトポロジーのような分野での進展が極めて限定的である理由を説明しているのかもしれません。数学的な観点からこの主張を検証する立場にはありませんが、この例は非常に的確に状況を表しています。
私にとってこれは、「特異点(シンギュラリティ)が目前にある」という議論全体を批判するのは容易いことでした。しかし、それはエラス・マスク氏がアストラに対して行った発言への反応であり、あまりにも過剰な解釈です。逆に、ここでの明確な進歩、そしてその速さ、さらに領域の拡大という要素を認めた上で、よりニュアンスのある視点に立つならば、単に一蹴するのではなく、考慮に値する軌道があると考えられます。
AI 全般における大きな議論の一つに、「アクセスの民主化」があります。私はかつてソフトウェア開発者としてコードを書くのが得意ではありませんでしたが、今では「バイブコーディング」でアプリを自在に作成し、家のあちこりでさまざまなことを試せるようになりました。
ここでも同様の議論が成り立つでしょうか?数学的な直感はあるものの、学術的な数学の形式言語や訓練を受けていなかった人々が、モデルにアクセスして分野を前進させることができるようになるのか。専門家からの批判を無効にするのはまさにこれです。「お前たちだけが金と教育のおかげでこのツールを使えていた」という批判に対し、「今ははるかに多くの人々が使えるようになった」と反論できるからです。
しかし、これはまた二面性のある話になります。広い意味では「Yes」ですが、私が話を聞いた数多くの数学者たちは、むしろこの状況に疲れ果てているようです。理論上はアクセスの民主化を歓迎していますが、AI 生成または AI 支援による論文があらゆる出版物や、彼らが利用するプレプリントサーバーを埋め尽くしている現状にはうんざりしています。
私が話を聞いた人々の中には、「今週だけで『これは本物なのか?』と尋ねるメールが3通も届いた」という声もありました。ChatGPT や Claude で何かを解決したと思い込んだものの、実際にそれを解けたかどうかを検証する数学的スキルを持っていないからです。
一方で、こうした成果には「有能な学部生が通常では成し得ないことを達成し、大学院レベルの作業をこなして、正当性のある論文を執筆した」という側面もあります。より大きな視点で見れば、私に話を聞いた数人の研究者は、「多くの研究機関はいわば象牙の塔の中にあり、こうしたリソースを世界的に利用可能にすれば、アクセスの格差が大幅に解消されるだろう」と述べていました。
しかし、コストの問題も無視できません。これらのシステムを実行するには多額の費用がかかります。ChatGPTやClaudeなどの無料版を利用しているときは、ついこの点を忘れがちです。上位レベルの利用では、やはりお金がかかるのです。必ずしも巨額というわけではありませんが、OpenAI は「10 件の結果に対して約 2,000 ドル」と主張していました。ただし、これは成果を得た後の他の一切のコストを考慮していないため、非常に寛容な見積もりです。
この数字さえも受け入れるなら、数学は極めて裕福な機関であっても、資金面で最も厳しい分野の一つだと言えます。セント・アンドルーズ大学のコルバ・ロニー=ドゥガル氏に話を聞いた際、「多くの場合、私は研究助成金を申請する手間すら省いています。必要ないからです。黒板があれば十分です」と語っていました。つまり、助成金さえ得ていない研究者にとって、2,000 ドルは大きな負担となります。このようにして、資金力のある機関であっても、研究者の門戸が閉ざされる可能性があります。
さらに、こうした変化のスピードも課題です。まだ誰も、これらのシステムを研究助成金の提案に組み込むことができていません。これも私が頻繁に目にしたテーマの一つでした。
実は、ロニー=ドゥーガル氏は、AI ラボの性質や数学への言及について、もう一つの素晴らしい発言をされています。彼女は「彼らは私たちの学問分野を広告の遊び場として扱っている」と述べています。多くの数学者が署名した「ライデン宣言」(Leiden Declaration) は、AI に関する過剰な期待に踊らされないよう誓う公開書簡です。
しかし、これらの動きは真っ向から対立しています。AI ラボがあらゆる学問分野を広告の材料として使い続けるのをやめる気配はありませんし、「過剰な期待には乗らない」と宣言する数学者たちの声も、その hype を止めるには至っていないのが現状です。
ここには、既存の分野を根底から覆そうとする動きに対する組織的な専門家の抵抗という側面があります。ご指摘の通り、数学はこれまで比較的安価に運営できる分野でしたが、今やさらにコストが下がり、アクセスしやすくなり、あるいは日々その地位が脅かされる可能性すらあります。数学者たちは、こうした抵抗が効果を生むと考えているのでしょうか?歴史的に見れば、数学者は政治的な駆け引きには長けていませんでした。そのため、「彼らは単に踏み潰されてしまうのではないか」と懸念する自分もいます。
数学の歴史を振り返れば、多くの数学者は実際には非常に賢明だったと思います。その際、常に思い浮かぶのがアイザック・ニュートンですが、彼は政治的には小物な側面も持っていました。しかし、これが現在の懸念です。
私が話を聞いた数人のうち、一人が特に印象的でした。彼によると、「数学は知識の頂点にある」という考えがしばしば存在するそうです。しかし、その人物は「それは嘘だ」と言い放ちました。この意見に同調する人も決して少なくありません。
とはいえ、数学は展示やデモには適しており、学問として整理されており、コストも安く済みます。マーカー氏が IBM のワトソンやがん治療への野望について言及したことを思い出してください。あれには人間を含む多くの複雑な実験が必要ですが、数学ではそのような手間がかかりません。そのため、この分野は参入しやすく、影響力を振るった後、より収益性の高い分野へ移ることも容易です。
彼らがそうしていると言っているわけではありません。これらの企業の多くの人々は採用された専門家たちです。その経歴や動機について疑うつもりはありません。しかし、長期的に見てこれが持続可能かどうかという疑問は残ります。はっきり言っておきましょう。この業界全体として見れば、数学者がこれらの企業にとって非常に収益性の高い顧客になるとは想像できません。
「この分野を推進し、経済的に実現可能なものに変えることが、もしかすると最も収益性の高いことかもしれません。数学という分野を前進させることで、その数学に基づいた工学や物理学の画期的な発見が生まれ、それが例えばロケット打ち上げの新たな方法につながるのです。
そこでは何らかの循環が起きているようですが、私は完全に理解していません。しかし、これがイノベーションの歴史です。研究から工学へ、そして収益を生む製品やサービスへと至る道筋です。これらの数学者たちは、ここで境界を押し広げることが、根本的に経済的な利益をもたらす何かの上流にあることを意識しているのでしょうか?
すぐに思い浮かぶ反論は、多くの人がこの分野が閉ざされることを恐れているという点です。つまり、定義上、「ああ、これはここでも機能する」と言えるような、驚くべき新しさにつながる画期的な発見がもう起こらなくなる可能性があります。むしろ、数学をこの新しい分野全体に応用するという収益性の高い取り組みは、道を開くのではなく閉ざすことで実現可能かどうかという未解決の問いとして残ります。
正にそうです。もしコンピュータがすべての未解決問題を解いてしまうため、個人が豊かになる動機が失われるなら、未解決の問題を解くことに対する経済的インセンティブは低下します。そこでは非常に根本的なものが崩壊するのです。
多くの場合、重要なのは単に問題を解くことだけではありません。前述の通り、これらの取り組みの本質は、その問題が他の分野において何を明らかにするかという点にあります。もしこれらの問題が極めて収益性の高いものであれば、数十年間放置されていたはずの問題を、もっと多くの人々が取り組んでいるはずです。
これは多くの純粋科学に共通する性質であり、現在のトランプ政権の科学政策に対するより広範な批判にも通じます。彼らのアプローチは非常に実用主義的でありすぎているからです。純粋研究を行うことには、定義上全く予測不可能でありながら、潜在的に莫大な利益をもたらす可能性があるという側面があります。それを計画することはできないのです。
私が数学において懸念するのは、これらの問題をすべて解決し、かつ新たな研究分野を開拓せずにそれを行えば、最終的に何が残り続けるのかということです。より収益性の高い視点や「何を狙うか」という観点から見たとしても、新しい研究領域を開拓せず、古い問題にチェックマークをつけるだけなら、その後に何が残るのかという大きな疑問符が残ることになります。
現状の進展に非常に興奮している人たちに話を聞いても、彼ら自身も何が起きているのかを本当に理解していないと口々に言います。個人的には「もしかしたらこれができるかもしれない」「あれができるかもしれない」という期待から興奮しているようですが、それでもなお、「これでこの分野はどうなるのか」という不確実感が残っています。
特に純粋な学問分野である数学研究においては、今後どうなるかを断言するのは難しい。他の科学分野では、「エンジニアリング問題や応用へとシフトするだろう」と言い切れるが、基礎的な問題をすべて解決し、そこから積み上がるものがなくなれば、その先はどこへ向かうのか。
The Verge の素晴らしい点の一つは、読者が広く、知識も豊富だということだ。あなたの記事に対して、ある数学研究者から寄せられたコメントがあり、ぜひ紹介したい。このコメントが適切な枠組みと言えるか、あなたのお考えを伺いたい。
そのコメントとはこうだ。「これらのモデルが分野に大きな変化をもたらすことに疑いの余地はないが、現状ではまだ廃れさせるには至らないだろう。むしろ、私たちの技の引き出しの中で特に有用な位置を占めることになるだろう。私が不安を抱くのは、これらがどこまで到達するのか見通せないことだが、全体的には楽観視している。適切に活用されれば、AI は数学にとって正味の利益をもたらすと思う」
「適切に活用されれば AI は X にとって正味の利益をもたらす」という考えは、人生の多くの場面でたどり着く結論のように思えるが、私が聞いた中で最も楽観的な反応だ。「うまくいけば素晴らしいことになる」。それが今の雰囲気なのか、それともまだパニック状態にあるのか。
やはり「戦場後の心的外傷」が最も強く残る印象だ。その直感的な反応は、「結局、誰にとってのプラスになるのか」「『適切』とは何を指すのか」といった疑問を招くだろう。これらはすべて、非常に正当な問いかけである。
オンライン上に投稿されたエッセイで見た大学院生たちの一部からは、暗い反応も聞こえてきた。「将来の研究員としての彼らの居場所はここにあるのか?」「これは単に AI の証明チェックをする仕事に格上げされるだけなのか?」そうだとすれば、それは非常に不満の残るキャリアになるだろう。
あるいは違うかもしれない。私にはわからない。いずれにせよ、適切に使われれば、それは確実に利益をもたらすはずだ。だが結局のところ、「適切」という言葉が何を意味するのか、そして誰のために語っているのかという点に行き着く。
不条理にも、ここで話を終わらせるのが最も適切な場所だと感じる。なぜなら、私たちがどちらも答えを知っているわけではなく、今後 1 年ほどのうちに状況が明確になっていくからだ。いずれ OpenAI は、自社のモデルで何を行ったかを人々に示さざるを得なくなるだろう。そしてより重要なのは、他の研究機関もこれらの結果を再現するか、あるいはさらに先へ進むことを示そうとするはずだ。必然的に、それはより多くの透明性を生み、さらに多くの数学者があなたたちに対して危機感を抱くことになるだろう。
ロバート、番組にお越しいただき、本当にありがとう。またすぐにお招きします。
お招きいただきありがとうございます。
*ご質問やコメントは decoder@theverge.com までご連絡ください。私たちはすべてのメールを必ず読んでいます!*
Decoder with Nilay Patel
*The Verge* が提供するポッドキャスト。大きなアイデアやその他の課題について語る番組です。
今すぐ購読する
https://pod.link/decoder
このストーリーのトピックや著者をフォローして、パーソナライズされたホームフィードで類似記事をもっと見たり、メール更新を受け取ったりしましょう。
- Nilay Patel
原文を表示
Today on *Decoder*, I’m talking with Robert Hart, *The Verge*’s London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis many lead mathematicians are having about it.
OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in the field. It caused a huge debate in the math community, and Rob spent some time talking to some of the most accomplished mathematicians of our time about it.
It’s funny that AI systems are all still pretty bad at elementary school arithmetic, but getting increasingly good at very high-end abstract math. That raises some big questions for the field of advanced math.
If AI can do math of this caliber, does that mean AI labs can transfer those skills to other domains? What good are academic grants and university programs training new generations of human mathematicians to identify new problems as they try to solve existing ones, if frontier models simply answer all the outstanding questions?
What if all this attention around math is just a big marketing exercise for frontier AI labs, which couldn’t care less what happens to one of the oldest and most fundamental academic disciplines there is?
There’s a lot here, and Robert has talked to a lot of people with a lot of views on all of it.
Okay: *Verge* AI reporter Robert Hart on what AI is doing to math. Here we go.
*This interview has been lightly edited for length and clarity. *
Robert Hart, you’re our London-based AI reporter here at *The Verge*. Welcome to *Decoder*.
Thank you for having me.
I am very excited to talk to you. There’s a lot going on in particular with AI and math that you recently dove into. You spoke to a lot of leading mathematicians about the crisis in mathematics due to AI.
It feels like a lot, and also like there’s a lot yet to know and discover about the interaction of these two things. A full existential crisis, which is pure *Decoder* bait. Broadly tell us what’s going on.
I think “a lot” sums it up quite well. Basically a bit of an existential crisis within, “what is mathematics? What are mathematicians doing, and what is the role of mathematicians going forward?”
A lot of that has been spurred by a phrase transition in what AI is capable of that has exploded in the last six months to a year. AI went from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time. It’s a lot of what these other fields have been struggling to deal with for the last five years in a very compressed period of time.
I would put that next to software engineering. We’ve been living through the AI crisis in software engineering for some amount of time. But as recently as 2024, even last year, the conventional wisdom was that AI models were particularly bad at math. The famous example is that these models could not count the number of R’s in the word strawberry. Even just counting eluded them.
What has happened to make them better at math? Are they still bad at general arithmetic and they’re good at advanced math, or is it something in between?
They are still truly, truly terrible at some areas of math. I did check, they can do strawberries now. I think someone’s tweaked it.
I think strawberry’s hard-coded. I want to be very clear, my conspiracy theory is that the strawberry thing is hard-coded into all the models.
I think so too. That is a conspiracy I’ll buy into.
But yeah, it’s still terrible at those kinds of things — math, arithmetic, even the days of the week. My boyfriend was saying the other day, “It keeps thinking it’s Wednesday. It’s not Wednesday.” Or time. Elissa Welle for us a few months ago wrote that ChatGPT can’t tell time. Still can’t. That’s not all of math.
So there’s this disconnect. To be good at math, you’ve got to be good at counting, or adding, or multiplying. A lot of it is actually reasoning. If you look at academic math papers, a lot of the time you won’t see numbers, which sums that one up, I think. So they’re still terrible, but they’re now also very good at this other part.
As to why, at some point you reach a critical mass of what these systems can do. We saw it with writing, we’ve seen it with programming. They’re very good at forging connections between different areas, applying old methods in new ways, those kinds of things. It appears that the newer models they’re training have apparently reached that level where it clicks, and now it can do math.
It’s important to say as well that we speak of math as a unitary discipline, especially from the outside. But imagine, say, biology. You’ve got something that would range from literally watching animals and describing behavior all the way through to cellular mechanisms and biochemistry. Math is not a unitary discipline either. AI is really good at some bits. Some bits like counting, it’s still really bad at.
Even on the more abstract levels, mathematicians have floated topology as one area that AI apparently still quite bad at. I can’t verify that, to be honest. It’s beyond my area of expertise. Still, it’s a bit of a mixed bag.
So you’ve described mathematics as a huge field, obviously, with many, many academic areas of interest. There are some parts where the models have gotten quite good. There are other parts, maybe the basic parts that people think of as math, which is simply counting, where they’re still struggling, and then there’s a wide range in the middle.
Is it the wide range in the middle where the existential crisis is, people don’t know what’s going to happen? Or is it at the parts where it’s really good?
A bit of both, which I feel is going to be a running theme through this. No one’s really afraid of it being a mediocre mathematician, but obviously there’s a huge element of what this field does. In terms of the research elements or the cutting edge, as we see with a lot of the results that generate hype, what can it do? There are areas now where it seems to be producing work that is on par with good mathematicians, alongside other parts where yeah, it can’t count.
All caught up in that is whether it’s going to rewrite employment structures or funding structures. You also raised the murky question of, what is mathematical knowledge? And the roles that these workers will be doing as well. It’s all of that wrapped into one.
I think that tracks broadly with the rise of AI in every field. Where you can just add horsepower or compute to a problem, and there’s some kind of verifiability, it seems like the models continue to get better. Everything in the middle where you might need some world knowledge or the models might need some actual intelligence about the world itself, they seem to struggle.
Those parts of math, at least reported out in your piece and what the labs are talking about, seem to be almost entirely self-contained theoretical problems. That’s where the models can generate a proof or solve a problem that no one’s been able to solve, and then try to verify that that has existed and they can just run it again and again and again.
That brings us, I think, to May of this year where an internal OpenAI model, which we have not really seen, disproved the unit distance conjecture, which is an 80-year-old problem. And then just recently we heard about Astra from OpenAI. Astra is the one where it seemed like the switch flipped, and everyone decided it was an existential crisis.
What did Astra achieve, and why is it a big deal?
I’m also pretty sure that Astra was probably behind the early one as well. OpenAI just listed it as an unnamed internal model. It’s probably Astra. They didn’t answer me when I asked.
OpenAI a few weeks ago dropped a blog along with a lot of paperwork proving it, I think several hundred pages. They called it “10 Advances in Mathematics and Theoretical Computer Science.” It was basically an array of disciplines that they claimed the newest model Astra had solved in some capacity. I think one was in quantum game theory, which I don’t know how to begin to explain. Even harder to explain is there was sphere packing in higher dimensions, so more than three dimensions, and there were a lot of other different disciplines as well.
It caused a lot of stir in the community. It was a bit of a bombshell. As we’d said, there’d been these individual breakthroughs that had happened, but OpenAI dropped 10 in one go, and they were quite big ones. Researchers had told me that if a human had solved these, we would be impressed. If a human had solved all 10, we probably wouldn’t believe it.
They’re problems mathematicians actually care about as well, which is an important point. A lot of previous breakthroughs have been accused of being in areas mathematicians didn’t really bother with. These are ones that mathematicians, good mathematicians, have spent a lot of time trying to solve and hadn’t.
Let’s talk about these 10. You reported them out. OpenAI did produce some documentation. But they’re not all entirely horsepowered out of nothing, right?
They’re based on previous work. There’s some question of attribution. What was the response? Is it, “Oh, the models did this?” Or was it the response we see to so much AI work, which basically boils down to, “Well you stole this and didn’t attribute anyone, and you’ve built on the shoulders of giants without mentioning it.” How did the response land?
By and large, the reaction was generally one of being quite impressed, from the people I spoke to. As I said, these are problems mathematicians care about. There was one that drew particular attention for how they credited it and also how they’d announced all of this in their blog post. OpenAI initially had said that these are 10 problems, and there had been no progress in the last 10 years. And then if you actually read the papers, one of them quite clearly says, “Oh, we build on progress from these two researchers.”
So that was later changed quite quietly. But a few of the researchers I spoke to were quite unimpressed, and they did feel it was an element of, “Well, yeah, you’ve not credited something that you’ve used heavily here, and by your own acknowledgement.” That said, one of the researchers I did speak to who was one of the ones named, and who was a bit ambivalent on the whole thing as well, so it was a real mixed bag.
The general impression was that it was quite impressive. These were actual breakthroughs that bothered people, and it did move the field forward in a way that, if a human mathematician had done these, several researchers actually said that, “Well, if a researcher had done any one of these problems, they’d probably be set for an academic career.”
It’s funny, credit and attribution in academia is the whole game. And it seems like the AI companies get away with being sloppy in a way that no human would be able to get away with being sloppy.
Did the scale of the discovery or the work overcome the sloppiness? If a human had accomplished the same goals and had been as sloppy, would the reaction be the same?
Part of me always wants to lean on the whole, “Oh, it looks like plagiarism,” element. But if you actually read the papers they produced — and one of the researchers I spoke to said there’s probably about 50 people in the world who are going to bother reading through this in depth — it is very clear. It doesn’t attempt to plagiarize. I think it was just a poor press release, to be honest.
And as much as I love to go in on it sometimes, having covered science for a decade-plus, I find the press releases are often overselling what discovery has actually been made and the import of it and the novelty of it. I think that’s just another case of what happened here.
Does this seem repeatable? There’s some proof that they provided that they solved 10 problems that were unsolvable. Do they provide any proof that they can solve another 10?
That’s the question. So the big unknown from the near-dozen people I spoke to for this was, “Well how many did they try to get these 10?” Who knows? They know, but they won’t say.
But that is the big question here. It’s unclear quite how many attempts it took to get these 10. I would be very impressed if it was the first thing they went after, and then out come these 10 impressive results.
It’s unclear what areas they would focus on next and why. There are, I imagine, business reasons behind which problems they are choosing to publicize that their models can do. All the AI labs have been hiring a cohort of senior mathematicians behind the scenes, so it’s anyone’s guess as to whether they do it again.
On the question of proof, math is a bit of an odd discipline in that repeatability is not the same as in the experimental sciences. A proof is a proof, and if it works, it works. The problem here is, well, can people follow through what they’ve done? Each field is quite highly specialized, so there’ll be individual mathematicians who are in those fields that go through the work. Those I spoke to that worked in some of the fields that were covered here say it all looks very legit.
In math, there’s a programming and computational proving language called Lean, where you can basically codify the mathematical proofs and run them through and it, well, proves it. I keep saying prove a lot here. But it will test the rigor and the assumptions of everything going on there, and they’ve published that as well. So it does appear to hold. Whilst they may say that the press release has a lot of hype or there’s a lot of hype around it, no one I’ve spoken to seems to be doubting the essential breakthroughs that they’re claiming here.
I want to stay on this subject for one more second. There’s the mathematical proof. We’ve generated a proof, and that is, as you say, just repeatable in a way that math is just logic. You can just go through the steps and say, “This proof worked,” and anybody listening to this who had to suffer through writing a proof in calculus in high school probably remembers that process. There’s something there that’s pure logic.
Then there’s a part of it that is software code, as you’re describing in Lean, where you can take the pure logic, you can express it in code, and you can run it to see if it works. I understand how AI is theoretically good at all of that. You’re just going to run the reasoning, and the reasoning is going to generate some code. You’re going to run the code, you’re going to get some verifiability. We’ve seen this play out in software engineering, where the code runs or not, it’s verifiable or not, and the models can just reason out about it.
Then there’s, to me, the big question that you alluded to. How many times do you have to run this? Can we verify that the models did this, and they weren’t directed by human mathematicians who’ve been hired at high rates by the labs in a way that suggests the field is going topsy-turvy?
You have a quote here from James Maynard, who has won the Fields Medal, the highest prize in mathematics, who said he’s been soul-searching. I keep looking at that quote. You’ve got similar quotes from all these other mathematicians in the piece, and it seems like they’re soul-searching against a thing that hilariously they cannot verify, which is, “How did the models do this, and is that thing scalable in a way that threatens mathematics?” What do we know about how the models did this?
I’d say as much as we normally do and do not know about this. There are a few issues there. One is the nature of the models. Well, this is an unreleased model, so good luck to anyone wanting to independently test it. The same goes with anything proprietary really. That said, I am inclined to almost give the benefit of the doubt that they’re not lying in some capacity about the models they’re using.
As for the other part, it’s perhaps more noteworthy to ask how did we prompt the model, or how is it being guided? Is that by a mathematician who knows what they’re doing? That’s probably a key factor here, according to a lot of mathematicians I’ve spoken with who’ve tried using these. This is often the consumer models, but still it peaks to a broader landscape.
They say that if you know what you’re doing and you can point things out, it’s good. Or you can use it as a tool in a way that you want, and in a way that you wouldn’t be able to if you didn’t really know how to fact-check it. I’ve had situations where I’ve had ChatGPT doing a basic sum, and I say, “That number is not right.” It responds, “Wait, so sorry. You’re right, it’s this,” and it’s still wrong. But that’s still needed at this level as well.
It does allude to a broader problem. As you said, this almost soul-searching of, “Well, what if we can automate that away, and what if it gets to a point where we don’t understand it?” And that really cuts to a deeper question of, “Well, what is mathematics? Why do we do it? Why do we value it as a field?”
Everyone will have different answers to that. But a fear of a lot of people I spoke to was that this might move beyond a realm of human interest, and in which case, well, maybe we just won’t engage with it. Or it’ll be something that interested people will go through, and then the rest will continue as normal.
You’ve got a quote here from a researcher in Zurich named Johannes Schmitt who says, “We might be headed toward a situation where the math problems get ‘mowed down’ by AI, but we don’t actually push the field forward because humans are taken out of the loop and they’re not either checking, or they don’t understand it, or they don’t know what the future breakthroughs might be.” How likely does that feel? Is that a big concern?
There is an element of the mowing down of the problems. Especially those that are used as a training field for younger mathematicians coming up and cutting their teeth, so to speak. But I also think that this idea, that math is problem-solving, is very much an outsider’s perspective of the mathematical endeavor.
So a lot of the mathematicians I spoke to found that the ticking boxes part is the least interesting and valuable part of the field. The areas that are valuable for them aren’t the, “Oh, you’ve solved something, or you’ve proven something.” It’s what happens from that.
I think it was James Maynard that said that the most interesting discoveries in the field aren’t that you’ve solved something, it’s what evolves from that. Sometimes those solutions open entire new fields of research that no one ever thought were possible, or, “Oh, this is a new tool that you can apply everywhere in fun and exciting ways.”
I think if we look at the popularization of math — even what I’m thinking of as those theorems that people have posed — it’s the questions that endure, not the solutions. It’s always Fermat’s Last Theorem, not like, “Well, here’s the solution to whatever the last guy proposed.”
I think the concern here is that well, they’re going to tick off all of these questions. Normally in the process of doing so, one would hope they would branch out into all of these new exciting areas or pose new questions. But AI won’t do that, and that’s the concern. And then that would leave the field quite sterile, and it will have all of these things that have been done, and maybe nothing left to pursue.
The general consensus was, well, the jury’s out. It’s too early to tell. Even with human mathematicians, it takes a lot of time to realize the impact of these kinds of things. As I said, it’s exploded in the last six months to a year, and math is not a fast-moving discipline at the best of times. But it’s too early to tell really whether that will be a concern. But it is a concern, and a big one.
You have another quote from Maynard here saying, “If the standard for a publishable paper in math is something that an AI cannot do, particularly when a PhD is typically four years, the challenge is you’re not trying to come up with a problem that AI can’t do now, it’s an AI in four years’ time.”
So this is really related to the rate of improvement of the models, which as you say, particularly in math, seems to be increasing, but not at an even rate across all of the domains of mathematics.
That appears to be what is causing the soul-searching. If you’re a student and you start today, and you pick some obscure domain that maybe the AI isn’t good at, sometime halfway through your PhD thesis or your PhD research, the AI will just solve it and you’ll be done. That is a real problem for you.
Has the field reacted to that yet, or are they just still in the shock of, “Oh, the models can start to do things that we didn’t think they were capable of”?
Yeah, I think it’s shock, really. I keep a little notebook to the side to just write down broad feelings whenever I do these interviews, and I’ve written “shell shock” in it. It’s far from universal, but it feels like it’s happened so quickly that it has just taken a lot of people by surprise. Even if they knew in theory that, well, this is coming.
They’ve seen all these AI math startups going. They’ve seen colleagues moving around to different labs or areas of work. But it just happened very, very quickly. So it’s given them very little time to figure it out. It’s not necessarily even the fields that it might be good at, it’s more just like, “What can it do?”
But I can’t imagine if something like this had come out when I was studying and literally in the space of half a year, it just upended what was possible. And over the summer as well. So students are possibly coming back to a completely different discipline after a break.
There’s some skepticism here in the world. Gary Marcus is a reliable skeptic of AI, and he pointed out that over and over again, what you see is that AI accomplishes something in one domain and then it’s used to generalize AI’s ability across every domain.
There’s a good quote from Marcus here: “As we learned a decade ago from AI’s shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine, success in one domain does not guarantee success in all.”
I can read this two ways. One, “AI has solved math,” which is not true as you’ve pointed out in several ways, all the way down to how it’s still bad at counting. And then there’s, “AI has solved math, and that means necessarily it’s going to come for everything else.” It will come for physics. It will come for law. It will come for whatever you want in the world. You can see if there’s any verifiability, AI can solve it, because you can just run it in this way.
I understand both sides of that argument, that obviously success in one domain does not guarantee success in every domain. And then the arc of AI is, well, it keeps collecting domains. If there is any verifiability, it is more likely to connect those domains than not. How do you see it?
I’m not entirely sure that the people that Marcus is criticizing here have actually said quite what he says they are saying. There’s a lot to say about the hype. But yeah, success in one domain does not even equal success throughout that domain, let alone in other domains.
That said, there has been an undeniable trajectory in the last few years of a broadening capability increase. I don’t think you need to be on the whole AGI train to acknowledge that, and to acknowledge that that will have an impact.
I think it was Andras Juhasz, one of the professors at Oxford, that I quoted in the story. Something else he’d said to me was that he’s been toying around with ChatGPT a bit, and he’s like, “I don’t think it has any geometric intuition whatsoever. Which might explain why there has been a very limited amount of progress in fields like topology.” I am in no position to verify that claim in terms of the math of it, but I think it illustrates it quite well. They call it the jagged edge.
It felt like an easy argument for me: “Let’s criticize the whole ‘the singularity is near.’” That’s what Elon Musk said in response to Astra, which is a lot. I also think it’s perhaps the least generous interpretation of that argument you can take to argue against. If you take a more nuanced element that does acknowledge that there has been clear progress here, and quite quickly, and as you said, it is racking up domains, I feel there’s a trajectory there that is a reasonable one to consider, rather than just dismiss out of hand.
One of the bigger arguments about AI in general is that it democratizes access. I was not a great software developer in my days trying to write software code, and now I can vibe-code apps at will to do all kinds of dumb stuff in my house.
Is there a similar argument here where a bunch of people who have mathematical intuition, but did not have the formalized language or training of academic mathematics, can now access a model and push the field forward? Because that is usually the thing that undercuts the criticism from the professionals is, well, many, many more people now have access to this thing that only you had access to because of your money and your training.
Annoyingly, I am going to say it’s a two-pronged thing again. But yes, in a broad sense, yes, it is. A lot of the mathematicians I spoke to were almost quite weary of this, actually. They love the idea, in theory, of democratizing access. They’re also quite fed up with AI-generated or -assisted papers that are flooding every publication imaginable, as well as the pre-print servers that they use in these fields.
Some of those I spoke to said things like, “Oh, I got three emails this week alone with people being like, ‘Hey, is this legit?’” Because they thought they’d solved something with ChatGPT or with Claude, and they also don’t have the mathematical skills to check whether they’ve actually solved something.
On the flip side, there are parts where they said, “Well, we’ve got a talented undergrad who’s done something that a talented undergrad would probably have never managed, and here they are doing grad-level work and they’ve produced a paper that is legit.” And in the bigger scheme of things, a few I spoke to said, “Well yeah, a lot of these are in the ivory tower. Having access to this kind of thing globally could really boost access to the kind of things here.”
On the flip side, the cost. These things cost a lot to run. It’s always easy to forget when you use, say, a free version of ChatGPT or Claude or something. At the higher levels, these things cost money. They may not necessarily cost a lot of money — OpenAI claimed, I think it was $2,000 for these 10 results, but that doesn’t factor in literally anything else once they’ve got these, so it’s a very generous number. But even taking that figure, math is quite a poor discipline, even at very well-off institutions.
Colva Roney-Dougal at St. Andrews, who I spoke to, said, “Well, a lot of the time I don’t bother getting a research grant. I don’t need one. I just have a blackboard.” And so if you’re not even getting a research grant, $2,000 is a lot to put up. So it could lock out researchers that way, even at quite well-funded institutions. Not to mention the speed at which this is happening, that virtually no one would’ve been able to bake any of this into a grant proposal yet, was another theme that I came across a lot.
Actually, Roney-Dougal has another great quote in your piece about the nature of the AI labs and how they are talking about math. She said, “They’re treating our discipline as an advertising playground.” A bunch of mathematicians have signed something called the Leiden Declaration, which is an open letter to then pledge not to buy into hype around AI.
These things are running right at each other. The AI labs are not going to stop using every discipline as an advertising playground. And a bunch of mathematicians saying, “We refuse to buy the hype,” certainly does not seem to be stopping the hype.
There’s just a piece of this that is organized professional resistance to a thing that is upending a field that has, as you say, been pretty cheap to operate, and now might be getting cheaper or easier to access or easier to upend, day by day. Do mathematicians feel that that is going to be effective? Historically, mathematicians are not savvy political operators. There’s a part of me that says, “Oh, they’re just going to get run over.”
I don’t know. In the history of math, actually, I think a lot of them were quite savvy. Isaac Newton is the one that always comes to mind for that — though quite a petty political operator as well. But yeah, that is the fear. A few that I spoke to, and one really comes to mind, mentioned that there’s often this belief that math is the pinnacle of knowledge. But he was like, “Well, that’s bullshit.” And he wasn’t alone in illustrating that sentiment.
But it is good for showcasing, and it’s a lot neater as a discipline, and a lot cheaper. You mentioned Marcus referencing IBM’s Watson and the curing cancer ambition. Well, that involves lots of messy experiments, including on people. You don’t need that in math, so it’s a really easy discipline to come in, throw your weight around, and then move to somewhere more lucrative if that’s what you want.
I’m not saying that that’s what they’re doing. A lot of the people at these companies have been hired. I don’t doubt their credentials for sure, and I don’t doubt their motivations as well. It does raise a question long-term as to how viable this is. Because let’s be clear, as a field goes, I cannot imagine mathematicians being a very lucrative enterprise customer for these companies.
The thing that might be lucrative is pushing a field forward to turn it into something economically viable. We push mathematics forward as a field, that turns into some engineering or physics breakthrough based on that mathematics, and that turns into, I don’t know, yet another way to launch rockets.
Some circle happens there that I don’t quite understand, but that is the history of innovation, from research, to engineering, to products or services that make money. Is that on the minds of any of these mathematicians, that pushing the boundaries here is upstream of something radically economically lucrative?
The immediate counter that would come to mind here is that a lot are scared that it’s closing off the field. So by definition, those breakthroughs that lead to something surprising and new that you can say, “Oh, this works here,” may not be happening anymore. If anything, that lucrative endeavor of applying math to this entire new field that may have a lot of money in it remains an open question as to whether anything like that would be possible if we’re closing off avenues, rather than opening them up.
Right, if the economic incentive of solving the unsolved problem is reduced because you personally won’t get rich if a computer is just solving every unsolved problem, something very fundamental breaks there.
A lot of it comes down to the fact that it’s not just solving problems. With a lot of these things, as we said, it’s about what solving that problem tells you elsewhere. If these were very lucrative problems to be solving, I imagine that more people would be trying to solve them than have left them for decades.
This is the nature of a lot of pure science, and it’s a broader criticism of what is going on perhaps with the Trump administration’s approach to science policy at the moment, in that it’s very applications-focused. There is something to doing pure research that can yield potentially very big dividends that is, by definition, utterly unpredictable as well.You cannot plan for it.
The fear I think with math is that in solving all of these problems, and then also doing so without opening up new areas of research, what are you left with? Even if it’s from a more lucrative, “What are you going after,” point of view, if you’re not opening up new areas of research and you’re just ticking off old ones, it just leaves a big question mark as to what might be left in its wake.
Even from those I spoke to that were very excited about what’s happening, they said that even they don’t really know what’s happening. And they’re excited from a personal level because, “Oh, we might be able to do this, might be to do that.” But there was still this lingering uncertainty of like, well, where does this leave the field?
Especially for more pure disciplines like research mathematics, it’s tougher to say what comes next. Because in a lot of the other sciences you can say, “Well okay, well they’d shift onto more engineering problems, or applying that.” But if you solve all the problems at the ground and there’s nothing being built up from that, where do you go from there?
One great thing about *The Verge *is our commenters are vast, they’re very knowledgeable, and there was a comment from a mathematics researcher on your story that I just want to read to you, and see if you think this is the right framework.
Here it is: “I have no doubt these models will bring massive change in the field, but in their current state, they won’t yet drive us to obsolescence. Just occupy a particularly useful spot in our bag of tricks. My apprehension comes from not knowing where these things will peak, but overall I remain optimistic. I think AI will be a net boon for math when used properly.”
I feel like “AI will be a net boon for X when used properly” is just where you land in life in a lot of things, but that’s the most optimistic response that I’ve heard: “If we get it right, it’s going to be great.” Is that the vibe, or is it still more shell-shocked than that?
I’d say shell shock is still the overriding impression. I think perhaps the gut response to that is like, “Well, it will be a net positive for whom, and what is ‘properly’?” All of those are quite legitimate questions here.
There were some bleak responses from graduate students I saw in essays posted online.Where’s their place in this as future researchers? Do they have a place in this? Is it as glorified AI proof checkers? That will be quite an unsatisfying career, I imagine.
Or maybe not, I don’t know. We will see. I think anything used properly will be a net boon. But yeah, I think it all comes down to what “properly” means, and for whom we’re talking about.
I feel like, oddly, that is an excellent place to leave it. Because I don’t think either one of us knows, and I suspect over the next year or so, things will come into focus. Because at some point, OpenAI will have to show people how they did the things of the models. And perhaps more importantly, the other labs are going to want to either replicate these results or show that they can push farther, which will necessarily have to lead to a little bit more transparency and yet more mathematicians having a crisis with you.
Robert, thank you so much for being on the show. We’ll have you back very soon.
Thank you for having me.
*Questions or comments? Hit us up at decoder@theverge.com. We really do read every email!*
Decoder with Nilay Patel
A podcast from *The Verge* about big ideas and other problems.
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
- Nilay Patel
-
-
-
-
-
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み