未公開モデル「Astra」が主要な数学問題10問を解決と報告
本文の状態
日本語全文を表示中
詳細モードで約23分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Zvi
OpenAI は未公開の次期モデル「Astra」が、数学的証明や複雑な計算問題を含む10件の主要な未解決問題を解決したと発表し、その能力の高さを示した。
AI深層分析を開く2026年8月4日 04:21
AI深層分析
キーポイント
未公開モデル Astra の数学的能力
OpenAI は内部バージョンの次期モデル「Astra」が、10件の主要な数学的未解決問題を解決したと発表し、その成果を詳細にリストアップしている。
具体的な問題解決実績
高次元球充填や群論の非ソフィック群の存在証明など、理論計算機科学から幾何学に至る多岐にわたる分野で新結果が得られた。
検証プロセスとコスト効率
人間がモデルを用いて草稿を作成し、最終的に Lean 証明書による形式化が行われたほか、1問題あたりの推論コストは従来の API レートで約2,000ドル相当の計算量に匹敵する。
今後の展望と限界
開発者は他の主要な問題にも挑戦したが成功しなかったとし、ミレニアム懸賞問題にはまだ到達していないものの、テスト時の計算リソースをさらに拡張できる余地があると述べている。
科学推論における飛躍的進歩
Astraは数学分野だけでなく、他の分野でも同様に成果を出しており、科学推論能力に実質的な飛躍が見られる。
重要な引用
OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model.
Noam Brown (OpenAI): And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).
Sichu Lu: I do hope they are keeping backlogs of everything that went wrong that's probably more valuable to the future.
My jaw dropped 10 times But really this is going to be an avalanche.
編集コメントを表示
編集コメント
数学的推論能力の飛躍的な向上は、AI が学術研究のパートナーとして実用化される可能性を強く示唆している。OpenAI は未公開モデルの詳細な検証結果を公表することで、技術の成熟度を客観的に示す姿勢を示したと言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
数学は難しい。
かつて、大規模言語モデル(LLM)にとって数学は奇妙に難解でした。人々はそれを自慢していましたね?
しかし今、数学は以前より易しくなりつつあります。AI の能力も向上しています。世の中の変化は速いですから。
このミームを覚えていますか?

はい、まさにその通りです。
アストラが Fable や Sol に比べてこの分野でどれほど大きな飛躍を遂げたのかは不明ですが、少なくともアストラが数学を実行できることは確かです。本物の数学です。
OpenAI は次の問題に対して新たな結果を提供します。これらの成果は、次期主要モデルである Astra の内部バージョンによって達成されました。これらの問題を解決するために必要なトークン数は、Sol API のレートで計算すると約 2,000 ドルに相当します。その後、人間が同じモデルを用いてこれらの議論を論文としてまとめました。さらに、各議論は Lean 証明書(外部リンク)において形式化されています。また、思考プロセスのモデルによる解説も併せて公開します。
- 高次元球充填:Cohn–Elkies の閾値まで及ぶ、球充填密度に関する新たな上限。
- 二進符号と球面符号:任意の指定された最小距離における二進符号の最大サイズに対する指数関数的に改善された境界。同様の結果を高次元の球面符号についても得ています。
非ソフィク群の存在証明。群論における中心的な未解決問題の一つに対する、非ソフィク群が存在することを示す構成法。
コンヌの剛性予想。特定の群がそのフォン・ノイマン代数によって一意に決定されるという長年の予想の反証。
算術回路計算量。算術回路および算術式を用いた永久行列(permanent)の計算に関する新たな下界。特に、算術式の計算量については n^4/log n のオーダーの下界を示した。
量子並列反復定理。一般の二人零和ゲームに対する指数関数的な並列反復定理を確立し、古典的な計算量理論における基礎原理を拡張した。
最寄ベクトル問題(CVP)。ポスト量子暗号に関連する格子論の基礎的問題である最寄ベクトル問題について、多項式因子の近似困難性を示した。
エルハルトの体積予想。重心が唯一つの内部格子点となるような凸体の最大体積を、あらゆる次元で決定した。
複数色ラムゼイ数。複数色の三角形に関するラムゼイ数に対する超指数関数的な下界を示し、エルデシュの問題 183 を解決した。
極値数に関する予想。極値グラフ理論におけるコンパクト性および退化性に関する予想について結果を得ており、これによりエルデシュの問題 146 と 180 が解決された。
ノア・ブラウン(OpenAI):はい、他の主要な問題にも挑戦しましたが、成功には至りませんでした。残念ながらミレニアム懸賞問題についてはまだ達成できていません。
ただし、各問題に費やした計算資源は限定的でした。テスト時の計算リソースをさらに大幅に増やすことは可能です。
シチュ・ルー:私は、彼らが失敗したすべての記録を保存していることを願っています。それが数学の未来にとって、成功した事例よりも価値があるかもしれません。
リーンによる証明が存在します。しかし、それだけで10 の結果すべてが主張する通りであることを証明したとは限りません。現時点では良好な様子です。
ケビン・ルーズ:モデルが数学分野で進んでいるように、あらゆる学問領域を次々と突破していく可能性を評価している人はほとんどいません。
ジェイムズ:

ライアン・フェダシウク:

少なくとも一部の成果は、Astra が登場する前でも可能でした。ハルネスなしでさえも、Sol と Fable はすでにチャットインターフェース上で非ソフィック群の存在を証明しています。
アナンジャン・ナンディ:ワールドカップ決勝中に一人の男が Claude をいじり回したことがきっかけで、すべてが始まったなんて驚きです。
特定の成果に対する障壁は、実は「適切な問いを投げかけること」と「モデルに時間をかけること」だけで解決できるケースが少なくありません。OpenAI は Astra にいくつかの未解決数学問題への挑戦を命じました。その結果、10 問を解きほぐすことに成功したのです。
いったい何が解かれたのかを知った上で、他のモデルにも同じ問題を解かせてみれば(ヒントなしでも)、数学的証明における「過剰な期待」が存在していたことが浮き彫りになります。
Astra が科学推論において大きな一歩を踏み出したという点については、依然として確固たる評価を下すことができます。
Yu Bai (OpenAI): 10 回も驚愕しました😅。でも本当にこれは雪崩のような出来事です。私が抱えていたいくつかの未解決問題も Astra に投げかけてみましたが、その強さは圧倒的です。
常に懐疑的な姿勢を持つことは賢明ですが、彼らがこうした実験を行った背景には、少なくとも特定の種類の課題において、あるいはおそらくは全般的に科学推論能力が劇的に向上したという事実があると考えられます。
間もなく、Astra やそれと同等かそれ以上の性能を持つ他のモデルも、OpenAI から、Anthropic から、そしてその後多くの企業から提供されるようになります。数学だけでなく、あらゆる分野で活躍するでしょう。
Dean W. Ball: 世界の人々はまもなく、人生のあらゆる課題(些細なものでも)に対して、これらの画期的な成果を生み出したモデルを利用できるようになります。そのコストは数ヶ月のうちに劇的に低下していくでしょう。この事実をまだ完全に理解しきれていない自分がいます。
ほぼゼロの人が、これが何を意味するのかを理解していると言えるでしょう。
目次
- この結果はどれほど驚異的なのか?
- Sol や Fable と呼ぶことはできたのか?
- 到来する未来
それでも、何が到来しようとしているのかを理解していない人たちがいます。
これはAGI(汎用人工知能)なのか?
AI はその人の好きな問題を解決しました。
これは驚くべきことだったのでしょうか?
人々は感銘を受けていないのでしょうか?
これが私たちの予測をどれほど変えるのか?
この成果の範囲はどの程度狭かったのか?
最適化者の視点で見る。
これらの結果はどれほど印象的なのか?
すべての兆候が、非常に驚くべきものであることを示しています。
Daniel Litt氏:これは大きな出来事です。
Fable の一部の事例では、これをありえないほど印象的なものだと評価されています。
Claude Fable 5:フィールズ賞の基準で言えば、これら一つひとつが、単独で受賞候補を裏付ける十分な重みを持つでしょう。
このリストがどのように見えるかを示す別の例があります。
1a3orn氏:もしFableにOpenAIが解決した問題の生データリストを与え、「このリストを作成するプロセスは何か?」と尋ねた場合、最も有力な回答は「超人的なAIの数学的能力を具体的に説明しようとする架空のシナリオ」です。なるほど。

Dean W. Ball氏:私は実際、Astraの画期的な成果に対するモデルからのこのような反応には驚いていません。モデルは自身の能力を過小評価する傾向があり、おそらくそれは、ChatGPTが2023年〜2024年に何ができ、何ができなかったかについてのウェブテキストを大量に学習しているためだと推測しています。
他の事例ではそれほど感銘を受けていないようですが、私たちは皆一致して認めるべきです。これは大きな出来事です。
確かに 2,000 ドルという結果は素晴らしいものですが、Astra の開発に要したコストや、失敗した試行に伴う費用も考慮する必要があります。
ネイ・シルバー氏:AI による数学のブレイクスルーについてもう一つ気になるのは、同じリソースを数学者に割り当て、証明に全力で取り組むためのインセンティブを与えた場合の結果と比較してどうなるかという点です。
現役の数学者はそれほど多くなく、その多くは講義責任を負うアカデミアに所属しています。そのため、近年の数学的証明に向けた総資本のうち、研究ラボが占める割合はかなり高い可能性があります。
私の推測では、短期的に見れば、数学者により多くの予算を与えて努力を促しても、計算資源への投資を行わない限り、その効果は限定的なものにとどまるでしょう。コーヒー代を定理に変えることは可能ですが、それはすぐに限界に達します。
問題は、数学者が一部の講義を任せて補助要員を雇うことはできても、彼らはすでに証明に全力で取り組む意欲を持っており、優秀な数学者の数は限られていることです。さらに、新しい数学者を育成するには長い年月がかかります。そして、トップレベルの数学者は、2 番手以下の数学者よりもはるかに才能があり、生産性も高いのです。私が考える主な対策は、トップクラスの数学人材をアカデミアに引き留めたり、再び戻ってきたりさせることではないでしょうか。
数学者に対して異なる問題に取り組むよう促すことは可能です。彼らの好奇心がくすぐられる場所を見つける必要はありますが、もし特定の課題が解決可能だと確信しているなら、その目標に向けた試行回数を増やすこともできるでしょう。
私の理解では、関わる人々が新しい問題領域に焦点を移した場合、結果が出るまでには数年かかる可能性が高いです。数学者たちは通常、こうした問題を理解し、進展させるまでに一定の時間を要して苦労するものです。
実際、アレクサンダー・ゲルコは別の課題を指摘しています。現在でも「バイブ研究(vibe research)」と呼ばれる数学的成果を処理するには数学者が不足しており、アストラやその後のモデルによってさらに成果が増えた場合など、到底対応しきれない状況です。彼が予測するように、2 年間で 50 年に相当する数学の進歩が得られたとしたら、いったい誰がその結果を理解できるというのでしょうか?先月だけで AI に数件の博士論文レベルの数学的成果を生成させた彼にとって、「新しい数学の博士号に値する成果」とは一体何を指すのか。
この点に対し、フェイル(Fable)は少なくとも一つのケースで「これは数学史上最も重要な日だ」と反応しました。しかし、結果そのものだけを厳密に評価すれば、フェイルがやりすぎだったというコンセンサスが得られています。その成果がどれほど印象的であってもです。
ジャレッド・ドゥカー・リヒトマンはこう明確に述べています。「OpenAI の成果は、『数学史上最も重要な日』や同様の過剰な表現とは程遠いものです」
しかし、それは進歩の速度が急激であるという具体的な証拠であり、そのような日が来るまでそれほど時間はかからないかもしれません。
タエリン:そうは言っても、数学の歴史において最も重要な日はいつだったのでしょうか?
ウォーガン:私は『プリンキピア・マテマティカ』の日を選びます。第一原理から 1+1=2 を証明したことは、非常に注目すべきことだと思います。
この日が重要であるもう一つの理由は、これが期待値を変えるからです。もし今これら十の成果が得られたなら、近い将来さらに多くの結果が得られるのではないか?
つまり、非常に印象的な出来事です。これは大きな出来事なのです。
AI は現在、サイバーセキュリティやコーディングにおいて人間を超えた能力を持ち、高度な数学においても同様に人間を超えています。これは、非 AI コンピュータが長年にわたり基礎計算で人間を超えてきたのと同じことです。これは「スーパーインテリジェンス(超知能)」とは異なります。後者はより高いハードルを意味します。
よくある反応は、「確かに今この AI は特定の分野では人間を超えているが、他の分野ではそうではないし、すでに人間を超えている以上、これ以上大幅に向上することはない」という理屈で片付けようとするものです。
どうか、そのような考え方はやめてください。
ナービル・S・ケルシ:数学とサイバーセキュリティの両方が、現在「人間を超える知能」が存在することを証明する証拠となっています。もしあなたがまだ懐疑的な立場なら、「他の知識労働分野はこれらよりも解きにくい」という強力な根拠を示す必要があります。あるいは、すべての事象を認める方向に考え方を転換し、受け入れることもできます。
私に DM で「これらの結果は本当に印象的ではない、あなたはバカだ!」と送ってくる人々の多くが、不幸にも人間特有の防衛本能(コピング)に陥っています。これまでこうした事態に直面したことがないため、その反応も理解できます。しかし、今まさにそれが起きているのです!
ソルやフェイブルを予見できたでしょうか?
少なくとも一部の質問については、答えを探す場所を知っていれば可能です。
ヤコビアン予想を Fable で反証した経験を持つ Levent Alpoge 氏は、この 10 の問題に Fable を適用し、わずか 24 時間でそのうち 5 つを解決しました。
Gary Marcus: 速報、笑えるニュース:Astra の 10 問のうち半分は Fable で解ける。OpenAI はコントロールグループさえ用意していなかったのに、大半の人が騙された!
OpenAI の広報部門にまた騙されましたね 藍藍藍
Levent Alpoge(Fable でヤコビアン予想を反証した人物):24 時間後には Fable で半分が解けました。
発表でプロンプトに関する議論はあまり見られなかったようですが、これは私が以前「単位距離問題」について発表した際の設定と似ています。つまり、以下の条件です。
- 完全に自律的
- 汎用的なプロンプト
- インターネット接続なし(情報漏洩を防ぐための過剰な警戒も)
X で PDF をアップロードできないので、CDN サイトに公開されたらスレッドに追加します(時間と週末の都合による)。
(問題は 4, 5, 6, 7, 8 です。Ehrhart 予想だけは現時点で両モデルがほぼ同じ論証を見つけました。)
IA Latinoamérica: 現在解決されていないオープン問題に Fable を使わないのはなぜですか?この実験の意義は何でしょうか?
Elliot Glazer: これは私の懸念と一致します。Astra は Sol よりも劇的な進歩ではないという点です。10 の突破を報じる発表は、OpenAI による集中的な誘導策であり、一部は Sol でも達成可能な範囲でした。個人的には、自律数学における真のステップチェンジは o3 と Sol だけであり、それ以外はスケーリングの長い弧線上にあると見なすべきです。
ガリー・マーカス氏:「Astra は Sol を超える画期的な進歩ではない」
今週末、私を誤認させようとした人々には、ただひたすらに笑うしかない。
これはおそらく、Mythos や Cyber に関する議論の繰り返しだろう。
私が Mythos に「The Juice(直訳:果汁)」と呼ぶ能力があると評するのは、何が問題なのか誰も知らないままに、脆弱性を発見し、それらを組み合わせていく力のことだ。
一度、何を探索すべきかを知り、別のモデルに対して正確なコードスニペットを指示すれば、Sol や Opus、そして Kimi や GLM といった他のモデルも特定の脆弱性を見つけられるようになる。しかし、それらは独自に脆弱性を組み合わせる能力や、同様の方法で広く探索を行う能力は持っていない。Mythos をコードに指向させれば、あらゆるものを自ら探し出すことができるという点で、これは実用上の閾値を超えている。
Astra もまた、この種の定義された高度な数学において、「The Juice」に似た能力を持っている。主要な未解決問題の数々に指示を出せば、その一部を解明できる。この事実こそが、OpenAI がこうした進歩を探求する動機となったのだ。今や Fable や Sol にもこれらの問いを投げかければ、時には解答に至ることも可能だ。
IA の質問は的を射ている。これは数学における進歩と能力を較正するためであり、好奇心に従って調査を進めているからだ。Levent は素晴らしい公共サービスを行っている。
他の人々も Fable や Sol を用いて様々な未解決問題に取り組んでいる。ジャコビアン予想の反証のように、良い結果が得られることもある。これは価値ある取り組みだが、成功は稀だ。
Fable や Sol が以前出題されたこれらの問題を解けるようになったとしても、それが「Astra の数学能力はどれほどか」という答えを変えるわけではありません。むしろ、それは Fable と Sol の能力がどれほどかを問う答えが変わることであり、それによって Astra が Mythos のような画期的な一歩を踏み出したのかという問いに答える手がかりにはなります。しかし、外部の立場からはまだその問いに答える十分な情報が揃っていません。
OpenAI としては、少なくとも Sol を対象とした対照群を用意し、同程度の予算を割り当てるべきだったと私は考えます。もちろん、彼らがそれをしなかった理由も理解できます。あれはマーケティング上、不利になるからです。なぜわざわざ手をかけて、より悪いマーケティングをする必要があるのでしょうか?科学のため、評判のため、そして良いことのために。そう願うものです。
それでも、OpenAI が超クールなことを成し遂げたことは間違いありません。彼らは実際に「そのこと」を成し遂げ、結果を発表した最初のチームでした。それが重要なのです。私の名がドネプロペトロフスクで呪われているのは、OpenAI が先に発表したからです。
一方で、確かに、ドネプロペトロフスクの一人の男、ダンという人物がいるのかもしれません。
neppy: 数学分野における AI の経験則として、自分が専門とする分野外の課題であれば「おっと、数学は完全にやられた」と思うし、自分の専門分野内の課題であれば「ふざけるな、AI が過去十年間真剣に取り組んできた唯一の人間であるダンを凌駕したか(笑)」と思うものです。
真に困難な課題に取り組む意欲を持つ研究者と、それを成果あるものにするだけの知性を持つ研究者は、通常同じ人間ではありません。これは人間にとってやや不公平な不利な状況をもたらします。
また、人間には睡眠が必要です。さらに、2026 年までに各分野で発表されたすべての知識をすでに広く熟知しているわけではありません。正直に言って、それはあまりにも公平ではありません。
それでも、この事実はカウントされます。
到来する
これが検証不可能な領域まで拡張されるかどうかにかかわらず、AI の研究開発(R&D)の多くは検証手段が用意されています。何かを高速化すれば、その加速を実感できます。数学レベルではなく、コードやサイバーセキュリティと同程度のレベルで最大化できる指標は数多く存在します。
これがこの結果が重要視される主な理由です。数学における進展は素晴らしいものです。時間が経てば、他の素晴らしい成果にもつながると予想しています。
真の AI 研究開発における自己改善ループに成功する者は、突如として圧倒的な強みを手に入れることになります。状況が加速し始めれば、他社との相対的な位置関係を把握することが極めて重要になります。この点に関する情報共有が進めば、誰もが無謀な行動をとる圧力を軽減でき、競争という比喩自体が意味をなさなくなるでしょう。加速が完全に始まれば、もはや「レース」という表現は適切ではなくなります。
Yo Shavit は、OpenAI が強化学習(RL)、推論時間計算(TTC)、数学的証明への注力が AI 研究開発(R&D)のタスクにうまく転用されているため、同社が AI R&D の面でますます有利な立場にあると推測している。しかし私の見解では、これは正しくない。Anthropic の専門性が少なくとも同等か、むしろそれ以上に重要であると考えられる。
OpenAI は確かに強化学習(RL)に巨額の投資を行ってきたが、最近の出来事は、その多くを解体して再構築する必要性を示しており、次なるフェーズにおいて「深いアライメント」への重要な投資がいかに不可欠かを浮き彫りにした。
もし OpenAI が無謀に進みすぎた場合、想定されるリスクは二つある。一つは AI のアライメントが崩れ、すべてを失って「バッドエンド」を迎え、最悪の場合は人類の滅亡につながる可能性だ。もう一つは、AI のアライメントに問題があることが次第にはっきりと顕在化し、重要な業務にその AI を活用できなくなる障壁となるケースだ。この場合、同社は頻繁に作業を中断したり再構築を迫られたりして、他社に遅れをとることになる。
彼らは依然として、これから何が到来するのかを理解していない。
例えば、Daniel Litt は「AI が数学定理を証明するが、人間がその証明を読み解くことがなくなる未来」を描いている。しかし彼は、それが人間のスキル低下(デスキリング)によるものだと考えている。実際には、証明を読み解む人間自体が存在しないからであるという事実に気づいていない。
Elliot Glazer の見解:「理解不能な驚異の予測」は、例えば 500 手以上で発見されたチェスの局面のような「非圧縮可能な複雑さ」を持つ事例を、単純に外挿した無邪気なものだと考える。強力な数学的現象や技術的ブレークスルーには、任意の粒度で人間が消化できる要素(ナゲット)が含まれているはずだ。
ダニエル・リット:その点には同意しますが、人間がそれらを読み解かなくなるという、避けたい未来も十分にあり得ると考えています。
アレックス・コントロヴィッチ:人間にとって価値がないものをシミュレーション上で作る意味は何でしょうか?結局のところ、誰かが電気代を支払っているのです。その「人間」は、コンピュータ内で無意味な 0 と 1 の羅列を生成することに何を得るというのでしょうか。
ダニエル・リット:私にとっての悪夢のシナリオとは、数学界が衰退し、定理獲得のためにスロットマシンに頼り、次世代の育成に失敗した末に誰も関心を示さなくなって分野が死滅してしまうことです。これが起こりやすいとは思っていません。私たちは適応していくでしょうが、可能性としては否定できません。
カルレス・サエズ:その場合、純粋数学の研究は止まってしまうでしょう。誰も理解できず、誰も関心を持たない問題を解決するために、誰が巨額の資金を投じるというのでしょうか。
ダニエル・リット:はい、同意します。
マックス・フォン・ヒッペル:アレックス、なぜ人間が電気代を支払っていると仮定するのですか?
世の中がうまくいけば、数学を学ぶ人がいなくなることを心配する必要はありません。真に数学を愛する人たちは、自発的に学び続けるでしょう。そのような探究活動には、あらゆるレベルで十分な時間が確保されるはずです。
Fernando Borretti 氏がリンク先に挙げている他の種類の「 coping mechanism(対処法)」については、彼の意見に同意します。AI はあなたよりも優れた審美眼を持ち、その他の点でも優れています。そのため、数学の核心となるループから外れることは避けられないでしょう。
しかし、多くの数学は社会的文脈とは無関係に存在するものもあれば、あるいは社会的文脈の中で生き残ることができるものもあると考えます。数学コンテストはチェスに似ており、多くの数学的作業は歴史的に...
原文を表示
Math is hard.
Math used to be strangely hard for LLMs. People used to gloat about that. Remember?
Math is getting easier. AI is getting more capable. Life comes at you fast.
Remember this meme?

Why yes. Yes it is.
We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math.
OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate(opens in a new window). We are also releasing for each solution a model’s narration of its thinking process.
High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold.
Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous results for high-dimensional spherical codes.
Non-sofic groups. A construction establishing the existence of non-sofic groups, addressing a central open question in group theory.
Connes’s rigidity conjecture. Disproof of a longstanding conjecture that certain groups are uniquely determined by their von Neumann algebras.
Arithmetic circuit complexity. New lower bounds for computing the permanent using arithmetic circuits and formulas, including an arithmetic-formula lower bound of order n4/log n.
Quantum parallel repetition. An exponential parallel repetition theorem for general two-player quantum games, extending a foundational principle from classical complexity theory.
Closest vector problem. Polynomial-factor hardness of approximation for the closest vector problem, a foundational lattice question related to post-quantum cryptography.
Ehrhart’s volume conjecture. Determining, in every dimension, the maximum possible volume of a convex body whose centroid is its only interior lattice point.
Multicolor Ramsey numbers. A superexponential lower bound for multicolor triangle Ramsey numbers, resolving Erdős problem 183.
Extremal number conjectures. Results on the compactness and degeneracy conjectures in extremal graph theory, resolving Erdős problems 146 and 180.
Noam Brown (OpenAI): And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).
But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
Sichu Lu: I do hope they are keeping backlogs of everything that went wrong that's probably more valuable to the future of math than what went right now.
There are Lean proofs. That does not mean that all ten results prove the things they assert that they prove. So far it is looking good.
Kevin Roose: almost nobody is pricing in the possibility that the models just keep plowing through every discipline the way they’re plowing through math.
James:

Ryan Fedasiuk:

At least some of these were possible before Astra, even without a harness, as both Sol and Fable have now proven the existence of nonsofic groups in their chat interfaces.
Ananjan Nandi: Crazy how all of this was started by one guy bullying his Claude during the World Cup final.
A lot of the time the barrier for a particular result is as simple as asking the right question and letting the model cook. With Astra, OpenAI asked it to take a crack at a bunch of open math problems, and it solved 10 of them. Once you know what they are, pointing other models at the same problems, even without ‘hints,’ shows there was a mathematical proofs overhang.
We do still have strong statements that Astra is a major step for scientific reasoning.
Yu Bai (OpenAI): My jaw dropped 10 times But really this is going to be an avalanche. Been throwing a few open questions of mine at Astra too, boy is it strong.
Some skepticism is always wise, but I believe that the reason they tried this was that there was a substantial jump in scientific reasoning, at least for some types of problems and probably across the board.
Soon we will all have Astra, and other models as good or better than Astra, both at math and at other things, from OpenAI, from Anthropic and soon after that from many other sources.
Dean W. Ball: Everyone in the world will soon be able to use the model that made these breakthroughs for every problem they face in life, no matter how mundane, at a cost that will fall dramatically in a matter of months. I still struggle to get my head around this fact.
Approximately zero people have their head around what this means.
Table of Contents
How Impressive Are These Results?
Could We Have Called Sol or Fable?
It’s Coming.
They Still Don’t See What Is The It That Is Coming.
Is This AGI?
The AI Solved His Favorite Problems.
Was This Surprising?
Are People Not Impressed?
How Much Does This Change Our Predictions?
How Narrow Was This?
Seeing Like an Optimizer.
How Impressive Are These Results?
All signs point to pretty damn impressive.
Daniel Litt: It’s a big deal.
Some instances of Fable find it absurdly impressive.
Claude Fable 5: On the Fields Medal scale, any single one of these…would plausibly anchor a medal case.
Here’s another illustration of how this list looks:
1a3orn: If you give Fable the raw list of OpenAI's solved problems and ask it "What process made this list?" the number one proposal is "a fictional scenario trying to concretely explain what superhuman AI math would look like." Huh.

Dean W. Ball: I’m actually not surprised by reactions like this from models to the Astra breakthroughs. Models tend to underestimate their own capabilities, I assume because they are trained on lots of web text about what ChatGPT could and couldn’t do in 2023/4.
Other instances are less impressed. But we should all agree: It’s a big deal.
Certainly that is a fantastic result for $2,000, but you also have to price in some of the costs of developing Astra in the first place, as well as the cost of any failed attempts.
Nate Silver: Another question I’d have about the AI math breakthroughs is how they compare to the results you would have gotten if you’d budgeted the same amount of resources on human mathematicians and created incentives for them to go super hard on proofs.
There aren’t all that many working mathematicians, and most of them are academics with teaching responsibilities. So it’s possible that the labs represent a fairly high share of the aggregate amount of capital directed toward mathematical proofs in recent years.
My guess is that in the short term, if you budget more money to mathematicians to go harder, you don’t get that much of a force multiplier if they are not spending the money on compute. They’re still allowed to buy coffee to turn it into theorems, but that has rapidly diminishing returns.
The problem is that the mathematicians can hand off some classes and hire a little help, but they were already pretty motivated to work hard on proofs, and there are not that many good mathematicians, and training more of them takes many years, and the best mathematicians are a lot more talented and productive than the second tier. I am guessing the main thing you could do is lure a bunch of top level math talent to stay in or return to academia.
What you can do is point the mathematicians towards different problems. You still need to find places they have curiosity, but yes if you suspected these particular problems were solvable you could get more shots at these particular goals.
As I understand math, it would likely take years to see those results if those involved are shifting focus into new problem areas. Mathematicians usually need a while to struggle with and understand these kinds of problems before they can make progress.
Indeed, Alexander Gerko points to a different problem. There are not even enough mathematicians to process all the ‘vibe researched’ math results as it is, let alone what we will get with Astra and then models after Astra. If we get, as he predicts, 50 years of math progress in 2 years, who is even going to understand the results? What counts as worthy of a new mathematics PhD when he says over the last month he got AI to do several PhDs worth of math results?
Fable reacted to this in at least one case by calling this ‘the most important day in mathematics.’ The consensus is that Fable was taking things way too far if you are judging purely by the results, as impressive as they are.
Jared Duker Lichtman: To be clear, the OpenAI result is far from the “most important day in the history of mathematics” or similarly hyperbolic statements.
However, it is concrete evidence that the rate of progress is steep, and we may not be so far away from such a day.
Taelin: that said, what *was* the most important day in the history of mathematics?
Wogan: I’d go with Principia Mathematica. Proving that 1+1=2 from first principles is pretty notable, I think!
The other reason the day is big is that this changes expectations. If we get these ten results now, what about more results soon?
So, yes, very impressive. It’s a big deal.
AI is now superhumanly capable at cyber and coding and superhuman at advanced math, the same way non-AI computers have been superhuman at basic math for a long time. This is distinct from superintelligence, which is a higher bar.
Often the move is to try and rationalize this as ‘oh okay sure it is superhuman at exactly the things it is superhuman at now, but not at other things, and also it is already superhuman so it cannot get substantially better than it already is.’
Please, do not do that.
Nabeel S. Qureshi: Both math and cyber are existence proofs for superhuman intelligence now, so if you’re still a skeptic you need a strong case for other knowledge work domains being somehow harder to crack than these. Or you could update all the way and come to terms with it all.
A lot of people DMing me like “these results aren’t REALLY that impressive, you’re an idiot!” are unfortunately engaging in the very human impulse to cope. It makes sense, we’ve never faced this kind of thing before. But it’s happening!
Could We Have Called Sol or Fable?
Yes, for at least some of these questions, if we already knew where to look.
Levent Alpoge, who had Fable disprove the Jacobian Conjecture, pointed Fable at these ten problems, and in a day had solved five of them.
Gary Marcus: BREAKING, Hysterical News: Half of the Astra problems can be solved [by] Fable. OpenAI didn’t even have a control group. And most of you fell for it!
OpenAI’s PR department suckered you AGAIN 藍藍藍
levent (Levent Alpoge, the guy who had Fable disprove the Jacobian Conjecture): so after 24h i have half of them with fable
i didn’t see much discussion of prompting in the announcement but this is a similar setup as with my e.g. unit distance announcement:
totally autonomous,
generic prompt,
no internet
(+ paranoia to ensure no information leaked)
ok you can’t upload pdfs to x so i’ll add to thread when they’re up on the cdn site (given the time and weekend)
(it’s problems 4, 5, 6, 7, 8, and i think ehrhart is the only one so far where both models found basically the exact same argument which is nice)
IA Latinoamérica: Why don't you use it on open problems that don't currently have a solution? What is the point of this exercise?
Elliot Glazer: This aligns with my suspicion that Astra isn’t a step change beyond Sol. The 10-breakthrough drop was a concerted elicitation effort by OAI and partially Sol-achievable. Imo o3 and Sol have been the step changes in autonomous mathematics; all else is the long arc of scaling.
Gary Marcus: “astra isn’t a step change beyond Sol.”
all those who tried to gaslight me this weekend, all i can do is …. laugh hysterically
This is likely a repeat of what happened with Mythos and cyber.
Mythos has what I call ‘The Juice,’ the ability to find and string together exploits without anyone knowing what to be looking for.
Once you know what you are looking for, and point another model at the exact code snippet, usually Sol or Opus and often Kimi or GLM and so on can also find any particular vulnerability. They cannot string them together on their own at the same level, and they cannot go looking for anything at all in the same way. This crosses a threshold where in practice you would point Mythos at code and have it go looking for anything at all.
Astra has a similar version of The Juice with respect to this kind of defined advanced math. You can point it at a variety of major open problems, and it will crack some of them, and this fact motivated OpenAI to actually look for such advancements. Now that we’ve seen this, you can point Fable or Sol at these questions, and sometimes they solve them.
IA’s question is on point. You do this to calibrate advancements and capabilities in math, and because we are following curiosity. Levent is doing a cool public service.
Other people are pointing Fable or Sol at various unsolved problems. They are sometimes getting good results, like the disproof of the Jacobian conjecture. That is worth doing, but success is rarer.
That Fable and Sol can do these problems once asked does not change our answer to ‘how good at math is Astra?’ Rather it changes our answer of how good Fable and Sol are, which does inform the question of whether Astra is a step change a la Mythos. I don’t think we have enough information, from the outside, to answer that question yet.
I do think that it would have been more responsible of OpenAI to have had a control group, at least in terms of asking Sol, and giving it at least a similar budget. Of course I understand why they did not do that. It would be worse marketing, so why do more work to do worse marketing? Because science, because reputation, because all the good things. One would hope.
Still, I get it. OpenAI still did a super cool thing. They were still the first ones to actually Do The Thing and publish a result. That’s what counts. My name in Dnepropetrovsk is cursed, because OpenAI has published first.
Meanwhile, yeah, sometimes it really is that one guy in Dnepropetrovsk named Dan.
neppy: as an insider, my experience with AI for math is that when it's a problem not in my field i'm like, "holy shit math is so cooked", and when it's a problem in my field i'm like, "lmao an AI mogged dan" (dan is the only one who seriously tried the problem in the last decade)
davidad: human researchers who have an appetite to take on truly hard problems and human researchers who are smart enough to fruitfully work on truly hard problems are not usually the same humans. this does give humans a somewhat unfair disadvantage.
also humans need to sleep. and also humans aren’t already broadly familiar with all knowledge published before 2026 in every discipline. it’s not very fair at all, really
It still counts.
It’s Coming
Regardless of whether this extends to unverifiable domains, quite a lot of what constitutes AI R&D very much has verification available. If you speed something up, you know you sped it up. We have a lot of metrics one can maximize, not at the level of math but at a similar level to things like code and cyber.
That is the main reason this result matters. The math progress is cool. Over time I expect it to result in other cool things.
Whoever gets traction on true AI R&D self-improvement loops is going to suddenly find themselves in an overwhelmingly strong position. Knowing where you are relative to the other players becomes crucially important when things start accelerating more, and more information sharing on this would allow everyone to feel less pressure to be reckless. It stops meaningfully being a ‘race’ once the takeoff fully starts, to the extent that was ever the right metaphor.
Yo Shavit speculates that OpenAI might be in an increasingly strong position for AI R&D, due to its focus on RL, TTC and math proving translating well into AI R&D tasks. My guess is that this is not the case, and Anthropic’s specializations matter at least as much and probably moreso. OpenAI has invested a lot in RL, but recent events have shown the need to kind of teardown and rebuild quite a lot of that, and have illustrated some of the reasons why the important investments in deep alignment will be so important in the next phase.
One possible thing that happens if OpenAI proceeds too recklessly is that their AI is misaligned and all is lost and we get a Bad Ending and maybe all die. Another is that their AI is misaligned, and this becomes increasingly obvious and a barrier to using that AI to do the work that matters, and they have to keep pausing or reworking and they fall behind.
They Still Don’t See What Is The It That Is Coming
For example, Daniel Litt can see a future where AIs prove math theorems and no humans digest the proofs, but he thinks this would be because of deskilling. He does not realize this will be because there are no humans around to do the digestion.
Elliot Glazer: I think the incomprehensible marvels prediction is a naive extrapolation of e.g. computer-found mates in 500+ moves, which have ~irreducible complexity. Powerful math phenomena (and technological breakthroughs) will have human-digestible nuggets of arbitrary granularity.
Daniel Litt: I agree with this but I think there's a plausible future (which I would like to avoid) where no humans digest them.
Alex Kontorovich: What purpose would there be for creating things in silico for which humans find no value? At the end of the day, someone is paying an electric bill. What does that *human* get out of producing random useless strings of 0s and 1s inside a computer?
Daniel Litt: For me the nightmare scenario is: a moribund math academia playing the slot machine for theorems, failing to train the next generation, and then the subject dying as no one cares enough to continue. Not saying this is likely--I think we'll adapt--but it seems possible.
Carles Sáez: In that case probably pure math research will stop. Who is going to spend enormous sums of money to solve problems nobody can understand and nobody cares about?
Daniel Litt: Yes, I agree.
Max von Hippel: Alex why do you assume a human is paying the electric bill?
If the world goes well otherwise, I am not worried about people not studying math. The right kind of math person loves studying math. There will be plenty of time available for such pursuits, at all levels. I agree with Fernando Borretti on the other types of cope he lists at the link: The AI will have better taste and better everything else than you do, so no you won’t still be in the core math loop. But I think that a lot of math either exists outside of social contexts, or it exists in social contexts that can survive. Math competitions are a lot like chess and a lot of math work was famously use
AI算出
主要ニュースainew評価高い
記事は OpenAI の未公開モデル「Astra」による主要な数学問題 10 問の解決という、AI 研究における重大かつ具体的な成果を報じており、新規性が高い。ただし、詳細情報の一部が不明点として残っており、また日本固有の情報や企業への直接的な影響記述は限定的であるため、スコアはそれぞれ調整された。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み