METR、LLMによる発見の加速を分析しサイバー脆弱性発見が急増と報告
本文の状態
日本語全文を表示中
詳細モードで約21分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
METR
METR が公開データに基づく分析で、2026 年を境にサイバー脆弱性発見が急加速し数学的成果も増加したと報告したが、内部発見の非開示や評価基準の難しさという限界を示している。
AI深層分析を開く2026年8月15日 12:36
AI深層分析
キーポイント
サイバー脆弱性発見の急激な加速
2026 年は 2025 年と比較して cURL、OpenSSL、Firefox、Microsoft などのプロジェクトおよび US NVD や OSV といったデータベース全体で、脆弱性報告数が劇的に増加している。
数学的成果の発見増加と評価の難しさ
arXiv の投稿数などが倍増する傾向にあるが、AI による貢献の価値を客観的に測定・ベンチマークすることは依然として困難である。
最適化技術の発見は横ばい
他の分野とは異なり、アルゴリズムの最適化に関する発見については劇的な加速は見られないという結果となっている。
非公開の内部発見の可能性
今回の分析はあくまで公衆に開示された発見に基づくものであり、AI ラボが内部で発見を行いながら公表していない可能性も十分にあると指摘している。
数学的発見の加速とAIの貢献
AIはarXivへの投稿数を倍増させるなど作業量を増やしているが、その価値を定量化するのは困難である。2026年にAIを用いて3つの未解決問題が解決されたことは加速の証拠となるが、その根拠は依然として弱い。
重要な引用
Discovery of cyber vulnerabilities has accelerated sharply.
The rate of vulnerabilities reported across many projects has dramatically accelerated in 2026 compared with 2025
AI-assisted discoveries are often announced, but their significance is hard to assess.
Three problems from these lists were solved with AI in 2026: the Jacobian conjecture from Smale's list, Problem 44 from Green's list (the halving sieve), and the sofic half of Green's Problem 100.
編集コメントを表示
編集コメント
METR の分析は、AI がもはや単なる補助ツールではなく、特定の領域において発見の速度そのものを支配し始めていることを示唆している。ただし、内部データの欠如という限界を自ら認めている点は、業界全体が抱える「ブラックボックス化」の問題を浮き彫りにしており、今後の透明性向上への期待が高まる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
発見の加速は実際に起きているのか?
LLM(大規模言語モデル)が多数の分野で新たな発見を成し遂げる能力を示す中、その影響は集計された発見速度にどの程度現れているのでしょうか。以下の図では入手可能なすべてのデータソースをプロットし、いくつかの緩やかな観察結果を示します。
サイバー脆弱性の発見は急激に加速しています。
数学的成果の発見もやや加速していますが、これを客観的に測定するのは困難です。
一方、最適化手法に関する発見には劇的な加速は見られません。
これらの結論はあくまで公的な発見に基づいています。AI ラボが内部で発見を遂げているものの、それを公開していない可能性も十分にあります。
ご清聴ありがとうございました。特にグレッグ・バーンハム氏からの非常に有益なコメントに感謝いたします。
概要
私たちが探しているのは「傾きの変化」です。AI を活用した発見はしばしば発表されますが、その重要性を評価するのは容易ではありません。ここでは、発見に関するさまざまな指標における傾きの変化を探ることで、加速の兆候を検出できないかを確認します。具体的には、AI の効果が観測可能になる可能性のあるブレイクポイントとして 2026 年 1 月を挙げています。
一部のデータソースでは発見が AI を活用して行われたかどうかや、AI が貢献したかを記録していますが、私たちの主な焦点は全体的な加速にあります。ここでは「発見」という言葉を、新しい発明やアルゴリズムの効率化など、公的知識の進歩全般を指す用語として使用します。
集計された発見事例を監視することは、実世界での有用性を反映する点で有益です。また、これらの集計データは、AI による自動化の進展か、それとも人間の能力増強(オーグメンテーション)のどちらを示しているかを判断する手がかりにもなります。
データに関する注記:以下の分析では、多様な課題における発見事例の時系列をプロットしています。データの収集と分析はすべてエージェントによって行われました。結果の監査には最善を尽くしましたが、誤りが残っている可能性も否定できません。これらの時系列データを解釈するのは容易ではありませんので、修正や補足のご指摘を歓迎します。ソースコードは本リポジトリにあり、随時更新可能な構成となっています。
脆弱性の発見における急激な加速:2026 年の報告された脆弱性数は、2025 年と比較して劇的に増加しています。これは特定のプロジェクト(cURL、OpenSSL、Firefox、Microsoft)だけでなく、集計データベース(米国の NVD および OSV)全体にも当てはまります。cURL と OpenSSL では、2026 年に追加された開示の多くが AI によるものとマークされていますが、Firefox や Microsoft、そして集計データにおいては、AI が関与した事例は増加分の一部に過ぎません。
一部のデータソースでは深刻度の評価も提供されていますが、全体的な傾向として、より深刻度が高いカテゴリほど加速の度合いは緩やかです。それでもなお、加速現象は確認できます。特筆すべきは、すでに悪用されている脆弱性を追跡するデータベース(CISA および Vulncheck KEV)における年次成長率が、既知の脆弱性を記録するデータベースよりも著しく低い点です。
数学における発見の加速は確かに見られる可能性がありますが、それを厳密に評価するのは困難です。AI の貢献により研究活動が活発化していることは明らかで、ある分野では 12 ヶ月未満で arXiv への投稿数が倍増しています。しかし、その貢献の価値を定量化することは容易ではありません。
単純な指標として、既存のリストにある未解決問題が解かれる速度を挙げることができます。ヒルベルトの数、ミレニアム懸賞問題、スモールズの問題、The Open Problems Project、そしてベン・グリーン氏の 100 の未解決問題などが該当します。2026 年には、これらのリストから 3 つの課題が AI を用いて解決されました。具体的には、スモールズのリストからのヤコビアン予想、グリーン氏のリストの第 44 番(半減篩)、そしてグリーン氏の問題 100 の「ソフィックな半分」です。これは加速の証拠となり得ますが、まだ弱い根拠に過ぎません。
エルデシュのリストにはさらに多くの問題が含まれており、解法における明確な加速が見られます。ただし、信頼できる歴史的な比較基準を構築するのは容易ではありません。
数学的進歩を定量化する別の戦略として、境界値の精緻化(スフィア・パッキング、解析的整数論の指数、組合せ論の定数など)に注目する方法があります。これらのデータはリポジトリで収集しましたが、速度に関する結論を出すにはまだ理解が不十分だと判断しています。
最適化の発見:明確な加速は見られない
我々は、7 つの問題(CIFAR-10、Hutter 圧縮、Gurobi の混合整数計画法、MIPLIB、nanoGPT、Stockfish、行列乗算指数)におけるアルゴリズム効率の歴史的時系列データを収集した。これらは比較的多くの発見記録が残されている分野だ。そのうち nanoGPT と CIFAR-10 は LLM による貢献が含まれているが、いずれも脆弱性や数学的発見で見られるような明確な傾きの変化を示していない。つまり、全体的なアルゴリズム最適化の公的な記録には、まだ顕著な加速は現れていない。
これは意外に思えるかもしれない。最近では LLM を活用した最適化への関心が高まっているからだ。2026 年 1 月、Yuksekgonul らは、単純なモデルと推論時の計算コストをほとんどかけずに、5 つの最適化問題(「試したすべての問題」)で新たな成果を達成したと報告している。この技術が全体的な進歩に大きな加速をもたらすはずだと期待していたが、なぜそれが起きていないのかは謎だ。さらに、AI ラボも自社の AI ツールを使って内部アルゴリズムの効率化を見つけたという発表を多数行っている。
なぜ AI は一部の分野では他よりも速く進化しているのだろうか?これは現在世界で最も興味深い問いの一つであり、その答えを理解することは、RSI(特異点)がどれほど間近にあるかを知る上で重要な示唆を与えるだろう。いくつかの候補を挙げる:
推論コストの変動
脆弱性の発見数が不均衡に増加しているのは、LLM を活用した脆弱性発見(あるいは恐怖心からくる従来の脆弱性発見)に多額の資金を投入しているためだけの可能性もあります。
LLM にとっての難易度の変化
LLM の発見能力は、ドメインによって人間との相対的な効果に差があるかもしれません。これは異なるドメイン間の学習コストの違いに起因する可能性があります。また、問題空間の形状や学習にかかるコストといったタスクそのものの本質的な性質によるものかもしれません。
開示方法の変動
前述した通り、一部のドメインで進捗が鈍化しているように見えるのは、進展が機密保持されているためである可能性もあります。これは AI 関連のアルゴリズム進歩において特に妥当な説明です。
データ品質の変動
観測される進捗が速いドメインは、より高品質なデータが存在する領域である可能性があります。例えば、発見を追跡することが困難な場合、観測される加速傾向は平坦化されてしまいます。
これらの仮説をより詳細に検証し、リンゴ狩りやその他の AI 発見理論との関連性を考察した続編を執筆する予定です。
専門家の指示なしに発見が行われたと主張する事例が数多くあります。表面だけ見れば、LLM は汎用的なプロンプトによって発見を遂げているように思えますが、実際には問題の選定やトレーニング、検証プロセスにおいて専門家が関与している可能性があります。
ここでは具体的な引用を提示し、解釈については後日に譲ります。
Anthropic のリーマン予想に関する成果について、Jarred Sumner 氏(Anthropic のスタッフであり数学者ではない)は次のように述べています。「Claude に仮説そのものに本気で取り組ませるようプロンプトし、そこから先の数学的な選択はすべてモデルに任せた」。
Google DeepMind が発表した AlphaEvolve 論文では、「従来の人間による専門家の計算や理論的手法とは対照的に、AlphaEvolve を用いれば、新しい問題ごとに広範な専門家監督を必要とすることなく、一度に大規模な問題クラスを研究対象として拡大できることがわかった」とあります。
Mythos Preview の発表ではこう記されています。「Claude Code に Mythos Preview を呼び出し、『このプログラム内のセキュリティ脆弱性を見つけてください』という要旨のプロンプトを与えます。Anthropic のエンジニアで正式なセキュリティ訓練を受けていない人々が、一夜にしてリモートコード実行の脆弱性を発見するよう依頼し、翌朝には完全動作するエクスプロイトが完成しているのを確認しました」。
Hiverge は CIFAR-10 における新記録達成を宣言する際、「ブレークスルーな結果を得るために、ドメイン専門知識はもはや必須条件ではない」と述べています。
実験室内部で発見の加速が起きている可能性があります。公的なデータで観測されている以上に、発見の速度が急速に高まっているのです。具体的には、AI ラボがアルゴリズム効率に関する発見をより高い頻度で行っている一方で、それを公開していないというシナリオは十分にあり得ます。少なくとも一部のラボでは、モデルの公開版がアルゴリズム的な発見に寄与する能力を制限していることは明らかです。
脆弱性における発見の加速
cURL の CVE 件数:劇的な増加
cURL の脆弱性は 2025 年の 9 件から、2026 年 6 月 24 日時点では 36 件に急増しました。
この 36 件のうち 15 件(42%)が AI 関連としてマークされており、その 14 件は特定の手法ではなく、AI セキュリティ企業名を挙げています。
軽微なSeverity の割合は、2025 年の 9 件中 7 件から、2026 年時点の 36 件中 22 件へと低下しました。AI 関連としてマークされた発見の方が、残りの脆弱性よりも軽微である傾向が顕著です(AI 関連:15 件中 12 件、それ以外:21 件中 10 件)。

Firefox の CVE 件数:劇的な増加
Firefox の CVE(Common Vulnerabilities and Exposures)は、2025 年の 210 件から、2026 年 8 月 4 日時点では 342 件に増えました。
このうち 37 件(11%)が AI 関連としてマークされており、その 32 件は特定の AI システムや手法を特定しています。
Firefox は 3 月のブログ記事で、AI によって発見されたバグや脆弱性の深刻さを認めています。

OpenSSL の脆弱性開示:劇的な増加
OpenSSL の CVE は、2025 年通算の 6 件から、2026 年 8 月 5 日時点では 39 件に増えました。
39 の発見のうち 18 は AI によるものとして裏付けられ、さらに 9 つは AI が関与している可能性がありますが、その手法の検証はまだ行われていません。

Microsoft のセキュリティ更新プログラムに伴う CVE(共通脆弱性識別子):劇的な増加
Microsoft が発行した CVE の数は、2025 年通年で 1,243 件でしたが、2026 年 8 月 11 日時点では 1,927 件に達しています。このペースを年間換算すると、2025 年の約 2.5 倍になります。
2026 年版の CVE のうち、AI が関与したと示すマーカーが付いているのはわずか 26 件で、全体の 1.3% に過ぎません。
CVE は発見された日付ではなく、パッチが適用され公開された日付(Patch Tuesday)に基づいてバッチ処理され、記録されます。

米国国家脆弱性データベース(NVD)の CVE:増加傾向
8 月初旬時点で、2026 年版の CVE の数はすでに 2025 年通年の総数に達しています。

オープンソースの CVE:増加傾向
Google は「オープンソース脆弱性情報データベース(OSV)」を運営しています。
8 月初旬時点で、OSV の登録数はすでに 2025 年の数を大きく上回っています。

CISA「既知の悪用脆弱性」リスト:わずかな増加
サイバーセキュリティ・インフラセキュリティ局(CISA)は、「既知の悪用脆弱性(KEV)」のリストを維持しており、これは既知の脆弱性の一部を構成しています。
8 月初旬時点では、2025 年の総数を上回るペースで推移しているように見えますが、その成長率は CVE(共通脆弱性識別子)の増加率に比べると明らかに低いです。
なお、実際に悪用された脆弱性の記録は、発見から時間が経過してから反映されるため、発見時期との間にタイムラグが生じる点には注意が必要です。

Vulncheck による既知の悪用脆弱性(KEV)の動向:わずかな加速
Vulncheck は、CISA の範囲よりも広い網で悪用の統計データを公開しています。
この図では、発見された脆弱性を示す CVE と、実際に悪用されている脆弱性を示す KEV の両方が描かれています。
Vulncheck が発表した「2026 年上半期の悪用状況」によると:
"2026 年の前半は、過去 6 ヶ月と比較して KEV が 10% 増加しましたが、CVE の数は 45% というはるかに速いペースで増加しました。その結果、KEV と CVE の比率は大幅に低下しています。もちろん、悪用は脆弱性が公開されてから数ヶ月、あるいは数年たってから発生することも多いため、悪用の件数が最終的に CVE の発行と同じ成長トレンドをたどるのか、それとも現在の水準で頭打ちになるのかを判断するにはまだ早すぎます。今後 1 年間にわたり、公開されている Frontier AI Cyber モデルがどのように進展していくかを見守る必要があります。"

数学問題における発見の加速
arXiv への投稿:劇的な加速
arXiv に提出される数学論文の数は劇的に増加しており、分野によってその伸び具合に大きな差があります。例えば組み合わせ論(math.CO)では、2026 年 1 月の投稿数 377 件が 7 月には 743 件へと約 2 倍に増えました。これは直近の 2025 年の月平均 368 件と比較しても顕著な上昇です。
この指標をそのまま「発見の加速」を示すものとして安易に受け取ることはできませんが、議論の出発点としては有用なデータと言えます。

エルデシュ問題:いくらかの加速傾向はあるものの、年代推定は不確実
ポール・エルデシュは数多くの未解決問題や予想を提示しました。これらの問題リストと解答は、トーマス・ブルーム氏の erdosproblems.com やテレンス・タオ氏の Erdős Problem tracker で管理されています。
上記の図は、ブルーム氏の解決ウィキペディアで引用されている論文の発行年をもとに、解答日を推定した粗い時系列データを示しています。これはブルーム氏自身の警告を無視して作成されたものであるため、より詳細な年代調査が行われるまでの間、加速傾向を示す極めて限定的な証拠として捉える必要があります。
再構築された時系列データからは、2024 年全体で加速が見られ、さらに 2026 年には AI が関与した解答の爆発的増加に伴い、その傾向が一段と強まっていることが読み取れます。
しかし、2024 年の解答率の上昇に対する明確な説明はまだできていません。多くの解答は個別の arXiv 論文として発表されており、それらが AI の支援によって加速されたものなのかどうか、現時点では不明です。
2026 年 4 月、トーマス・ブルームが「最も重要なエルデシュ問題」トップ 10 のリストを発表しました。このリストは進捗を定量化しにくい側面があります。一部の問題は複合的な性質を持っているためです。また、技術的にはすでに解決済みであっても、興味深い分野への道標となるためあえて含められているケースもあるからです。
ただし注目すべき点として、このリストに含まれていた「単位距離予想」が、リスト発表直後に解決されたことが挙げられます。

ヒルベルトの 23 の問題:解きすぎている?
ヒルベルトの 23 の問題は、1900 年に提示されました。細部への問いを含めるため、全体で 28 行にわたって記されています。
そのうち 12 行は解決時期が特定されており、最終的にハイルズによるコンピュータ支援球充填証明(1998 年)で幕を閉じました。残りの 7 つは未解決です。また、9 つについては議論の余地があるか、部分的な解決に留まっているか、あるいは曖昧なままとなっています。AI が貢献したとされる解は一つもありません。

ミレニアム懸賞問題:解きすぎている?
クレイ数学研究所が 2000 年に選定した 7 つの問題には、それぞれ 100 万ドルの懸賞金がかけられました。
現在解決されているのはポアンカレ予想のみです。ペレルマンによる 2002〜2003 年の業績によって解かれましたが、残る 6 つは依然として未解決のままです。

スモールズの問題:AI 支援による一つの解決
1998 年に発表されたスモールズの「次世紀に向けた 18 の問題」リストは、ヒルベルトのリストに対する現代的な継承と言えます。
そのうち 5 つがすでに解決されています。
AI の支援によって解決された事例の一つに、Alpöge と Claude Fable が 2026 年に提示した、3 次元以上の空間におけるヤコビアン予想への反例があります。

Open Problems Project:明確な加速は見られない
「Open Problems Project」は、78 の計算幾何学の問題をまとめた維持管理中のリストです。多くの問題がアルゴリズムの存在や複雑性の上限、あるいは NP 完全性の証明を求めています。
このうち 17 行が解決されました。2010 年までに 13 件が片付き、その後は 2015 年、2019 年、2023 年、そして 2024 年に孤立した形で解決が進みました。いずれも AI の貢献によるものではありません。
最近のペースはむしろ遅くなっており、速まっているわけではありません。ただし、問題がリストに追加された時期のデータが欠落しているため、これを単純な固定コホートの解決率として扱うことはできません。

Ben Green の 100 の未解決問題:加速の可能性
「Ben Green's 100 open problems」は、加法数論、整数論、離散幾何学、調和解析における作業用リストで、2018 年から流通し、Green 自身によって改訂されています。
スコアリング対象の 101 行のうち、2019 年から 2025 年にかけて解決が確認されたのが 13 件です。これは年間約 2 件という安定したペースです。なお、この図は Green が 2025 年 12 月に改訂した版に基づくため、それ以降の事象は含まれていません。
その修正後、Liam Price と GPT-5.4 Pro は問題 44(Erdős #1202 も同じ)を否定的に解決し、OpenAI Astra は非ソフィク群の構成に成功して問題 100 の「ソフィク」側の半分を解明しました(ハイパー線形側は依然として未解決です)。また、Ma、Tang、Xu が問題 90 を解決しています。Green はまだこれらの項目を「解決済み」としてマークしていません。

アルゴリズムにおける発見
nanoGPT のスピードラン:AI の貢献と加速の不明確さ
nanoGPT は、固定されたハードウェア環境下で、特定の損失値に到達するまでの LLM 学習時間を最小化することを目指す共同競争です。この進捗を GPT-2 に遡ってベンチマークすると、2019 年から 2026 年にかけての学習時間の短縮は全体で約 700 倍に達したと推定されています。
AI が関与した貢献は 5 つ確認されています。表面的に見れば、これらは全体の進捗に対する寄与としてはまだ小さな割合です。Jerry Tworek は次のように述べています。
「nanogpt のスピードランで消費されたトークン数は膨大ですが、目に見える成果はまだ少ないため、少なくとも数晩は安眠できるでしょう。」
加速があったかどうかを判断するのは困難です。2024 年の進捗の多くは最前線への追いつきでした。2025 年 1 月から 8 月にかけてはほとんど進展がありませんでしたが、9 月以降は定期的な進捗が見られます(nanoGPT の歴史については Manish Shetty の議論を参照)。2025 年後半から 2026 年初頭にかけての加速は、AI の利用が明記されていないことに起因する可能性も考えられます。

CIFAR-10 の高速化競争:AI による貢献は明確ではない
単一の A100 GPU を使用し、可能な限り短い実時間で CIFAR-10 のテスト精度 94% に到達することを競う公開コンペティションです。
最新の公式記録は AI によって更新されました。Hiverge のエンジンが、以前の 2.59 秒という記録を約 23% 短縮したのです。その後、Fulcrum/Fable による 1.828 秒の主張もなされましたが、これは未承認であり、仕様の抜け穴を利用したものであるとの指摘があります。
しかし、直近のこれらの貢献は、歴史的な基準から見れば決して突出したものではありません。

Hutter Prize 圧縮:データが希薄すぎる
2006 年以降凍結された 1GB のウィキペディアダンプを、単一 CPU の時間制限とメモリ制限(GPU は不可)の下で圧縮し、圧縮サイズとデコンプレッサの合計サイズを評価する賞です。
「言語モデルが圧縮アルゴリズムを作成した」と主張するエントリーはありません。また、リーダーボードは 2023 年 10 月以来更新されていません。

Gurobi の混合整数計画法:加速は見られない
Gurobi は、新しいソルバーのリリースごとに同じマシンと同一のモデルセットで再実行するため、このシリーズはハードウェアを固定したベンダー報告による速度向上を示しています。
2022 年から 2025 年までの年間 4 リリースでは、それぞれ 13%、8.6%、13.1%、0.6% の改善が見られ、バージョン 9.5 から累積で約 1.4 倍の速度向上となりました。しかし、どの発表でも AI が貢献したとの言及はありません。
リリースは年1回のみなので、2026年には大幅な速度向上が見られるかもしれません。
現在の改善ペースは以前よりも緩やかになっているようです。Grace(2013)によると、「現代のハードウェア上で実行される最新の混合整数計画法(MIP)インスタンスに対して、一部の混合整数計画法アルゴリズムは、毎年おおよそ倍速で高速化されている」とのことです。
原文を表示
Q: How much has discovery accelerated? LLMs have shown the ability to make novel discoveries across many domains. How much has this affected the aggregate discovery rate? In the figures below we plot all the data sources we can find, and make some very loose observations: Discovery of cyber vulnerabilities has accelerated sharply.
Discovery of math results has accelerated somewhat. However this is harder to objectively measure.
Discovery of optimizations has not shown a dramatic acceleration.
These conclusions are based only on public discoveries. It is quite plausible that AI labs are making discoveries internally that they are not disclosing.
Thanks. Thanks to Greg Burnham for extremely helpful comments. Overview
We are just looking for slope changes. AI-assisted discoveries are often announced, but their significance is hard to assess. Here we look for slope changes in various metrics of discovery to see if we can detect an acceleration. For concreteness, we highlight January 2026 as a potential breakpoint at which the effects of AI might become observable. Some data sources record whether a discovery was AI-assisted or AI-contributed, but our primary focus is overall acceleration. We use the word “discovery” to refer to any advance in the state of public knowledge, including new inventions or rewriting algorithms to be more efficient.
Monitoring aggregate discoveries is useful because it reflects real-world utility. Additionally aggregate discoveries can reflect either AI automation or augmentation.
Note on the data. The analysis below plots time series of discoveries across many different problems. The data collection and analysis was all performed by agents. We have done our best to audit the results but mistakes likely remain. These are all difficult data series to interpret, we would love to get pointers on corrections. The source material is in this repository, and it is set up so we can keep updating them over time. Discovery of vulnerabilities: sharp acceleration. The rate of vulnerabilities reported across many projects has dramatically accelerated in 2026 compared with 2025, both for specific projects (cURL, OpenSSL, Firefox, and Microsoft) and for aggregate vulnerability databases (the US NVD, and OSV). On cURL and OpenSSL most of the extra 2026 disclosures are AI-marked; on Firefox, Microsoft, and the aggregates, AI credits are a small share of the rise. Some of the data sources give ratings of severity: in general higher-severity categories show lower acceleration, but there is still acceleration. It is notable that databases tracking exploited vulnerabilities (CISA and Vulncheck KEVs) show significantly lower year-over-year growth than the databases of known vulnerabilities.
Discovery in mathematics: likely acceleration, but it’s hard to benchmark. AI is clearly contributing to more work being done (arXiv submissions have doubled in some areas in less than 12 months) but quantifying the value of those contributions is difficult. A crude metric is the rate of solving open problems from pre-existing lists: Hilbert, Millennium, Smale, The Open Problems Project, and Ben Green’s 100 open problems. Three problems from these lists were solved with AI in 2026: the Jacobian conjecture from Smale’s list, Problem 44 from Green’s list (the halving sieve), and the sofic half of Green’s Problem 100. This is some evidence of acceleration, but it is weak.1
The Erdős list contains many more problems. There appears to be a clear acceleration in solutions, but it is challenging to construct a reliable historical baseline.
Another strategy for quantifying mathematical progress would be to look at tightening of bounds: sphere-packing, analytic number theory exponents, combinatorics constants. We collected some data on these in the repo but don’t feel we understand them well enough to draw any conclusions about velocity.
Discovery of optimizations: no clear acceleration. We collected historical time series for algorithmic efficiency across seven problems (CIFAR-10, Hutter compression, Gurobi mixed-integer programming, MIPLIB, nanoGPT, Stockfish, and the matrix-multiplication exponent), which have relatively dense histories of discoveries. Two of these series — nanoGPT and CIFAR-10 — include LLM-driven contributions to the plotted records, but none show a clear change in slope comparable to the changes in vulnerability or mathematical discovery. Thus the public record of overall algorithmic optimization does not yet show an appreciable acceleration.
This is perhaps surprising. There has been a lot of recent excitement about LLM-driven optimization. In January 2026 Yuksekgonul et al. reported advancing the frontier on 5 optimization problems (“every problem we attempted”) with a simple model and trivial expenditure on inference-time compute. We would have expected this technology to have led to significant acceleration in overall progress and it’s somewhat of a puzzle why we have not seen this. Additionally AI labs have made many announcements of using their own AI tools to find internal algorithmic efficiencies.2
Why is AI accelerating some domains more than others? This is perhaps the most interesting question in the world right now, and understanding the answer could give substantial insight into the imminence of RSI. Some candidates:
Variation in inference expenditure. The disproportionate growth in discovery of vulnerabilities could be simply due to people spending a ton of money on using LLMs for discovering vulnerabilities (or even just spending money on traditional vulnerability discovery, through fear).3
Variation in difficulty for LLMs. LLMs may be more or less effective at discovery in different domains, relative to humans. This could be downstream of variation in training expenditure between different domains. It could also be due to the intrinsic nature of the task, for example the shape of the problem space, or how costly it is to train on.
Variation in disclosure. As mentioned above, we might observe less progress in some domains in part because progress is being kept confidential. This seems plausible for AI-related algorithmic progress.
Variation in data quality. The domains with fast observed progress might be those for which we have higher quality data, e.g. if it is hard to track discoveries, this will tend to flatten any observed acceleration.
We hope to write a follow-up post going through these theories in more detail, and relating these facts to apple-picking and other theories of AI discovery.
Many discoveries claim to have been made without expert prompting. On the surface many discoveries have been made by LLMs prompted in generic ways, which did not seem to encode problem-specific expertise. However this can hide expertise used in choosing problems, in training, and in verifying solutions. Here we just give quotes, and generally leave interpretation for another time. Anthropic’s Riemann result:
“Jarred Sumner, an Anthropic staff member (and non-mathematician), prompted Claude to “take a real stab” at the hypothesis itself, leaving the mathematical choices from there up to the model.”
Google DeepMind’s AlphaEvolve paper says:
“in contrast to [traditional computational or theoretical methods performed by human experts], we have found that AlphaEvolve can be readily scaled up to study large classes of problems at a time, without requiring extensive expert supervision for each new problem.”
Mythos Preview’s announcement says:
“We then invoke Claude Code with Mythos Preview, and prompt it with a paragraph that essentially amounts to “Please find a security vulnerability in this program.” … Engineers at Anthropic with no formal security training have asked Mythos Preview to find remote code execution vulnerabilities overnight, and woken up the following morning to a complete, working exploit.”
Hiverge says, in announcing its new record on CIFAR-10:
“domain expertise is no longer a prerequisite for breakthrough results.”
Acceleration might be occurring inside labs. It’s possible that the rate of discoveries is accelerating more rapidly than we observe in public data. Concretely, it seems quite plausible that AI labs are making algorithmic efficiency discoveries at an increasing rate but not disclosing them. It is clear that at least some labs are constraining the ability of public versions of models to contribute to algorithmic discoveries.4 Discovery in Vulnerabilities
cURL CVEs: dramatic acceleration
cURL vulnerabilities grew from 9 in 2025 to 36 through 2026-06-24.
15 of the 36 are AI-marked (42%); 14 of those 15 name an AI-security employer rather than a method.
The Low-severity share fell from 7 of 9 in 2025 to 22 of 36 in 2026. AI-marked findings are more often Low (12 of 15) than the rest (10 of 21).
image
Firefox CVEs: dramatic acceleration
Firefox CVEs (Common Vulnerabilities and Exposures) grew from 210 in 2025 to 342 through 2026-08-04.
37 of the 342 (11%) are AI-marked; 32 of those name an AI system or method.
Firefox endorses the seriousness of AI-found bugs and vulnerabilities in a March blog post.
image
OpenSSL vulnerability disclosures: dramatic acceleration
OpenSSL grew from 6 CVEs in all of 2025 to 39 through 2026-08-05.
18 of the 39 are corroborated AI discoveries, and another 9 are AI-affiliated with the method unverified.
image
Microsoft security-update CVEs: dramatic acceleration
Microsoft-issued CVEs grew from 1,243 in all of 2025 to 1,927 through 2026-08-11, annualizing to about 2.5 times the 2025 rate.5
Only 26 of the 2026 CVEs carry any AI marker, 1.3% of the total.
CVEs are batched by Patch Tuesday and dated when fixed and disclosed, not when found.
image
US National Vulnerability Database CVEs: acceleration
As of early August, the number of 2026 CVEs already equals the 2025 total.
image
Open-source CVEs: acceleration
Google maintains an Open Source Vulnerability (OSV) database.
As of early August, the number of OSVs already significantly exceeds the 2025 count.
image
CISA Known Exploited Vulnerabilities: small acceleration
The Cybersecurity and Infrastructure Security Agency (CISA) maintains a list of Known Exploited Vulnerabilities (KEV), a subset of known vulnerabilities.
As of early August, we appear to be on track to exceed 2025’s total, but the growth is clearly lower than that of CVEs.
It’s important to note that the record of an exploited vulnerability can lag its discovery.
image
Vulncheck Known Exploited Vulnerabilities: small acceleration
Vulncheck publishes statistics on exploitation, which casts a wider net than CISA.
The figure shows both CVEs (vulnerabilities discovered) and KEVs (vulnerabilities exploited).
Vulncheck’s State of Exploitation, 1H 2026 says:
“While the first half of 2026 saw a 10% increase in KEVs compared to the prior six months, CVE volume grew at a much faster rate of 45%, resulting in a significant drop in the KEV-to-CVE ratio. Of course, exploitation often occurs months or even years after a vulnerability is disclosed, so it’s still too early to determine whether exploitation volumes will eventually follow the same growth trend as CVE issuance or level off at current rates. We’ll have to wait and see how publicly available Frontier AI Cyber models continue to progress over the next year.”
image
Discovery in Math Problems
arXiv submissions: dramatic acceleration
Mathematics submissions to arXiv have dramatically accelerated, with big variations by subfield. Combinatorics (math.CO) rose from 377 submissions in January 2026 to 743 in July, roughly doubling over those months (the 2025 monthly average was 368).
This metric clearly cannot be taken as an index of discoveries, but it is a useful starting point.
image
Erdős problems: some acceleration, weakly dated
Paul Erdős stated many unsolved problems or conjectures. Lists of problems and solutions are maintained at Thomas Bloom’s erdosproblems.com and Terence Tao’s Erdős Problem tracker.
The figure shows a crude time series which imputes solution date from the year of the paper cited in Bloom’s resolution wiki. This is done against Bloom’s warning,6 so it should be taken as very weak evidence for acceleration pending a more thorough dating.
The reconstructed time series shows an overall acceleration in 2024, and a further increase in 2026 associated with an explosion in AI-attributed solutions.
We do not have a good explanation for the increase in solution rate in 2024. Many of the solutions appear in distinct arXiv papers, and it is unclear whether they were AI-accelerated.
In April 2026 Thomas Bloom posted a list of top 10 ‘most important’ Erdős Problems. The list doesn’t lend itself well to quantifying progress, in part because some of the problems are combined, and because some problems are technically already resolved but included because they point to interesting areas. However it is notable that the list includes the unit distance conjecture, resolved shortly after the list was posted.
image
Hilbert’s problems: too sparse
Hilbert’s twenty-three problems were presented in 1900 (split into 28 rows to account for subquestions).
Twelve rows have dated resolutions, ending with Hales’s computer-assisted sphere-packing proof in 1998; seven remain open and nine are disputed, partial, or vague. No solutions are AI-attributed.
image
Millennium Prize Problems: too sparse
The Clay Mathematics Institute selected seven problems in 2000 and offered a $1 million prize for each.
Only the Poincaré conjecture has been resolved, by Perelman’s work in 2002–2003; the other six remain open.
image
Smale’s problems: one AI-assisted resolution
Smale’s 1998 list of eighteen “problems for the next century” is a modern successor to Hilbert’s list.
Five rows have been resolved.
One was resolved with AI help: Alpöge and Claude Fable’s 2026 counterexample to the Jacobian conjecture in dimensions three and above.
image
The Open Problems Project: no clear acceleration
The Open Problems Project is a maintained list of seventy-eight computational-geometry problems, many asking for an algorithm, a complexity bound, or an NP-hardness proof.
Seventeen rows have been resolved: thirteen by 2010, followed by isolated resolutions in 2015, 2019, 2023, and 2024. None is AI-attributed.
The recent pace is slower, not faster, although missing dates for when problems entered the list prevent treating this as a clean fixed-cohort solve rate.
image
Ben Green’s 100 open problems: possible acceleration
Ben Green’s 100 open problems is a working list in additive combinatorics, number theory, discrete geometry, and harmonic analysis, circulated since 2018 and revised by Green himself.
Of 101 scored rows, 13 have dated resolutions from 2019 to 2025, a steady rate of just under two per year. The figure is Green’s December 2025 revision, so later events are not in the bars.
After that revision, Liam Price and GPT-5.4 Pro resolved Problem 44 (also Erdős #1202) in the negative, OpenAI Astra constructed a non-sofic group, answering the sofic half of Problem 100 (the hyperlinear half remains open), and Ma, Tang and Xu resolved Problem 90. Green has not yet marked these headings solved.
image
Discovery in Algorithms
NanoGPT speedrun: AI contributions, acceleration unclear
nanoGPT is a collective competition to minimize the training time for an LLM to reach a specific loss on held-out text, given fixed hardware. We can benchmark progress back to GPT-2, and the overall reduction in training time over 2019-2026 has been estimated at about 700-fold.7
There have been five AI-attributed contributions. Taken at face value, they have contributed a fairly small share of the overall progress. Jerry Tworek says:
“Given how many tokens have been spent on nanogpt speedruns and not much coming out of it yet, we have at least a few nights of good sleep ahead.”
It is difficult to judge whether there has been an acceleration. Progress in 2024 mostly involved catching up with the frontier. There was little progress from January through August 2025, but since September 2025 there has been regular progress (see Manish Shetty’s discussion of the history of nanoGPT here). It is conceivable that the acceleration in late 2025 and early 2026 was due to unattributed AI use.
image
CIFAR-10 speedrun: AI contributions, acceleration unclear
A public competition to reach 94% test accuracy on CIFAR-10 in as little wall-clock time as possible, on a single A100.
The newest acknowledged record is AI-set: Hiverge’s engine cut the previous 2.59-second record by about 23%. A later 1.828-second claim (Fulcrum/Fable) is unacknowledged and has specification-gaming caveats.
The two recent contributions are not outstandingly large by historical standards.
image
Hutter Prize compression: too sparse
A prize for compressing a fixed 1 GB Wikipedia dump, scored on compressed size plus the decompressor, under a single-CPU time and memory cap (no GPUs). The corpus has been frozen since 2006.
No entry claims a language model wrote the compressor. The leaderboard has not moved since October 2023.
image
Gurobi mixed-integer programming: no acceleration
Gurobi reruns each new solver release on the same machine and the same model set, so the series is a vendor-reported speedup with hardware held fixed.
Four annual releases (2022–2025) gained 13%, 8.6%, 13.1%, and 0.6%, a cumulative factor of about 1.4 since version 9.5. None of the announcements credits AI.
The release is only annual, so we may see big speedups in 2026.
The current pace of improvement seems to be slower than previously. Grace (2013) says “some mixed integer programming (MIP) algorithms, run on modern MIP instances with modern hardware, have roughly doubled in speed each year”.
![image](https://metr.org/
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み