METR、Anthropic研究者の生産性向上を2倍以上と推定
本文の状態
日本語全文を表示中
詳細モードで約25分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
METR
METR の Thomas Kwa は、Anthropic がコード生成量で8倍向上した事実を経済モデルに適用し、研究者の生産性向上率が少なくとも2倍以上であると推定する分析結果を発表した。
AI深層分析を開く2026年8月1日 14:09
AI深層分析
キーポイント
経済モデルによる生産性推定
Cobb-Douglas や CES 関数を用いた分析では、コード出力が8倍になった場合、研究者の総生産性は2.33倍から2.91倍の範囲で向上すると予測される。
推定値を低下させる要因
AI 生成コードの冗長性、実質的な価値が低いコードの増加、研究者の非合理的な行動といった要因により、実際の効果は推計値より低くなる可能性がある。
R&D 全体への影響
Anthropic の Mythos システムカードでは R&D 全体の向上が2倍未満とされるが、これは計算資源の増加が含まれないためであり、研究者単体の生産性向上はさらに大きい。
Cobb-Douglasモデルによる研究者の生産性向上
コーディングに費やす時間が半分と仮定すると、コード出力が8倍になることで研究者の総生産性は約2.83倍向上する。
CESモデルにおける代替性と補完性の影響
コードと非コードの成果物が互いに補完的か代替的かによって、生産性向上の度合いが異なるため、両者の関係性を考慮した分析が必要である。
重要な引用
all models predict researcher uplift at Anthropic from coding agents alone is >2×
Cobb-Douglas predicts that if pre-AI time spent coding is β=50% and code output increases by a factor M=8, then researcher uplift is M^β = 2.83
Anthropic's Mythos system card claims that their overall R&D uplift is 'well short of' 2x
If researchers spend roughly half their time coding (β = 0.5), then U = √8 ≈ 2.83.
編集コメントを表示
編集コメント
この分析は特定の研究者の意見であり、METR 内の他のメンバーは異なる見解を持っている点に注意が必要である。しかし、コード生成ツールの進化が研究プロセスそのものに与える影響を定量的に捉えようとする試みとして、業界全体にとって示唆に富む内容となっている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
注記:本稿のモデル化における前提と結論はトーマス・クワの意見であり、METR の他のメンバーはこれに同意していません。また、数値計算は Claude によって確認されましたが、人間の専門家による再検証は行われていません。
導入
Anthropic が発表した RSI ブログ記事によると、2026 年第 2 四半期には、Anthropic の貢献者が 1 日に処理するコード量が 2021〜2024 年の期間と比較して 8 倍に増加しました。これは、研究者の総体的な有効出力が何倍になったか(シリアルの研究者アップリフト)について、何を意味しているのでしょうか。
もちろん、コード量が増えたからといって研究量がそのまま 8 倍になるわけではありません。コーディングは業務の一部に過ぎないからです。しかし、各コード行(LoC)の品質や記述性が 2025 年以前のものと同等であると仮定し、標準的な経済モデルの前提を適用すれば、驚くべき結論が導き出されます。すべてのモデルが示すところによれば、コーディングエージェント単独による Anthropic の研究者アップリフトは 2 倍以上です。(なお、コード以外のタスクにおける向上分が含まれていないため、実際の数値はこれよりも高くなる可能性があります。)
コブ・ダグラス関数を用いた予測では、AI 導入前にコーディングに費やしていた時間が全体の 50%(β=50%)であり、コード出力が 8 倍(M=8)になった場合、研究者アップリフトは M^β = 2.83 倍と計算されます。
CES(一定弾力性の代替生産関数)を用いた分析では、興味深い数学的な偶然により、コードが均質であると仮定した場合、その推定値は約 [2.75, 2.91] の狭い範囲に収束します。
もしコードが CES 型でありながら非均質である場合、つまり AI が低リスクのコード処理を高リスクのものよりもより劇的に高速化する場合、推定値は [2.33, 2.66] の範囲となります。これは前述の数値より低いものの、依然として 2 倍を超えています。
しかし、新しいコードが古いコードと同等ではない理由はいくつかあり、これらは少なくとも一部は Anthropic の内部データによって解決できる可能性があります。
冗長性:M は、真の品質調整済みコード出力を大幅に過大評価しています(例えば、同じ機能を実現する AI によるコードは、人間によるコードよりも 2 倍以上冗長である可能性があります)。
ほとんど価値のないコード:AI は、AI 以前には決して書かれなかった低リスクなコードにおいて人々の作業速度を劇的に向上させています(20 倍以上)。しかし、これらのコードの価値は通常のコードの 5%〜20% に過ぎません。
研究者の非合理性:Anthropic の研究者たちは、研究価値に寄与しないコードを生産しています(例えば、「雰囲気コーディング」が楽しいという理由から)。
Anthropic の Mythos システムカードでは、全体の R&D 向上率は「2 倍には遠く及ばない」とされています。これは、R&D が労働と計算資源の両方に依存しており、計算資源の増加は数値に含まれないため、研究者の向上率が 2 倍以上であっても矛盾しないことを示しています。私は、全体の R&D 向上率は研究者の向上率が約 3.5 倍になった時点で 2 倍に達するだろうと推測します。これはおそらく来年頃のことでしょう。したがって、ラボはコード出力を単なる行数(LoC)だけで評価するのではなく、より高品質なデータとモデルを取得し、それを研究者の向上率に関連付けることが重要です。
経済モデル
核心的問いは、ある研究者がコード生成量を 8 倍に増やした場合、全体としてどれほどの研究価値を生み出せるかです。これは、思考、執筆、実験、コミュニケーションなど研究者が行う他の活動と比較してコードがどの程度重要かを考えることに依存します。また、これらの活動が互いに代替関係にあるのか、補完関係にあるのかも鍵となります。つまり、ある活動を増やすことが他方の価値を高めるのか、それとも低下させるのかという点です。
これを分析する標準的な経済学的ツールとして「生産関数」があります。これは各投入要素(ここではコード出力と非コード出力)の量を基に、総価値を算出するものです。最も一般的なモデルはコブ・ダグラス型ですが、本稿ではコブ・ダグラス型に加え、それを一般化した CES 型(代替弾力性一定)も検討します。
コブ・ダグラスモデル
ある財の生産量が複数の投入要素に依存する場合、その生産量を予測するための最もシンプルな経済モデルは「コブ=ダグラス関数」と呼ばれます。これは AI Futures Project が労働力と計算資源に対して仮定しているモデルと類似しています。
コブ=ダグラス関数の式によれば、研究出力 Y は以下のように表されます。
Y = q_n^β × q_o^(1-β)
ここで、
β はコーディングに割かれる時間の割合(time share on coding)
q_n と q_o はそれぞれコードと非コードの生産量です。
これに基づき、シリアルの研究者向上率 U を、AI 導入後の出力を AI 導入前の出力で割った比率(U = Y_post / Y_pre)として定義できます。仮に悲観的に、非コード分野での向上は全くない(つまり q_o は一定である)とすれば、以下の関係が導かれます。
研究者がコーディングに費やす時間が全体の約半分(β = 0.5)だと仮定すると、U は √8 ≒ 2.83 となります。β = 0.5 という値は、私が調査した人々が妥当と考える範囲の中央値に近いものです。β の真の値については大きな不確実性がありますが、ここでは簡略化のため一律 0.5 と固定します。より厳密な分析では、このパラメータを変動させるべきでしょう。
CES モデル(定数弾力性の代替モデル)
コブ・ダグラス関数から一般的な CES モデルへ移行すると、投入要素間の代替可能性をモデル化できるようになります。コードと非コードが強い補完関係にある場合(例えば左靴と右靴のように)、コードの出力を倍増させても非コードの出力が増えなければ、全体の生産性はほとんど向上しません。一方、両者が代替関係にある場合(バターとマーガリンのように)、一方の投入量を倍にすれば、ほぼ全体価値も倍になります。AI 研究においてどちらの関係が成り立つかはまだ不明なため、ここでは様々なケースを想定して検証を行います。
設定
ある研究者が、コーディングと非コーディングの 2 つのタスクに固定時間 T を配分します。出力 Y は CES(恒等弾力性)関数に従うと仮定します。
Y = [β q_n^ρ + (1-β) q_o^ρ]^{1/ρ}, ρ = (σ - 1)/σ
ここで、
βは AI 導入前のコーディングに割く時間の割合です(CES 関数では、AI 前後で時間配分が変化する可能性があります)。
q_n と q_o はそれぞれ生成されるコード量と非コード量を表します。
σは、コーディング出力と非コーディング出力の間の代替弾力性を示します。
このとき、研究者の生産性向上率 U(両期間で時間を最適に配分すると仮定)は以下のように計算できます。
U = [β g_n^{σ-1} + (1-β)]^{1/(σ-1)} という式において、g_n は AI がコーディングに提供する 1 時間あたりの速度向上率(コーディング・アップリフト)を指します。
我々は g_n を直接観測することはできません。代わりに観測するのは M = 8 です。これは Anthropic の貢献者によるコード出力の平均増加量です。6 この値が異なるのは、AI によってコーディングのコストが下がると研究者らが時間を再配分するためです。実は、M、U、σ、g_n は比較的シンプルな方程式 M = g_n^σ / U^{σ-1} で関係付けられており、5 M を観測することで U の推定値をかなり頑健に得ることができます。
8 ≈ e² であるため、この推定は σ に依存しにくいことがわかります!
Cobb-Douglas モデルではコードと非コードの代替性が 1(σ=1)であると仮定しました。しかし、これらが補完関係(σ < 1)または代替関係(σ > 1)にある場合どうなるでしょうか?結論はほとんど変わりません。
重要な理由は、コード出力が8倍になったことが、2つの自由パラメータ(コーディングの向上度 g_n とコードの限界価値)の取りうる値を大幅に制限している点にあります。もしコーディングの向上度とコードの限界価値の両方が大きければ、コード生成が容易でかつ極めて価値が高いため、出力は8倍を超えていたはずです。逆に、g_n と σ の両方が小さければ、コード生成コストが高く価値も低くなるため、出力は8倍未満になっていたでしょう。
コード出力が8倍に増加したという事実から導き出せるのは、2つのパラメータの中間的な組み合わせのみです。(a) コーディングの向上度が高くコードの限界価値が低い場合、あるいは (b) コーディングの向上度が低くコードの限界価値が高い場合のいずれかです。
世界 (a) では、コードと非コードは強い補完関係にあります(概ね σ < 0.5)。これは研究者たちがコーディングから時間を切り離し、他のボトルネック解消に注力することを意味します。この状況でコード出力が8倍にとどまっているのは、コード生成の速度向上が極めて大きく、少なくとも23倍であるためです。これにより非コード作業に割ける時間が大幅に確保され、研究者たちは U=2.75 という向上度を達成することができました。
世界 (b) では、コードと非コードがほぼ完全な代替関係にあり(σ=3.0)、研究者は生産性が高く非コードタスクを代替できるため、ほとんどすべての時間をコーディングに費やします。しかし、M=8 倍のコード出力増加があった場合、彼らがコーディングに費やす時間がほぼ倍増したことを考慮すると、1 時間あたりのコード生産量は g_n=4.2 倍にしかなりません。したがって、アップリフトは依然として U=3.07 に留まります。
どちらの場合も、U は M^β = 2.83 と概ね近い値になります。
window.addEventListener('message', function (e) { var iframe = document.getElementById('uplift-explorer'); // Only accept the height message from our own iframe, same-origin, with a // finite numeric payload, and clamp it so a bad value can't blow up layout. if (!iframe || e.source !== iframe.contentWindow) return; if (e.origin !== window.location.origin) return; var h = e.data && e.data.upliftExplorerHeight; if (typeof h !== 'number' || !isFinite(h)) return; iframe.style.height = Math.max(200, Math.min(h, 8000)) + 'px'; });
なぜ M を観測した後に U が σ にあまり影響されないのでしょうか。lnU の式を導出し、σ=1 の周りで摂動すると次のようになります。
7
Δ ln U ≈ β(1-β)-ln M + (ln M)²/2
この式で括弧内の項は、U と M^β の百分率差を決定する部分ですが、M = e² ≒ 7.39 でゼロになります。M = 8 の場合、その値は -2.08 + 2.16 = 0.08 となり、ほぼゼロです。つまり、ln U が σ に対して持つ一次の感度は、σ の単位あたりわずか 0.25 × 0.08 = 0.02 に過ぎません。
結論として、U は σ ∈ [0.5, 2] の範囲で 2.83 から±3% の間にとどまります。
σ が示唆する AI コーディングシェア後の g_n と U の関係
σ | g_n | AI コーディングシェア後 | U
---|---|---|---
0.0 | — | 不可能 (M ≤ 2) | —
0.3 | ≒114 | 3.5% | 2.56
0.5 | 23.3 | 17% | 2.75
0.7 | 12.5 | 32% | 2.80
1.0 | 8.0 | 50% | 2.83
1.5 | 5.7 | 70% | 2.86
2.0 | 4.8 | 83% | 2.91
3.0 | 4.2 | 95% | 3.07
M = 8 は σ の下限も示唆する
M = 8 を観測することは、σ に依存しないアップリフトの推定値を与えますが、同時に σ にも制約を課します。
M = g_n · (t_n^postAI / t_n^preAI) であることを思い出してください。ここで、g_n はコーディングにおける生産性向上(スピードアップ)率、t_n はコーディングに費やされる時間を表します。
σ < 1 の場合、研究者はコーディングから時間をシフトさせます(t_n^postAI / t_n^preAI < 1)。このため、観測された 8 倍のアウトプットを達成するには、時間あたりのスピードアップ率 g_n が M を超える必要があります。σ が小さいほど、g_n はより大きくなければなりません。具体的には σ = 0.3 の場合、g_n ≒ 114 が必要になります。また、σ = 0.3 という値は、Anthropic の研究者が 2026 年第 2 四半期にコーディングに費やす時間が全体の 3.5% に過ぎないことを意味しますが、これは現実的ではありません。
σ = 0(レオンティエフ型、完全補完財)の場合、生産関数は Y = min(β q_n, (1-β) q_o) となり、出力はどちらかの投入要素がボトルネックになります。たとえ g_n が無限大に近づいても、最大で M = 1/(1-β) = 2 までしか伸びず、M = 8 はあり得ません。直感的には、非コード側が厳格なボトルネックである場合、コーディング速度をいくら上げても、総コード出力は非コード出力の最大増加分である 1/(1-β) を超えることはできません。
現在の状況では、時間あたりのコーディング速度向上率が 23 倍未満であり、Anthropic の貢献者がコーディングに費やす時間が全体の 17% よりも大きいと考えるのが妥当です。これは σ ≥ 0.5 を意味し、コーディング研究と非コーディング研究の間に強い補完関係があるという仮説を否定します。この推定はさらに精密化できます。Anthropic の時間配分を実測すれば、それが AI 導入前の配分と似ているなら σ ≈ 1 となり、もしそれより高いか低いかによって、コードへの依存度が増減する方向に応じて、σ が前節で述べた理由により低下または上昇します。
コードの多様性モデル:AI による高速化が、リスクの高いコードでは鈍化する可能性は?
上記のモデルは、すべてのコードが均一に高速化されると仮定しています。しかし実際には、AI による速度向上効果は、リスクの低いコードにおいてより顕著である可能性が高く、そのようなコードは一般的に行ごとの価値も低くなる傾向があります。
これをモデル化する際は、コードを「リスクの低い部分 (L)」と「リスクの高い部分 (H)」に分割して考えます。このモデルの詳細な定義は付録に記載されていますが、簡潔に言えば、外層ではコードと非コードの間で Cobb-Douglas 関数が用いられ、内層ではリスクの低いコードと高いコードの間で CES(定弾力性の代替)関数が適用されます。ここでαは、AI 導入前の全コードに対するリスクの低いコードの時間シェア割合を表します。
仮に、リスクの高いコードの対数向上率がリスクの低いコードの 1/3 になるとした場合(つまり g_high = g_low^(1/3))、かつ他のパラメータも妥当な値を選べば、以下の表のような結果が得られます。
α (低リスクシェア) | g_low | g_high | U
0.9 | 9.4 | 2.1 | 2.33
0.7 | 13.8 | 2.4 | 2.39
0.5 | 22.6 | 2.8 | 2.49
0.3 | 44.9 | 3.6 | 2.66
0.1 | ~139 | 5.2 | 3.08
α の値を [0.3, 0.9] の範囲外に設定することは、現実的ではないとして却下できます。なぜなら、α が低い値であれば g_low > 44.9 が必要となり、逆に α が高い値であれば AI 導入前の研究者のコーディング時間の 90% 以上が低リスクタスクに費やされていたことになり、これはフロンティア AI 研究におけるコードレビューや大規模実験の頻繁な実施などを考慮すると不自然だからです。
このモデルは σ_LH に対してある程度の頑健性を示しますが、均質な CES モデルが σ に対して持っていたほどではありません。α = 0.9 の場合、σ_LH が [0.5, 2] の範囲で変化する際に U は ±11% 変動します(これは ±3% に比べて大きいです)。なお、α = 0.5 のときは偶然にも±2% の変動しか起こらず、α = 0.3 では±17% となります。
注意点:アンソロピックの向上率が 2 倍未満になる可能性はありますか?
結果に実質的な影響を与えるデータや方法論上の課題として考えられるものは 5 つあり、そのうち 3 つは十分に現実的だと判断できます。
現実的な理由
冗長性
AI の導入により、同じ機能を実現するために開発者がより多くのコード行数を書くようになる可能性はありませんか?METR が 2025 年初めに実施したランダム化比較試験(RCT)では、オープンソース開発者を AI 利用可能グループと不可グループに無作為に割り当てました。その結果、AI を使用できる場合、開発者はより多くのコード行数を記述することが確認されました。具体的には、AI 利用可否の両方の課題を抱えた 10 人の開発者において、完了した PR に追加されたコード行数の幾何平均は、AI 許可時の課題で 1.22 倍から 2.57 倍(95% 信頼区間)に達しました。信頼区間の幅が広く、Claude 3.7 Sonnet や o1 エラ時代の AI と現代の AI の間に差異があるため、これだけでは冗長性に関する確定的な結論は出せません。しかし、アンソロピックにおける冗長性の係数を 1.83 倍(やや不確実な中央推定値)と仮定すれば、コード出力量は依然として約 4.4 倍(8/1.83)増加することになります。また、低リスクシェア α が [0.3, 0.9] の範囲にあるコード異質性モデルを用いると、研究者の向上率は 1.84 から 2.08 の範囲に収まり、これは 2 倍という閾値のすぐ近くとなります。
増加の形状を見ると、冗長性がコード出力の数値に 2 倍を超える影響を与えているようには思えません。2025 年第 4 四半期(Anthropic がおそらく Opus 4.5 と 4.6 にアクセスできていた時期)には、1 人あたりのコード行数(LoC/person)はベースラインの 2.5 倍でしたが、2026 年第 1 四半期(Mythos Preview)には 5.8 倍に跳ね上がりました。2.5 倍の時点で、すでにコードの大部分が AI によって書かれている状態です。また、Fable や Mythos が Opus よりも著しく冗長であるという話は聞きません。したがって、2.5 倍から 5.8 倍、そして 8.0 倍への増加は、主に冗長性によるものではないと考えられます。
Anthropic の 1 人あたりのコード出力が指数関数的な傾向を維持し続けるなら、もはや冗長性を過度に心配する必要はないはずです。なぜなら、すでにほぼすべてのコードが AI によって生成されているからです。
これよりも強い主張には自信がありません。というのも、冗長性はコーディングの生産性向上(uplift)自体の一部であるためです。例えば、コーディングの生産性が非常に高い場合、研究者はコードを一切読む時間がなくなり、それがさらにコードを肥大化させる可能性があります。
ほとんど役に立たないコード
冗長性とは、同じ機能を実現するためにコード行数が増えることを指しますが、もし研究者が、研究の進展にあまり寄与しない新しい機能を備えたコードも書いているとしたらどうでしょうか。Tom Cunningham はこれを「キャデラック・タスク」と呼んでいます。このようなタスクが存在することは、新しいタスクにおける生産性向上(uplift)が、実際の価値向上を過大評価していることを意味します。
METR での経験則として、人々は手書きでは決して書かないような、ほとんど役に立たないコードを大量に生成しています。具体例としては以下のようなものがあります:
- プロジェクト内のタスク間のすべての依存関係を示すプロジェクト DAG の可視化ツール
- リモート開発環境をバックエンドとする、長時間実行されるエージェント・ランのための Web インターフェース
統計手法の根本見直し:1 回の検証のためにプロジェクト全体を再構築する
CES(コブ・ダグラス型生産関数)は、コード量が極端に少ない場合の限界効用逓減が著しいため、ある程度の「ほぼ役に立たないコード」の影響を既に考慮しています。したがって、CES の予測を超えるほど大量の「ほぼ役に立たないコード」が生産されている理由が CES 以外にある場合にのみ、推定されたアップリフト値をさらに引き下げるべきです。
その要因の一つとして、非合理的な時間配分が挙げられます。もう一つは、限界価値が機会費用を下回る前に、CES が予測するよりも多くの「ほぼ役に立たないコード」を書き溜めてしまう可能性です。
「ほぼ役に立たないコード」の有用性には下限があります。なぜなら、私たちがそれらに対して得られる速度向上は 25 倍未満であることが多く、結果として、手作業で書く限界となるコアコードに比べて、少なくとも単位時間あたり 0.04 倍の価値があるからです。今後のモデル化では、この事実を用いてより保守的な下限値を設定することが可能です。
非合理的な時間配分や、AI を使用したいという本能的な欲求など
本記事におけるすべてのモデルは、研究者がコードと非コード業務の間、および低リスク・高リスクのコード作成の間で、研究進展を合理的に最大化するように時間を配分していることを前提としています。しかしこれはやや疑わしい仮定です。なぜなら、時間配分には他にも様々な要因が影響するからです。例えば、利便性や楽しさ、組織の方針、あるいは非合理的な行動などがそれらに含まれます。
昨年の同一研究において、METR はオープンソース開発者が AI によって作業速度が約 20% 向上したと認識していた一方、実際には平均して約 20% 遅延していたことを発見しました。この現象の一部は AI への不慣れさに起因しますが、Anthropic の貢献者には当てはまりません。ただし、コーディングが以前よりも楽しくなったことで、結果として作業に費やす時間が増えているという点は十分にあり得ると考えられます。
非合理的な時間配分がコード出力に与える影響は、冗長性(定数倍の要因に近い)とは異なり、向上率が大きくなるほど悪化する可能性があります。
考えにくい理由
極端な不均一性
AI が、事前 AI 時代のコーディング時間の 25% 未満を占める低リスクのコードでは劇的に速度を向上させる(20 倍以上)一方で、高リスクのコードや非コードタスクではほとんど効果がないとしたらどうでしょうか。
私はこれを不自然だと考えます。その理由は以下の通りです。
- 今日では手書きでコードを書くことは極めて稀であり、AI の利用による速度向上が必ず存在するはずです。
- METR は、多様なエンジニアや研究者を対象とした調査における自己申告の向上率が平均して約 2 倍であったと報告しており、その効果は大きすぎて開発者が AI を使用することを望むあまり、ランダム化比較試験(RCT)を実施することが不可能でした。
Anthropic のグラフにおいて、分母の「アクティブコントリビューター」とは「直近 12 ヶ月間に活動した別個の著者」を指します。つまり、ランダムな営業担当者がコーディングを始めたり、Anthropic の平均的な人材レベルが低下したりした場合、彼らがエンジニアや研究者の平均よりも多くのコードを書かない限り、全体の平均値は下がってしまいます。
分母が時間とともに大幅に縮小した可能性(例えば、頻繁でない貢献者がコーディングを止めた場合など)は考えられますが、それはあまり起こりにくいことのように思われます。
ディスカッション:全体の上昇率を見積もるにはコード出力重視で
CES モデルにおいて、コードの増加量のみに基づく上昇率の見積もりは分散パラメータσに対して頑健ではありません。一方、コード出力に基づく見積もりは頑健です。その主な理由は、コード出力が時間配分の変化を通じてコードの限界価値を考慮している点にあります。
(この頑健性は、出力係数 M がおよそ e² の付近で最も高くなりますが、いずれにせよコード出力の方がコード増加量よりも優れています。)
時間配分が限界価値を正確に反映している限り、コード出力はコードアップリフトよりも優れた指標と言えます。前述のような「不合理」なコーディング時間の偏り(限界価値の変化がないにもかかわらず AI への時間配分が増える現象など)の影響を受けやすい側面はありますが、コード出力の変動要因としては二次的なものだと考えています。したがって、どちらか一方を選ぶなら、コードアップリフトよりもコード出力の方がより適切な指標である可能性が高いです。また、コードアップリフトから研究者全体の生産性向上を推定しようとしても、コード出力から推定する場合と同様のデータ上の課題に直面し、大きな利点はないでしょう。
では、なぜアンソロピック社自身の試算はこれほど低いのでしょうか?
2026 年 4 月に発表された「Mythos Preview」のシステムカードには、「AI の加速は、我々の AI 進歩全体のペースを AI に起因する持続的な倍増にまで高めるにはまだ遠い。この加速は主にエンジニアリングの実行段階に集中しており、研究判断の領域には及んでいない」と記されています。アンソロピック社が「R&D(研究開発)速度が 2 倍になった」と主張した際の算出方法は公開されていないため、その手法を批判することはできませんが、この差はおそらく我々が推定している対象の違いによるものだと考えられます。
具体的には、私が推定しているのはシリアルの研究者一人あたりの生産性向上(serial researcher uplift)であり、アンソロピック社が推定しているのは計算資源の投入量も考慮した研究開発全体の速度向上(overall R&D speedup)です。
原文を表示
Note: the modeling assumptions and conclusion are Thomas Kwa’s opinion, and others at METR disagree.1 Also, the math was checked by Claude but not a second human.
Introduction
Anthropic’s RSI blog post reported that in Q2 2026, Anthropic contributors merged 8× as much code per day as in the 2021-2024 period. What does this imply about the factor by which a researcher’s total effective output increased — the (serial) researcher uplift2?
Of course, 8× more code doesn’t mean 8× more research, as coding is only part of the job. However, if we assume each line of code (LoC) has equivalent quality and verbosity to pre-2025 code and make standard economic modeling assumptions, we can conclude a surprising amount: all models predict researcher uplift at Anthropic from coding agents alone is >2×. (Researcher uplift could be even higher, because these numbers assume no uplift on non-code tasks.)
Cobb-Douglas predicts that if pre-AI time spent coding is \(\beta=50\%\) and code output increases by a factor \(M=8\), then researcher uplift is \(M^\beta = 2.83\).
CES (constant elasticity of substitution) production functions, due to a fun mathematical coincidence, infer a narrow range of about \([2.75, 2.91]\) if code is homogeneous.
If code is CES but non-homogeneous, such that AI speeds up low-stakes code more than high-stakes code, we obtain a range of \([2.33, 2.66]\), lower but still over 2x.
However, there are several reasons new code may not be equivalent to old code, which would be at least partially resolvable with internal Anthropic data.
Verbosity: \(M\) substantially overstates true quality-adjusted code output (e.g. perhaps AI code is >2x more verbose than human code for the same functionality).
Barely-useful code: AIs are speeding up people enormously (>20x) on low-stakes code that would never have been written pre-AI, but is only 5%-20% as valuable as normal code.
Researcher irrationality: Anthropic researchers are producing code that doesn’t contribute to research value (e.g. because vibe coding is fun).
Anthropic’s Mythos system card claims that their overall R&D uplift is “well short of” 2x, which is consistent with >2x researcher uplift because R&D depends on both labor and compute (and compute increases don’t count towards the number). I’d guess that overall R&D uplift will probably hit 2x somewhere around 3.5x researcher uplift, which could happen in the next year or so. Therefore, it’s important that labs obtain sufficiently high-quality data and models to measure code output beyond just LoC and relate it to researcher uplift.
Economic models
The core question is: if a researcher3 produces 8× more code, how much more research value do they create overall? This depends on how important code is relative to everything else a researcher does (thinking, writing, experiments, communication), and on whether those activities are substitutes or complements — i.e., whether doing more of one makes the others more or less valuable.
A production function is the standard economic tool for this. It takes the quantities of each input (here, code output and non-code output) and returns total value. The most common production function is Cobb-Douglas; we consider both Cobb-Douglas and a slight generalization CES (Constant Elasticity of Substitution).
Cobb-Douglas model
The simplest economic model for predicting how much of a good is produced, if it needs more than one input, is called Cobb-Douglas. This is similar to what the AI Futures Project assumes for labor and compute. The equation for Cobb-Douglas gives research output \(Y\) as:
\[Y = q_n^\beta q_o^{1-\beta}\] where:
\(\beta\) is the time share on coding4
\(q_n\) and \(q_o\) are the quantities of code and non-code produced.
We can then define (serial) researcher uplift \(U\) as the ratio of post-AI to pre-AI output, \(U = Y_{\text{post}} / Y_{\text{pre}}\). Assume pessimistically that there is no non-code uplift, which means \(q_o\) is constant. Then it turns out that:
\[U = M^\beta\] If researchers spend roughly half their time coding (\(\beta = 0.5\)), then \(U = \sqrt{8} \approx 2.83\). (\(\beta=0.5\) is roughly the median of what people I ask find reasonable. I have substantial uncertainty about \(\beta\), but fix it at 0.5 for simplicity throughout. A more thorough analysis should certainly vary it.)
CES model
When we move from Cobb-Douglas to the general CES case, we gain the ability to model substitutability: if code and non-code are strong complements (like left and right shoes), doubling code output without more non-code output barely helps. If they’re substitutes (like butter and margarine), doubling one input nearly doubles total value. We don’t know which is true for AI research, so we test across a range.
Setup
A researcher splits fixed time \(T\) across two task types: coding and non-coding. We assume the output \(Y\) follows the CES function:
\[Y = [\beta \, q_n^{\,\rho} + (1-\beta) \, q_o^{\,\rho}]^{1/\rho}, \quad \rho = \frac{\sigma - 1}{\sigma}\] where:
\(\beta\) is the pre-AI time share on coding4 (in CES, it is possible for time shares to change between the pre-AI and post-AI periods);
\(q_n\) and \(q_o\) are the quantities of code and non-code produced.
\(\sigma\) is the elasticity of substitution between coding and non-coding output.
Now, it can be calculated that the researcher uplift \(U\) (assuming researchers allocate time optimally in both periods) is:
\[U = [\beta \, g_n^{\,\sigma-1} + (1-\beta)]^{1/(\sigma-1)}\] where \(g_n\) is the per-hour speedup AI provides on coding, which we will refer to as the coding uplift.5
We don’t directly observe \(g_n\). We observe \(M = 8\), the mean code output increase of Anthropic contributors.6 (These differ because researchers reallocate time when AI makes coding cheaper.) It turns out that \(M\), \(U\), \(\sigma\), and \(g_n\) are related by a fairly simple equation \(M = \frac{g_n^\sigma}{U^{\sigma-1}}\),5 and observing \(M\) gives a fairly robust estimate for \(U\).
Because 8 ≈ e², the estimate is robust to σ!
In Cobb-Douglas, we assumed that code and non-code have unit substitutability (\(\sigma=1\)). What if they’re complements (\(\sigma < 1\)) or substitutes (\(\sigma > 1\))? The conclusion actually changes very little.
The key reason is that 8x code output substantially constrains the possible values of our two free parameters: coding uplift \(g_n\) and marginal value of code (which is increasing in \(\sigma\)). If coding uplift and marginal value of code were both large, then the amount of code produced would be larger than 8x, as code would be both cheap to produce and super valuable. Alternatively, if \(g_n\) and \(\sigma\) were both small, then output \(M\) would be less than 8x, as code would be expensive to produce and not valuable. Only intermediate choices of the two parameters — (a) high coding uplift and low marginal value of code, or (b) low coding uplift and high marginal value of code — are possible given that code output has increased by 8x.
In world (a), code and non-code are strong complements (roughly \(\sigma < 0.5\)). This means the researchers shift their time away from coding to other bottlenecks, and we only see code output 8× because code speedup is extremely high, at least 23x. This frees up so much time for non-code that researchers still achieve uplift of \(U=2.75\).
In world (b), code and non-code are near-perfect substitutes (\(\sigma = 3.0\)); researchers spend almost all their time coding, because it’s more productive and substitutes for non-code tasks. But then an \(M=8\)x code output increase means that researchers must have produced only \(g_n=4.2\)x code per hour, because they have almost doubled their time spent coding. So uplift is still only \(U=3.07\).
In both cases, \(U\) remains fairly close to \(M^\beta = 2.83\).
window.addEventListener('message', function (e) { var iframe = document.getElementById('uplift-explorer'); // Only accept the height message from our own iframe, same-origin, with a // finite numeric payload, and clamp it so a bad value can't blow up layout. if (!iframe || e.source !== iframe.contentWindow) return; if (e.origin !== window.location.origin) return; var h = e.data && e.data.upliftExplorerHeight; if (typeof h !== 'number' || !isFinite(h)) return; iframe.style.height = Math.max(200, Math.min(h, 8000)) + 'px'; }); Why is \(U\) relatively unaffected by \(\sigma\) after observing \(M\)? If we obtain an expression for \(\ln U\) and perturb it around \(\sigma = 1\):7
\\Delta \ln U \approx \beta(1-\beta)\left[-\ln M + \frac{(\ln M)^2}{2}\right\] The bracketed term — which drives the percent difference between \(U\) and \(M^\beta\) — equals zero at \(M = e^2 \approx 7.39\). At \(M = 8\), it equals \(-2.08 + 2.16 = 0.08\) — almost zero. So the first-order sensitivity of \(\ln U\) to \(\sigma\) is only \(0.25 \times 0.08 = 0.02\) per unit \(\sigma\). The upshot: \(U\) stays within ±3% of 2.83 for \(\sigma \in [0.5, 2]\).
\(\sigma\) implied \(g_n\) post-AI coding share \(U\)
0.0 — — impossible (\(M \leq 2\))
0.3 ~114 3.5% 2.56
0.5 23.3 17% 2.75
0.7 12.5 32% 2.80
1.0 8.0 50% 2.83
1.5 5.7 70% 2.86
2.0 4.8 83% 2.91
3.0 4.2 95% 3.07
M = 8 also implies a lower bound on σ
Observing \(M = 8\) gives a \(\sigma\)-robust estimate of uplift, but it also constrains \(\sigma\).
Recall that \(M = g_n \cdot \frac{t_n^{\text{postAI}}}{t_n^{\text{preAI}}}\) where
\(g_n\) is the coding uplift/speedup
\(t_n\) is the quantity of time spent on coding
When \(\sigma < 1\), researchers shift time away from coding (\(\frac{t_n^{\text{postAI}}}{t_n^{\text{preAI}}} < 1\)), so the per-hour speedup \(g_n\) must exceed \(M\) to produce the observed 8× output. The smaller \(\sigma\) is, the larger \(g_n\) must be — at \(\sigma = 0.3\), you need \(g_n \approx 114\). A value of \(\sigma = 0.3\) also implies that Anthropic researchers only spend 3.5% of their time coding in Q2 2026, which is implausibly low.
At \(\sigma = 0\) (Leontief / perfect complements), the production function becomes \(Y = \min(\beta \, q_n, \, (1-\beta) \, q_o)\), so output is bottlenecked by whichever input is scarcer. Even \(g_n \to \infty\) can only produce \(M = 1/(1-\beta) = 2\), so \(M = 8\) is flatly impossible. Intuitively: if non-code is a strict bottleneck, no amount of coding speedup can raise total code output by more than \(1/(1-\beta)\), the maximum increase in non-code output.
It is reasonable to believe that the per-hour coding speedup is <23× and Anthropic contributors now spend >17% of their time coding, which implies \(\sigma \geq 0.5\). This rules out strong complementarity between coding and non-coding research. We could refine this estimate further by measuring Anthropic’s time shares; if they are similar to pre-AI time shares, \(\sigma \approx 1\), whereas if they are higher or lower, the shift away from or towards code would mean lower or higher \(\sigma\) for the reasons in the previous section.
Code heterogeneity model: What if AI speeds up high-stakes code less?
The above model assumes all code is uniformly sped up. But it is probably true that AI speedup is higher on low-stakes code, which also tends to have lower value per LoC. We can model this by splitting code into low-stakes (\(L\)) and high-stakes (\(H\)) components.8 The full definition of this model is in the appendix, but briefly, the outer layer is Cobb-Douglas between code and non-code, while the inner layer is CES between low- and high-stakes code, with \(\alpha\) being the pre-AI time share of low-stakes code as a fraction of all code.
If high-stakes code gets 1/3 the log-uplift: that is, \(g_{\text{high}} = g_{\text{low}}^{1/3}\), and we make other reasonable parameter choices, we get the following table:
\(\alpha\) (low-stakes share) \(g_{\text{low}}\) \(g_{\text{high}}\) \(U\)
0.9 9.4 2.1 2.33
0.7 13.8 2.4 2.39
0.5 22.6 2.8 2.49
0.3 44.9 3.6 2.66
0.1 ~139 5.2 3.08
We can reject any value of \(\alpha\) outside \([0.3, 0.9]\) as implausible — low values would require \(g_{\text{low}} > 44.9\), while high values would mean that >90% of researcher coding time pre-AI was spent on low-stakes tasks, which I find implausible given the prevalence of code review, large experiments, etc. in frontier AI research.
This model is still somewhat robust to \(\sigma_{LH}\), but not as much as the homogeneous CES model was to \(\sigma\): at \(\alpha = 0.9\), \(U\) varies by ±11% across \(\sigma_{LH} \in [0.5, 2]\) instead of ±3%. (At \(\alpha = 0.5\) it happens to vary by only ±2%, but at \(\alpha = 0.3\) by ±17%.)
Caveats: How could Anthropic’s uplift be less than 2x?
I can think of five data and methodological issues that could meaningfully affect the results, of which three are plausible.
Plausible reasons
Verbosity
What if AI causes contributors to write more lines of code for the same functionality? In METR’s early-2025 uplift RCT, where open-source developers were randomly assigned AI and non-AI issues, developers wrote more lines of code when they were allowed to use AI. Specifically, for the 10 developers with both AI and non-AI issues, the geomean LoC added in completed PRs was somewhere between 1.22x and 2.57x as large (95% CI) for AI-allowed issues. Due to the wide CI and the differences between Claude 3.7 Sonnet/o1 era AI and modern AI, we can’t prove anything about verbosity without further investigation. But if we assume that Anthropic’s verbosity factor is 1.83x (the sketchy central estimate), code output will still have increased \(8/1.83 \approx 4.4\)x, and the code heterogeneity model with low-stakes share \(\alpha \in [0.3, 0.9]\) gives researcher uplift in the range \([1.84, 2.08]\) — right around the 2x threshold.
The shape of the increase makes me think verbosity doesn’t affect the code output number by more than ~2x. In Q4 2025 (when Anthropic probably had access to Opus 4.5 and 4.6), LoC/person was 2.5x baseline, but in Q1 2026 (Mythos Preview), it jumped to 5.8x. At 2.5x, the majority of code is already AI-written, and anecdotally Fable/Mythos is not much more verbose than Opus, so the increase from 2.5x to 5.8x to 8.0x is mostly not verbosity. If the exponential trend in Anthropic’s per-capita code output continues, we should become less worried about verbosity because almost all code is already AI-written.
I’m not confident in any statement stronger than this, because verbosity is partly a function of coding uplift itself. E.g. if coding uplift is very high, researchers no longer have time to read any code, which could bloat it further.
Barely useful code
Verbosity means more LoC for the same functionality, but what if researchers are also writing more code with new functionality that doesn’t create much research progress? Tom Cunningham calls these Cadillac tasks, and their existence means that uplift on new tasks is always an overestimate of value uplift.
Anecdotally at METR, people generate lots of barely useful code they wouldn’t write by hand. Some examples:
A project DAG visualizer for all the dependencies between tasks in the project
A web interface to long-running agent runs backed by a remote dev instance
Redoing a project’s entire stats methodology for one sanity check
Barely useful code is somewhat accounted for by CES (which has strongly diminishing marginal returns to code when \(\sigma \ll 1\)) so one should only discount uplift estimates further if there is some reason beyond CES that lots of barely-useful code is being produced. One factor could be irrational time allocation; another could be that it’s simply possible to write a larger volume of barely-useful code than CES predicts before the marginal value drops below one’s opportunity cost.
The utility of barely useful code can be bounded below, because most of us get less than 25x speedup on them, and therefore they’re at least 0.04x as valuable per unit time as the marginal core code we’d write by hand. Future modeling efforts could use this to get a more conservative lower bound for uplift.
Irrational time allocation, intrinsic desire to use AI, etc.
All modeling in this post assumed that researchers allocate their time between code and non-code, and between low- and high-stakes code, in a way that rationally maximizes research progress. This is a somewhat dubious assumption, because there are various other factors that determine time allocation: convenience, fun, organizational policy, irrational behavior, etc.
In the same study last year, METR found that open-source developers thought that AI had sped them up ~20%, even when AI had actually slowed them down by an average of ~20%. This particular effect is partially due to inexperience with AI, which does not apply to Anthropic contributors, but I can certainly believe that writing code is often more fun than it used to be, leading people to spend more time on it.
The impact of irrational time allocation on code output probably gets worse the larger uplift is, unlike verbosity, which is closer to a constant factor.
Unlikely reasons
Extreme heterogeneity
What if AIs are speeding up people enormously (>20x) on low-stakes code that makes up less than 25% of pre-AI coding time, and basically not at all on high-stakes code or non-code tasks?
I find this implausible because:
Very little code is written by hand these days, so there must be some speedup from AI use.
METR found that self-reported uplift among survey participants, which averaged over a wide range of engineers and researchers, was around 2x, and uplift was large enough that developer preferences to use AI precluded conducting an RCT to measure it.
Changing denominator
Anthropic’s graph says “active contributor” in the denominator means “a distinct author in the trailing twelve months”. This means that if random salespeople started coding, or Anthropic’s average talent level went down, they would bring the average down unless they wrote more code than the average engineer/researcher. It is possible that the denominator has greatly shrunk over time e.g. if infrequent contributors have stopped coding, but this doesn’t seem likely.
Discussion
Prefer code output over code uplift, for estimating overall uplift
In the CES model, uplift estimates based on code uplift alone are not robust to \(\sigma\), but estimates using code output are. The key reason is that code output takes into account marginal value of code, through time allocation changes. (The robustness is highest around output factor \(M \approx e^2\), but is always better for code output than code uplift.)
To the extent time allocation accurately reflects marginal value, code output is better than code uplift. It is susceptible to “irrational” changes in coding time allocation such as discussed above (which might make time allocation shift towards AI in the absence of marginal value changes); however, I expect these to be secondary drivers of code output. So overall, code output is probably a better metric than code uplift if we had to pick one. I expect that modeling overall researcher uplift from code uplift would have most of the same data issues as from code output, and not provide much advantage.
Why is Anthropic’s own estimate much lower?
The Mythos Preview system card (April 2026) stated that AI acceleration is “well short of a sustained, AI-attributable doubling of the overall pace of our AI progress. The acceleration is concentrated in engineering execution rather than research judgment.” Anthropic’s methodology for their claim of «2x R&D speedup is not public, so I am not able to critique it, but the difference is probably just that we’re estimating different quantities.
Specifically, I estimate serial researcher uplift and they estimate overall R&D speedup, which also depends on compute a
AI算出
市場分析ainew評価標準
AI コーディングエージェントによる生産性向上率を推定する独自の経済モデル分析が含まれており、単なる事実報告ではなく深い洞察を提供しているため market_analysis に分類されます。また、METR の分析という一次情報に基づく新規性が認められますが、日本企業への直接的な影響や特定バージョンの発表はないため、関連性は低めです。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み