Anthropic、新生産コードの80%がClaudeによって作成されたと発表—企業も追いつく方法とは(7分読了)
Anthropic は自社の生産コードの 80% が AI モデル「Claude」によって作成されたことを発表し、これは業界全体が直面する新しい競争基準を示す画期的な事例である。
キーポイント
AI によるコード生成の圧倒的シェア
Anthropic の 5 月の生産環境へのマージコードのうち、80% 以上が人間ではなく AI モデル「Claude」によって作成されたことが報告された。
エンジニアリング生産性の劇的向上
AI オートメーションの導入により、エンジニアあたりの四半期当たり出荷コード量が 2021-2025 年のベースラインと比較して 8 倍に増加した。
進化のロードマップと段階的移行
Anthropic は、手書きからチャットボット支援、コーディングエージェントを経て、現在は自律的なデバッグやタスク委任を行う「自律型エージェント」へと至る明確な進化フェーズを提示している。
再帰的自己改善への道
この事例は、AI が自らを研究・アップグレードする「再帰的自己改善(recursive self-improvement)」の兆候を示しており、他企業の自動化戦略のロールモデルとなっている。
重要な引用
More than 80% of the code merged into Anthropic's production codebase in May wasn't authored by humans, but by its own AI model, Claude
This transformation has triggered an 8x increase in the volume of code shipped per engineer per quarter compared to the company's 2021–2025 baseline
For enterprise technical leaders, this is no longer a localized research curiosity; it's a new, aggressive competitive baseline
影響分析・編集コメントを表示
影響分析
この記事は、AI が単なる補助ツールから自律的な開発者へと役割を変化させた決定的な転換点を示しており、業界全体が「AI 主導の開発」への移行を迫られていることを意味します。今後、他社も同様の生産性向上を実現するために、エージェントベースのワークフロー再設計を急務とするでしょう。
編集コメント
80% のコードが AI 生成という数字は、単なる効率化の域を超え、開発パラダイムそのものの転換を示唆しています。他社もこの「新基準」にどう対応するかが今後の競争力を分ける鍵となるでしょう。
Anthropic の共同創設者兼 CEO、ダリオ・アモダイは その到来を予告していました、しかしそれでもなお、これは画期的なマイルストーンのように感じられます。今日、記録的な成長を遂げた AI スタートアップが共有した 新しいレポート によると、5 月に Anthropic の本番コードベースにマージされたコードの 80% 以上は人間によって作成されたものではなく、同社の独自 AI モデルである Claude によって作成されたものです。
この変革により、エンジニア一人あたりの四半期あたりに出荷されるコード量が、2021〜2025 年の基準と比較して 8 倍に増加しました。同社はこれにより、誰か、あるいは何かがレビューしなければならないコード量がさらに増えることを指摘しています。
企業における技術リーダーにとって、これはもはや局所的な研究の好奇心の対象ではなく、新たな、攻撃的な競争の基準線となっています。
最前線の AI 研究所が、そのエンジニアリング成果の大部分を自律型エージェントに委譲することに成功し、長年求められてきた AI の聖杯である「再帰的自己改善」(モデルが独自に研究し、自身をアップグレードできる能力) の兆候を示している場合、他のセクターの企業がなぜ AI エージェントを用いて自社の内部ソフトウェア開発のより多くの部分を自動化できないのでしょうか?
もちろん、言うのは簡単ですが実行は容易ではありません。Anthropic は現在の第 1 世代 AI ブームの主要な創出者の一人であるため、彼らがこの技術を効果的に導入する方法を知っていることは予想通りです。
しかし、エージェントが処理するコードやワークフローの量を増やそうとする他の企業にとって、Anthropic の新しいブログ記事は、最新の AI の進歩を活用するために自社の運用とワークフローを再設計するための一般的な計画の概要を詳細に記述しており、他社も同様にこの計画を採用することができます。
他の企業が従うことができる Anthropic のロードマップ
人間中心のコーディングから自律的なオーケストレーションへの移行には、AI 機能の進化を理解することが必要です。Anthropic は、企業が自社のデジタルトランスフォーメーションのロードマップにマッピングできる明確な歴史的連続性を提示しています:
- 2021–2023(手動記述):エンジニアはローカルのテキストエディタ内でコードとドキュメントをネイティブに記述します。
- 2023–2025(チャットボット支援):開発者は初期モデルを使用して短いコードスニペットを生成し、その出力を手動でコピー&ペーストして環境に組み込みます。
- 2025–2026(コーディングエージェント):能力の高いエージェントが自律的にファイル全体を書き換え・編集します。
- 今日現在(自律型エージェント):エージェントはコードを独立して実行し、ライブ環境のデバッグを行い、数時間にわたる作業ストリームを専門的なサブエージェントに委任します。
この急速な進化は外部ベンチマークによって裏付けられています。SWE-bench といったソフトウェアエンジニアリング評価フレームワーク(複雑なオープンソースコードベース内の実際のバグレポート解決をモデルに課すもの)は、2 年間の期間で飽和状態に達しました。
さらに、長時間実行能力の評価では、Claude Opus 4.6 が 12 時間にわたるタスクを確実に持続できることが示され、Claude Mythos Preview は 16 時間を超える継続的な問題解決の領域へと押し広げています。
内部技術的には、その飛躍はさらに顕著です。明確な仕様が一見存在しない非常に複雑で開放的なエンジニアリング課題において、Claude の成功率は 2026 年 5 月に 76% に達し、これは 6 ヶ月間で 50 ポイントの増加となります。
モデルに AI モデルのトレーニングコードを高速化させることを課す孤立した最適化ベンチマークでは、Anthropic の内部モデルである Mythos Preview が 52 倍の速度向上を達成しました。
比較のために、熟練した人間の開発者が、全く同じコードベースにおいて単に 4 倍の速度向上を達成するために通常必要とするのは、手動のリファクタリングに 4 から 8 時間です。
より完全なプロダクションコード自動化のための 3 ステッププラン
企業が Anthropic の 80 パーセントというマイルストーンを再現するためには、技術的な意思決定者は「開発者アシスタント」という思考モデルを捨て、「自動工場」アーキテクチャへと移行する必要があります。この転換は、製品管理、運用、そして開発ワークフローの 3 つの側面に明確な影響を与えます:
1. コード実行からアーキテクチャ監督へシフトする
人間の時間におけるコード生成コストがほぼゼロに近づくと、主要なエンジニアリングの役割はソフトウェアを書くことから、目標を指定し出力を検証することに移行します。企業のリーダーは、開発者がシステムアーキテクトおよび審査員として行動するように再教育する必要があります。この転換の実務的な側面について、ある Anthropic の従業員は次のように述べています:
「今日の状況の形はおおよそ『人間がアイデアを持ち、モデルがそれを実装し、テストし、評価する速度が以前よりも 1 つオーダー(桁)速い』というものです。」
2. コードレビューのボトルネックを克服する
組織に膨大な量の AI 生成コードを注入することは、必然的に運用上の摩擦を生み出します。
Amdahl の法則によれば、あらゆるプロセスの速度向上は、その直列化された非自動化のボトルネックによって厳密に制限されます。
Anthropic では、システムに合成コードを洪水のように流し込むことで、即座に人間のコードレビューが重要なボトルネックとなりました。
これに対抗するため、企業チームは、自動化された AI コードレビューアーを継続的インテグレーション/継続的デリバリー(CI/CD)パイプラインに直接展開する必要があります。
Anthropic は、プルリクエストのマージ前にアーキテクチャ上の欠陥、セキュリティ上の欠陥、回帰バグを分析するよう任務を与えられた自動化された Claude レビューアーを実装しました(商用利用向けに 3 月に展開された公開アクセス可能なバージョン Claude Code Review)。これと同様に、Qodo のような専門企業も、この目的のために特別に設計されたツールを提供しています。
Anthropic の事例では、事後分析により、自動化された層が旗艦サイト claude.ai で過去に発生した停止の原因となった生産環境バグの約 3 分の 1 を検出していたことが示されました。
3. 高ボリュームの運用負債への対応
企業は頻繁に、レガシーコードの保守と長期間先送りされた技術的負債によって足止めを食らっています。エージェントに推測的な新機能の作成を任せるのではなく、技術リーダーは自律型エージェントを、閉ループかつ手間のかかるクリーンアップ作業に向けさせるべきです。
2026 年 4 月、Anthropic のエンジニアは Claude を活用して、永続的な API エラーのクラスを解決しました。自律的に動作したこのモデルは、800 件以上の個別修正を実行し、エラー率を 1,000 分の 1 にまで劇的に低下させることに成功しました。
監督していたエンジニアによれば、膨大で見知らぬコードの文脈を同時に頭の中で保持するという認知的負荷(cognitive load)のため、人間が開発者であれば同じ作業を実行するには 4 年もの完全な期間が必要だったと推定されています。
主に AI 生成コードによって構成される時代における企業への考慮事項
AI が主に作成したコードベースを運用することは、企業の法務チームやセキュリティチームが対処しなければならない独自のガバナンス課題をもたらします。
寛容な MIT ライセンスやコピレフト GPL フレームワークなどのオープンソースライセンスモデルとは異なり、独自 LLM インフラストラクチャを利用する企業向けコードベースは、それぞれの AI ベンダーの商用利用規約(terms of service)の対象となります。
自律型エージェントの導入には、コンプライアンス、セキュリティ、知的財産保護を確保するための厳格な検証プロトコルが必要です。
- コードの品質と保守:Anthropic の内部データによると、2025 年後半には AI が作成したコードは客観的に人間による出力より質が低かったものの、2026 年半ばにはおおよそ同等レベルに達し、その年内には人間の基準を上回る見込みです。企業ガバナンスは、自動化された出力のベースライン品質が平均的な手動コーディングよりも構造的に優れているという現実に対応する必要があります。
- スケールしたセキュリティ監査:大量の自動コード生成には、自動化された脆弱性発見が不可欠です。Anthropic の「Project Glasswing」プロジェクトはこの課題の規模を示す事例であり、「Mythos Preview」を活用することで、最初の数週間でグローバルなデジタルインフラストラクチャ全体で 10,000 件以上の高・重大度ソフトウェア脆弱性を特定しました。これにより、企業のサイバーセキュリティ上の課題は、脆弱性の発見からパッチ展開の速度へと完全にシフトしました。
- アライメントカスケードのリスク:技術リーダーは厳格な検証ゲートを維持しなければなりません。企業が AI システムを使用して自社固有のソフトウェアインフラストラクチャを継続的に修正・保守・拡張する場合、検出されないエラーや微妙なアライメントのズレが、 successive エージェントセッション(連続的なエージェント実行)を通じて蓄積し、システム整合性を徐々に損なったり、人間の注意を逃れるセキュリティエクスプロイトを導入したりする可能性があります。
Brace for internal enterprise culture disruption
The transition to an AI-dominated codebase is altering the cultural dynamics of engineering teams, introducing both unprecedented efficiency and deep psychological friction.
Publicly, Anthropic framed these metrics as a harbinger of a broader transformation. In an official statement on X, the company observed:
"Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It's happening faster than we thought, and the implications deserve greater attention."
They expanded on the immediate productivity implications shortly thereafter:
"Today, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025... Many engineers also say Claude's code quality is now on par with human code; we expect it to be better within the year."
Behind these corporate metrics lies a complex human reality. Internal employee communications reveal a distinct erosion of traditional workplace collaboration, as peer-to-peer developer interaction is systematically replaced by asynchronous agent calls:
「仕事(そして人生)は、人間同士の間で行われる小さな親切の贈与経済によって成り立っていた。『このスクリプトを動かせる?』……それぞれがわずかな負債と、相互の意識を生み出していた。しかし Claude はその親切を奪い取った。それは速く、負債をゼロにするが、それぞれの行為は人間同士の協働への機会を失うことを意味する。」
個人貢献者にとって、主要なスキルセットの完全自動化は、関連性とシステムに対するコントロールに関する深刻な職業的不安をもたらす:
「私は約 1 年前から本格的に『Claudifying』に取り組み始めた。それは驚くべき冒険であり、私が自分でコードを書いたのは約 5 ヶ月前が最後だ。」
「すべてがうまくいく日には、自分が何をしているのか意味がないと感じず、すべてが自動化され、私が決して及ばないほど速く、良くなっていると考えるしかない。しかし一方で、すべてが壊れてなぜそうなるのかわからない日もあり、自分が何をしていたのかもう全くわからないと気づかされる。」
Anthropic の技術的スピードに追いつこうとする企業リーダーは、これらの心理的なダイナミクスを無視する余裕はない。
コードベースの 80 パーセントを自動化するには、API トークンの購入やエージェントループの設定だけでなく、文化の全面的な見直し、開発者の陳腐化不安を緩和するための戦略、そしてソフトウェアスタックに対する最終的な人間のコントロールを維持するための厳格で自動的な検証ガードレールの導入が求められる。
原文を表示
Anthropic co-founder and CEO Dario Amodei said it was coming, but it still feels like a milestone: More than 80% of the code merged into Anthropic’s production codebase in May wasn't authored by humans, but by its own AI model, Claude, according to a new report shared by the record-breaking AI startup today.
This transformation has triggered an8x increase in the volume of code shipped per engineer per quarter compared to the company’s 2021–2025 baseline, which the company notes means even more code someone or something must review.
For enterprise technical leaders, this is no longer a localized research curiosity; it's a new, aggressive competitive baseline.
If a frontier AI laboratory can successfully offload the vast majority of its engineering output to autonomous agents — showing signs of the long-sought AI Holy Grail of "recursive self-improvement," models that can independently research and upgrade themselves — what's preventing enterprises across other sectors from automating more of their internal software development with AI agents, too?
Obviously, it's easier said than done. Anthropic is one of the principle creators of the current gen AI boom, so you'd expect them to know how to deploy the technology effectively.
But for other enterprises looking to bump up the amount of code and workflows handled by agents, Anthropic's new blog post details the outlines of a general plan they too can adopt to re-engineer their operations and workflows to take advantage of the latest AI advances.
Anthropic's roadmap that other enterprises can follow
The transition from human-centric coding to autonomous orchestration requires understanding the evolution of AI capabilities. Anthropic outlines a clear historical continuum that enterprises can map onto their own digital transformation roadmaps:
- 2021–2023 (Manual Writing): Engineers write code and documentation natively within local text editors.
- 2023–2025 (Chatbot Assistance): Developers use early models to generate brief code snippets, copying and pasting outputs manually into their environments.
- 2025–2026 (Coding Agents): Capable agents actively write and edit entire files autonomously.
- Present Day (Autonomous Agents): Agents execute code independently, debug live environments, and delegate multi-hour work streams to specialized sub-agents.
This rapid evolution is validated by external benchmarks. Software engineering evaluation frameworks like SWE-bench—which tasks models with resolving real bug reports in complex, open-source codebases—have saturated over a two-year window.
Furthermore, long-duration capability evaluations demonstrate that models like Claude Opus 4.6 can reliably sustain operations on 12-hour tasks, while Claude Mythos Preview pushes past 16 hours of continuous problem-solving.
Internally, the technological leap is even more stark. On highly complex, open-ended engineering problems where clear specifications are initially absent, Claude’s success rate climbed to 76% in May 2026 — a 50-point increase in a six-month window.
In isolated optimization benchmarks, where models are tasked with accelerating AI model training code, Anthropic’s internal Mythos Preview model achieved a 52x speedup.
For comparison, a skilled human developer typically requires four to eight hours of manual refactoring to achieve a mere 4x speedup on the exact same codebase.
3-step plan to more complete production code automation
For an enterprise to replicate Anthropic's 80 percent milestone, technical decision-makers must abandon the "developer assistant" mental model and transition to an "automated factory" architecture. This shift impacts product management, operations, and developer workflows in three distinct ways:
1. Shift from Code Execution to Architectural Oversight
When code generation costs near zero in human time, the primary engineering role shifts from writing software to specifying goals and reviewing outputs. Enterprise leaders must retrain developers to act as systems architects and judges. As one Anthropic employee noted regarding the operational reality of this shift:
"The shape of stuff today is roughly ‘humans have ideas, and the models are able to implement, test and evaluate them an [order of magnitude] faster than before.’"
2. Overcome The Code Review Bottleneck
Injecting vast quantities of AI-generated code into an organization inevitably creates operational friction.
According to Amdahl’s law, the speedup of any process is strictly limited by its serial, non-automated bottlenecks.
At Anthropic, flooding the system with synthetic code instantly turned human code review into a critical bottleneck.
To counter this, enterprise teams must deploy automated AI code reviewers directly into their Continuous Integration/Continuous Deployment (CI/CD) pipelines.
Anthropic implemented an automated Claude reviewer (a publicly accessible version, Claude Code Review rolled out for commercial usage in March) tasked with analyzing every pull request for architectural defects, security flaws, and regression bugs before merging. Other dedicated firms like Qodo offer tools tailor-made for this purpose, as well.
In Anthropic's case, retrospective analyses indicated that the automated layer caught approximately one-third of the production bugs responsible for historical outages on the flagship claude.ai website.
3. Target High-Volume Operational Debt
Enterprises are frequently paralyzed by legacy code maintenance and long-deferred technical debt. Rather than deploying agents to write speculative new features, technical leaders should direct autonomous agents toward closed-loop, painstaking cleanup operations.
In April 2026, an Anthropic engineer deployed Claude to resolve a persistent class of API errors. Operating autonomously, the model shipped more than 800 individual fixes, successfully reducing the error rate by a factor of 1,000.
The supervising engineer estimated that a human developer would have spent four full years executing the same work, due to the cognitive load of holding massive, unfamiliar code context in their head simultaneously.
Considerations for enterprises moving forward in an age of primarily AI-generated code
Operating a codebase predominantly authored by AI introduces unique governance challenges that enterprise legal and security teams must navigate.
Unlike open-source licensing models (such as the permissive MIT license or copyleft GPL frameworks), enterprise codebases utilizing proprietary LLM infrastructure remain subject to the commercial terms of service of the respective AI vendor.
The deployment of autonomous agents requires rigorous verification protocols to ensure compliance, security, and intellectual property protection:
- Code Quality and Maintenance: Anthropic’s internal data indicates that while AI-authored code was objectively lower in quality than human output in late 2025, it reached rough parity by mid-2026, with expectations to surpass human standards within the year. Enterprise governance must adapt to a reality where the baseline quality of automated output is structurally superior to average manual coding.
- Security Auditing at Scale: The sheer volume of automated code creation demands automated vulnerability discovery. Anthropic’s Project Glasswing illustrates the scale of this issue: utilizing Mythos Preview, the project identified more than 10,000 high- and critical-severity software vulnerabilities across global digital infrastructure within its first few weeks. This shifted the enterprise cybersecurity challenge entirely from vulnerability discovery to patch deployment velocity.
- The Risk of Alignment Cascades: Technical leaders must maintain strict verification gates. If an enterprise uses an AI system to continuously modify, maintain, and expand its proprietary software infrastructure, undetected errors or subtle misalignments can compound over successive agent sessions, gradually corrupting system integrity or introducing security exploits that escape human notice.
Brace for internal enterprise culture disruption
The transition to an AI-dominated codebase is altering the cultural dynamics of engineering teams, introducing both unprecedented efficiency and deep psychological friction.
Publicly, Anthropic framed these metrics as a harbinger of a broader transformation. In an official statement on X, the company observed:
"Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention."
They expanded on the immediate productivity implications shortly thereafter:
"Today, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025... Many engineers also say Claude’s code quality is now on par with human code; we expect it to be better within the year."
Behind these corporate metrics lies a complex human reality. Internal employee communications reveal a distinct erosion of traditional workplace collaboration, as peer-to-peer developer interaction is systematically replaced by asynchronous agent calls:
"Work (and life) ran on a gift economy of small favors between humans. ‘Can you help me get this script running?’ [...] each one created a little debt, a little mutual awareness. Claude has eaten the favors. It’s faster, it creates zero debt, but each of these is a lost bid for human collaboration."
For individual contributors, the total automation of their primary skill set introduces acute professional anxiety regarding relevance and systemic control:
"I started leaning hard into Claudifying about a year ago. That’s been a crazy adventure and it’s now been ~5 months since I last wrote any code myself."
"On days where everything works well, I can’t help but think nothing I do matters, everything is automated and better and faster than I ever will be. But then there are days where everything breaks and I don't understand why and I realize I have no idea what I’ve been up to anymore."
Enterprise leaders aiming to match Anthropic’s technical velocity cannot afford to ignore these psychological dynamics.
Achieving an 80 percent automated codebase requires more than purchasing API tokens or configuring agent loops; it demands a total cultural overhaul, a strategy for mitigating developer obsolescence anxiety, and the implementation of rigorous, automated verification guardrails to maintain ultimate human control over the software stack.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み