自律型企業の未来を語る
TLDR AI は、自律型企業(Self-Driving Company)の概念を解説し、AI エージェントが人間の監督なしに複雑な業務を自律的に実行する未来像とその実現に向けた技術的・組織的な課題について分析している。
キーポイント
自律型企業の定義と進化
従来の自動化ツールを超え、目標を設定し、計画を立て、実行し、結果を検証して自己修正する完全な AI エージェントによる企業運営の概念を提示している。
技術的実現可能性の現状
LLM の推論能力やマルチモーダル処理の進歩が、自律的な意思決定を支える基盤となっている一方、完全な自律化にはまだ信頼性と安全性の課題が残っている。
組織変革と人間の役割
企業が自律型になる過程で、人間は「実行者」から「監督者」や「戦略家」へと役割を転換する必要があり、組織文化の変容が不可欠であると指摘している。
リスクと規制の課題
自律的な意思決定に伴う責任所在の曖昧さやセキュリティリスクへの対応策として、ガバナンスフレームワークの整備が急務であるとしている。
重要な引用
The Self-Driving Company is not just about automation; it's about autonomy.
Agents must be able to plan, execute, and verify their own work without constant human intervention.
Human oversight shifts from doing the work to defining the goals and auditing the outcomes.
影響分析・編集コメントを表示
影響分析
この記事は、AI の役割をツールからパートナーへと昇華させるパラダイムシフトを示唆しており、企業経営や組織論における重要な転換点を指摘しています。技術的な実現可能性だけでなく、人間と AI の協働関係の再定義を迫る内容であり、今後数年以内に多くの企業が自律化への移行を検討する際の指針となるでしょう。
編集コメント
「自律型企業」という概念は、単なる効率化の枠を超え、企業のあり方そのものを変える可能性を秘めています。技術的な進歩と並行して、人間がどう関わるかという経営視点の重要性が浮き彫りになった記事です。
過去6ヶ月間、Replitのエンジニアたちはコード出力を約3倍に増やしました。レビュー所要時間は横ばいを維持し、リバージョンや製品インシデントも発生件数は変わりません。品質指標は向上し、リリース速度も加速しています。通常予想されるようなトレードオフ(何かを犠牲にして何かを得る関係)は一切起きていません。
コードそのものが目に見える部分ですが、表面の下で何が起こっているのかの方が、はるかに興味深いです。
エージェントたちは現在、生産環境でのインシデント調査、プルリクエストのレビュー、質問への回答、ビジネスデータの分析、サポートチケットのトリアージ(優先度付け)、営業先アカウントのリサーチ、そしてReplit Agent自体を動かすシステムの改善などを行っています。
まるで一人のマスター知能が全従業員に浸透しているかのような感覚がありますが、実際はそうではありません。これは会社全体で展開される拡張されたエージェントシステムです。人々から目標を受け取り、文脈を集め、作業を実行し、結果を検証し、人間の判断が必要な場合はエスカレーションします。
私たちはこれが、新しい組織形態の始まりだと考えています。それが「自律型企業(セルフ・ドライビング・カンパニー)」です。
自律型企業が人間を不要にするわけではありません。人間が依然として目的地を選びます。どの問題に優先順位をつけるか、難しいトレードオフをどう決着させるか、審美性やセンスを発揮するか、そして結果に対する責任を負うかは人間が決めます。
しかし、その目標に至るために必要なすべての手順を人間が一つずつ実行するわけではありません。
この転換点は昨年後半に始まりました。AI分野で働く多くの人々と同じく、私たちはクリスマスの休暇から戻り、何かが根本的に変化したと感じました。モデルは、はるかに長い時間軸にわたって作業を継続できるようになっていたのです。
繰り返し失敗していたタスク、例えばアラートの選別や根本原因の調査などが機能し始めました。AI は、これまで解決が難しかったバグの一部を処理できるようになりました。そこで私たちは、エージェントをエディターやチャットウィンドウ内に存在する単なるツールとして扱うのをやめました。慎重に、そして確実に、それらを企業の骨格そのものに織り交ぜていったのです。
エンジニアリング部門で価値が実証されると、採用は自発的に広まりました。チームは次々と最も退屈な作業を任せるようになり、ビジネスを前進させるための戦略的思考や創造的な活動に時間を戻すことができました。人々は「自動化された」と感じるのではなく、「昇進した」と実感しています。
これが、AI が Replit での働き方を根本から変えた物語です。
エンジニアリング部門がまず影響を実感
1 月下旬末、私たちは内部のエージェント活用事例を迅速に実験できるようインフラの規模を拡大しました。エージェント用ハッチネス(harness)やマイクロ VM、リモートファイルシステム基盤を活用し、どのエンジニアでも並列で多数のエージェント群をオーケストレーションできる体制を整えました。その後、アクセスポリシー、トークンプロキシ、監査ログ、そしてゼロトラストネットワークによって全体を厳格に管理しました。これにより、業務遂行に必要な GitHub、GCP、Azure、Linear、Notion、Slack、ZenDesk など、あらゆるツールへのエージェントのアクセスを安全に行えるようになったのです。
システム間での文脈(コンテキスト)を共有できるようになったことで、生産性には飛躍的な向上が見られました。以前は失敗していた実験が容易になり、最も即座に現れた影響はコーディング統計において確認できました。
3 月の Agent 4 リリースを控えたスプリント週は、例年なら生産性が急上昇する時期です。会議がなくなり、範囲も明確になり、エンジニアリングチームは純粋な実行モードへ移行します(1 日最大 16 時間働くことも珍しくありません)。しかし今回は様子が違いました。私たちがこれまで見たことのない形で生産性曲線が上方にカーブし、その要因は新しい社内エージェントシステムの導入にあります。1 月初旬から 6 月末までの間に、貢献されたコード行数は 5.8 倍に増加しました。

この増加の一部は、優秀な人材の採用によるものです。新しいエージェントが即戦力となるまでの時間を短縮してくれたのは素晴らしいことですが、より純粋なデータを見るために採用効果の影響を除外してみましょう。著者(エンジニア)のコホートを一貫して維持した上で比較すると、以前と比べて 2.9 倍のコード量が生成されています。従来、チーム規模を拡大しても一人あたりの生産性を一定に保てれば「優秀」と評価されるのが通説です。しかし私たちは、チーム規模を 2 倍にしつつ、一人あたりの生産性も 3 倍に引き上げました。

この新しいコードを誰がレビューしているのか、あるいはレビュープロセスに新たなボトルネックが生じていないか気になるかもしれません。しかし、当社のコードレビューの遅延は横ばいであり、その主な理由はエージェントがレビュー業務を担当しているからです。現在、エージェントはリスクレベルを評価し、必要に応じてのみ第 2 の人間によるレビューを要請するようになっています。これにより、人間の PR(プルリクエスト)レビューに要する時間が 30% 削減され、さらにその割合は増加傾向にあります。

エージェントがより多くのコードを記述・レビューするようになると、品質への懸念が生じるのも当然です。しかし、PR の巻き戻し率(左側のグラフ)や発生したインシデント数の推移を見れば、どちらも横ばいであることがわかります。これは相対的に見ると、実際には品質が向上していることを意味します。

その理由の一つは、これらのプロセスもエージェント支援で行われている点にあります。人間のコードレビューには「エージェントによる共レビュー」という利点があり、これによりバグの発見率が向上しています。また、インシデント調査(意味のあるバグや実際の事象)においても、根本原因の特定を試みるエージェントが支援するため、平均緩和時間(MTTM)は短縮されています。

最終的な試練は、追加されたコード入力が実際に価値ある成果を生み出しているかどうかです。結局のところ、エンジニアリングの本質はユーザーのために機能を届けることにあります。私たちは Linear でプロジェクトを追跡しており、営業やマーケティングチームが新機能についてユーザーに伝えるべきタイミングを把握できるようにしています。その結果、コーディング量の増加とともに、プロジェクトの完了率が劇的に向上していることが確認できます。

自律型エンジニアリングチームは、品質を維持しながらより多くの成果物をリリースできるようになります。
エージェントの集合体が、大規模なループ工学を実現する
拡大して見てみると、その仕組みがどのようなものかイメージできます。エンジニアたちが生成ループを見つけ、検証可能なタスクを完了させるためにエージェント群を送り出すとき、最も劇的な変化が見られます。すべての従業員に、複数のエージェントを起動できる「マネージャー型エージェント」が用意され、あなたが代わってループ上で動作するエージェントの調整が可能になります。このループ処理により、以下のような独特なプルリクエスト(PR)のグラフが生まれました。

あるエンジニアが長期間停滞していた CSS システムの移行を完了し、その知見を共有しました。別のエンジニアは製品のローカライズを可能にする移行プロセスを自動化し、さらに別のエンジニアも不安定なテストのメンテナンスを自動化しました。そして CTO は、PSC と fd のシャットダウンに関連する当社の最も困難なネットワークバグの一つを、エージェント群によって解決しました。私たちが「何が可能か」について持っていた前提はすべて書き換えられました。


最も注目に値する自律型システムの事例は、AI チームによるものです。彼らはユーザーフィードバックを分析し、改善案を提案するとともに、ベンチマークと A/B テストを組み合わせて成果を検証する「継続学習システム」を構築しました。Replit Agent は自ら進化を続けるのです。

「自社開発か、購入か」の議論は変わった
私たちの新しい内部エージェントは、「自社開発か、外部製品か」という議論そのものを変えました。私たちは常に新しい AI ツールを試していますが、外部ソリューションを購入すればスピードアップにつながり、市場を常に見極めることもできます。しかし、自社で開発するほど、そうした外部調査の必要性は減っていきます。
現在、私たちの内部エージェントは「業界をリードしている」とされる製品よりも優れたパフォーマンスを発揮しています。先日も、7 桁(百万円単位)の SaaS ソリューションを廃止しました。Replit で完全に構築された自社アプリの方が優れており、従業員もそちらへ移行したからです。
突然、ツールがまるで「私たちのために作られたか」のように感じられるようになりました。ナレッジベースとの深い統合や、私たちが行ったカスタマイズのおかげで、他のソリューションは劣っているように見えてしまいます。
さらに驚いたのは、内部エージェントが業界特化型の製品さえも凌駕したことです。エンジニアがアラートを仕分けし、インシデントの根本原因を特定するためのツールは、同等の品質を持ちながら、自社エージェント上で稼働させるコストの 10 倍もの費用がかかりました。また、自動化されたペネトレーションテストを行うツールも、自社版の方が脆弱性を多く発見でき、かつコストは 10 倍高かったのです。
両方の自社バージョンは容易に本番環境へ導入され、インシデント対応の平均検知時間(MTTM)を短縮し、重要なシステムを攻撃から守る堅牢性を高めることに成功しました。

まだ学ぶべきことが多く、モデルも日々進化している現状を考えると、これはあくまで始まりに過ぎません。
エンジニアリングを超えて、事業全体へ
自律型企業(Self-driving company)はエンジニアリング部門だけで完結するものではありません。Replit におけるすべての機能が変化しています。
利用の広がりは急速に進み、そのきっかけの多くは Slack のインターフェースによるものでした。他部署の社員も、エンジニアがエージェントにタスクを割り当てている様子を見て興味を持ち、実際に試してみるようになりました。当初最も人気だったのは「質問への回答」です。ナレッジベースとコードベースの現状情報を組み合わせることで、エンジニアの介入を待たずに製品に関する期待値を確認できるようになりました。その後、社員たちは修正が必要なコピーやドキュメントの改善へとつなげることができました。これにより、ユーザーへの対応速度が劇的に向上したのです。


しかし、これはまだ序章に過ぎません。ここから社内のあらゆる部門が新たなスキルや連携機能の提案を次々と持ち寄るようになりました。
最初の大きな突破は、データチームによるものでした。彼らはエージェントにデータウェアハウス上のセマンティックレイヤー(意味層)を提供し、どのテーブルが真のソースとなるか、またそれらがどのように関連しているかを理解させました。
これにより、Replit の誰でもビジネスインテリジェンスに関する質問を行い、信頼できる回答を得られるようになりました。生データからチャートやプレゼン資料を作成することも可能です(この記事内のすべてのチャートも同様です)。データチームは теперь、単純な問い合わせへの対応に時間を割くのではなく、より難易度の高い課題に取り組むことに集中できるようになりました。
最近では、あるプロダクトマネージャーが複雑なローンチ分析を自分自身で実施できるようになっています。これは、エージェントがコードベース内のイベントを理解し、それが顧客データプラットフォーム上でどのように現れるか、そして複雑なサブスクリプション状態とどう結合するかを知っているからです。
営業部門でも同様の効果が見られます。営業開発チームは、汎用的なツールではアクセスできない社内知識を活用して、製品適合見込み顧客の特定と情報補完をエージェントに任せています。これにより、アプローチする際に文脈がより明確になり、成果につながります。
アカウントエグゼクティブ(担当営業)も、顧客との会話に臨む前にこのツールを利用します。誰が最も高い価値を得ているのか、どのプロジェクトが最も活発なのか、契約に対するクレジット使用状況はどうなっているのか——こうした情報を把握するために活用しています。その結果、各顧客ごとにカスタマイズされたブランドイメージの資料としてパッケージ化されます。
自律型営業チームは、顧客との接点をより多く持ち、かつその質も高めています。
マーケティングチームは、エンジニアリング部門や製品部門との会話やドキュメントを基に、単一のプロンプトだけでゼロから製品仕様書をドラフト作成できるようになりました。これにより、すべての会議に参加する必要がなくなり、リリースの準備を早期に開始し、最新情報を常に把握することが可能になります。結果として、チームは計画立案やクリエイティブな活動に時間を割けるようになり、世に出る製品のインパクトを高めることができます。


サポートチームには、問題調査や標準的な手順書の活用といったスキルがエージェントに付与されました。エージェントは、標準的なカスタマーサービスのトーンで回答を提供するか、あるいはチケットの概要と調査結果を添えてエンジニアリング部門へエスカレーションするかを選択できます。この自律型サポート体制により、人間が対応する難易度の高い案件(エスカレーションされたケース)の処理時間が 60% 短縮され、ユーザーはすぐに開発に戻ることができます。

どのケースでも、人間が自動化によって不要になったわけではありません。むしろ、彼らの役割は昇格したのです。「自律型」の企業では、実行する人から指揮する人が生まれ、成果を想定し方向性を設定できる人材こそが活躍しています。それが今、最も価値のある仕事です。
次のステップへ
生産性の向上自体は確かに魅力的ですが、Replit の人々を本当に動かしているのは「技術の民主化」です。
私たちはこの新しい働き方をすべてのユーザーに届けたいと考えています。現在、大規模展開に必要なポリシー、権限管理、セキュリティ、コスト制御といった要素を整え、実現に向けて全力で取り組んでいます。Replit を最も活発に利用するのは、実際にビジネスを構築している起業家やエンタープライズユーザーです。自己運転型の機能には、こうしたユーザーのニーズに応えるためのスケーラブルな安全対策が不可欠です。
その実現のために、私たちは今まさに開発を進めています。
上記のグラフをご覧いただければ、長く待つ必要はないはずです。
原文を表示
In the past six months, engineers at Replit have nearly tripled code output. Review times held steady. Reversions and product incidents have stayed flat. Quality metrics improved, and releases have accelerated. All the typical trade-offs you might expect have not occurred.
While the code is the visible part, what's happening under the surface is much more interesting.
Agents now investigate production incidents, review pull requests, answer questions, analyze business data, triage support tickets, research sales accounts, and improve the systems that power Replit Agent itself.
It feels like a single master intelligence threaded through every employee, even though it is not. It is an expanding system of agents operating across the company: taking goals from people, gathering context, performing work, checking the results, and escalating when human judgment is needed.
We think this represents the beginning of a new kind of organization: the self-driving company.
A self-driving company is not one without people. People still choose the destination. They decide which problems matter, make difficult tradeoffs, exercise taste, and take responsibility for the outcome.
But increasingly, they do not perform every step required to get there.
The shift began late last year. Like many people working in AI, we returned from the Christmas break feeling that something fundamental had changed. Models could sustain work over much longer horizons.
Tasks that had repeatedly failed, like alert triage and root-cause investigation, began working. AI started solving some of our most stubborn bugs. So we stopped treating agents as tools that lived inside an editor or chat window. We wove them, carefully, into the fabric of the company itself.
Once engineering proved the value, adoption took on a life of its own. Team after team started offloading their most tedious work, reclaiming time for the strategic and creative thinking that actually moves the business. People don't feel like they've been automated. They feel like they've been promoted.
This is the story of how AI has completely changed the way we work at Replit.
Engineering saw the impact first
In late January we turned up infrastructure to experiment with internal agent use cases quickly. We leveraged our agent harness, microVMs, and remote filesystem infrastructure so any engineer could orchestrate swarms of agents in parallel. Then we locked the whole thing behind access policies, token proxies, audit logging, and our ZeroTrust network. At that point we felt safe giving the agent access to all the things we use to get our jobs done: GitHub, GCP, Azure, Linear, Notion, Slack, ZenDesk, and more.
With context across systems, we saw a leap forward in productivity. Experiments that previously failed became easy. The most immediate impact was in coding stats.
We were in the sprint week leading up to Agent 4 release in March, where we typically see a big spike. Meetings disappear, scope is known, and engineering shifts into pure execution mode (often for up to 16 hours per day). But this time was different. Our productivity curve bent upward in a way none of us had seen before, which can be traced to the adoption of our new internal agentic system. From early January to late June, there was a 5.8X increase in the lines of code contributed.

Part of this increase can be attributed to hiring well. Our new agent accelerates time to productivity, which is great, but we can remove the hiring effect for cleaner data. Keeping a consistent cohort of authors, we see 2.9x as much code as before. Traditionally, it’s considered excellent if you keep output per engineer flat as you scale a team. We just tripled per engineer rate while doubling the team.

You might wonder who is reviewing all this new code and whether we’ve created a new bottleneck in the review process. Our code review latency is flat, largely because we put our agent to work in reviewing code. It’s now able to assess risk levels and only call in a second human reviewer when necessary. That means 30% (and growing) of human PR review time has been saved.

With our agent writing and reviewing more code, we should be worried about quality. If we look at PR reversion rates (left) and incidents opened, trends are flat. This means we’re actually improving on a relative basis.

One reason is that these processes are also agent assisted. Human code reviews have the benefit of an agentic co-reviewer, so more bugs get caught. Incident investigations (meaningful bugs or actual incidents) are assisted by an agent that attempts to find the root cause, so mean time to mitigation (MTTM) is going down.

The final test is whether additional code inputs represent real value output. At the end of the day, engineering is delivering features for users. We track projects in Linear so that sales and marketing teams know when to communicate with users about new features. You can see the rate of project completion is sharply up along with our coding volume.

A self-driving engineering team can ship more, while raising quality at the same time.
Our agent of agents is enabling loop engineering at scale
Zooming in gives us an idea of what this looks like. When engineers find ways to generate loops, sending a fleet of agents off to complete a verifiable task, we see the most dramatic change. Every employee gets access to a manager agent that can spawn multiple agents, enabling orchestration of agents working in loops on your behalf. Loops resulted in some very unique looking PR graphs, like these:

One Engineer completed a long stalled migration of our CSS system and shared his learnings. Another engineer automated a migration that enabled us to localize the product. Yet another automated flaky test maintenance. Our CTO finally cracked one of our hardest networking bugs related to PSC and fd shutdown with a swarm of agents. All of our assumptions about what is possible have changed.


The most exciting self-driving example comes from our AI team. They built a continual learning system that analyzes user feedback, proposes improvements, and uses a combination of benchmarks and A/B tests to validate the wins. Replit Agent is self improving!

The build vs. buy conversation has changed
Our new internal agent also changed conversations about whether we build or buy software. We regularly try out new AI tooling. Buying solutions can help us go faster, and we also assess the market constantly. But the more we build, the less of this we will need to do. Our internal agent now outperforms products we test that are seen as market leading. We just churned a seven-figure SaaS solution because our internal app, built entirely in Replit, was superior and employees had migrated over.
All of a sudden, tools feel like they are built for us. The deep integration with our knowledge bases, and customization we’ve done, makes other solutions feel inferior.
What surprised us more was that our internal agent also beat out vertical specific products we evaluated. A tool to help engineers triage alerts and root cause incidents came back with similar quality but at 10x the cost of running it on our agent. A tool that runs automated penetration testing found fewer vulnerabilities than our internal version at 10x higher cost. Both our versions were put into production with ease, reducing MTTM in incidents and hardening critical systems against attacks.


With how much we’re still learning, and how models are improving, it’s clear this is only the beginning.
Beyond engineering and into the whole business
A self-driving company doesn’t stop at Engineering. Every function at Replit is changing.
Usage spread quickly out of Engineering, mostly because of a Slack interface. The rest of the company noticed engineers tagging our agent with tasks and tried it for themselves. Initially, the most popular use case was asking questions. By combining our knowledge base with the state of the code base, anybody could clarify product expectations without waiting for engineering input. Those employees could then fix copy or documentation as a follow up. It was an immediate boost in being able to respond to users faster.


But that was just the beginning. From there, contributions of new skills and integrations started to come in from all parts of the company.
The first big unlock came from our data team. They gave the agent a semantic layer over our data warehouse, so it knows which tables are sources of truth and how they relate to one another.
Now anyone at Replit can ask business intelligence questions and get a reliable answer. They can build charts and presentations from live data (including every chart in this post). The data team spends its time going deeper on the hardest problems, instead of fielding requests. Recently, a PM was able to self-serve complex launch analysis because our agent understands events in the codebase, how they show up in our customer data platform, and how to join those with complex subscription states.


Sales found the same leverage. The sales development team uses the agent to find and enrich product qualified leads, drawing on internal knowledge that more generic tools can’t see, so outreach lands with more context. Account executives use it to prepare for customer conversations to understand who is getting the most value, what projects are most active, and how credit usage tracks against their contract. This is all then packaged up into branded slides customized to the account. A self-driving sales team has more, higher quality touchpoints with their customers.




Our marketing team can use the agent to draft product specs from scratch with a single prompt, based on conversations and documents products across engineering and product. This gives them the ability to start moving on launches sooner and stay up to date, without needing to be in every single meeting. They have more time to plan and be creative, which will ensure our releases have greater impact when they are out in the world.


Our support team gave the agent skills to investigate issues and follow standard playbooks. It can choose to offer a response in our standard customer service voice, or escalate to engineering along with a summary of the ticket and investigation. A self-driving support team closes the hardest tickets (those escalated to humans) 60% faster. Users get back to building sooner.

In every example, the human didn't get automated out. They got promoted. Self-driving turns doers into directors, and the people thriving are the ones who think in outcomes and set direction. That is the most valuable work there is now.
Where to next?
Making ourselves more productive is exciting, but what really motivates the people at Replit is democratizing technology.
We want to bring this new way of working to all of our users. We’re hard at work making sure we can do this with the policy, permissions, security, and cost controls needed to deploy this at scale. Replit’s most active users are entrepreneurs and enterprise users building real businesses. Self-driving needs safety measures that can scale to meet those users.
We’re hard at work building that now.
Given all the graphs above, you won’t have to wait long.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み