2026 年中時点の CLI コーディングエージェントの現状(37 分読了)
本文の状態
日本語全文を表示中
詳細モードで約45分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
2026年半ば時点のCLIコーディングエージェント市場は、IDE中心からターミナルやCI環境へシフトし、AnthropicやOpenAIなどの主要プレイヤーが標準化を競う過熱状態にある。
AI深層分析を開く2026年8月4日 17:34
AI深層分析
キーポイント
利用環境の劇的変化
2024年のIDE中心から2026年半ばにはCI、SSH、GUIのないマシンでの利用が主流となり、スクリプト性とウィンドウマネージャとの競合回避が重視されるようになった。
市場の過熱と多様化
2026年7月時点で35ものアクティブなCLIコーディングエージェントが存在し、CursorやAmpなどの大手から12のオープンソースチームまで参入が殺到している。
標準化とエコシステムの形成
Linux Foundation が「Agentic AI Foundation」を設立し、Anthropic の Model Context Protocol や OpenAI の AGENTS.md などを基盤とした業界標準の確立が進んでいる。
主要プレイヤーの技術的進化
Anthropic の Claude Code がアジェンティックループやプランモードを定義し、OpenAI や Google も CLI ツールを相次いでリリースして機能競争を展開している。
Agentic AI Foundation (AAIF) の設立と標準化の動き
2025年12月にLinux FoundationがAnthropic、OpenAI、Blockらと共にAAIFを設立し、MCPやAGENTS.mdなどのプロトコルを基盤とした業界標準を推進した。
重要な引用
By mid-2026 the serious usage ran from CI, SSH, and machines with no GUI at all.
35 actively maintained CLI coding agents as of July 3, 2026. Crowded field.
Anthropic's Claude Code research preview in February 2025 set the shape: agentic loop, file and shell tools, project memory file, permission prompts, plan mode, hooks, subagents.
December 2025 is when the standards people showed up: the Linux Foundation formed the Agentic AI Foundation (AAIF), anchored by Anthropic's Model Context Protocol, OpenAI's AGENTS.md convention, and Block's Goose agent
編集コメントを表示
編集コメント
2026 年半ばという未来の視点から、CLI エージェントが IDE の壁を破り、インフラ層で定着した様子が浮き彫りになっている。標準化団体による仕様の統一が進む中、開発者は自社のワークフローに最適なエージェントを選別する判断力が求められる時代へと突入している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ターミナルは予想外の勝者となりました。2024 年の時点では、IDE が主流になると見られていました。Copilot はすでにエディタに組み込まれ、Cursor の台頭も目立ち、「エージェント」という言葉はまだ、シェルアクセスを付加したチャット機能を指すものでした。しかし 2026 年半ばには、本格的な利用は CI や SSH、あるいは GUI を備えていないマシン上で行われるようになりました。スクリプト化が可能で地味な存在である一方、モデルがウィンドウマネージャと競合することもありません。
2026 年 7 月 3 日現在、活発にメンテナンスされている CLI コーディングエージェントは 35 あります。市場は非常に混雑しています。
短い歴史
第 1 波が到来したのは、これらを「エージェント」と呼ぶ前年の 2023 年でした。gptme(2023 年 3 月)は、モデルがターミナルからシェルコマンドを実行できるようにしました。Aider(2023 年半ば)は、変更の単位をアトミックなコミットとして git を中心に AI ペアプログラミングを実現しました。また Open Interpreter(2023 年 7 月)はこの考え方を一般化し、コンピュータ全体を制御する仕組みへと発展させました。これら 3 つは現在も生き残っており、gptme はデーモンとして、Aider はペアプログラマーとして、Open Interpreter は汎用的なコンピュータコントローラーとしてそれぞれ役割を果たしています。
2025 年 2 月、Anthropic が発表した Claude Code の研究プレビューが、その後の CLI コーディングエージェントの姿を決定づけた。アジェンティック・ループ、ファイルやシェルのツール、プロジェクトメモリ機能、権限確認のプロンプト、プランモード、フック、サブエージェント——これらが標準的な構成となったのだ。他社もこぞってこれを模倣した。
OpenAI は 2025 年 4 月に Codex CLI をリリース(後に Rust で書き直し)、Google も同年 6 月に Gemini CLI を発表し、積極的な無料枠を設けた。2025 年後半には、Cursor、Amp、Augment、Factory、Charm、そしてオープンソースチーム 12 社を含む多数の製品が、わずか 1 クォーターで相次いで登場した。
2025 年に Gemini CLI を標準化して導入したチームの多くは、昨秋の設定を引き継いでいます。しかし、2025 年 12 月になって業界の標準化を主導する動きが本格化しました。
この月、Linux Foundation は「Agentic AI Foundation(AAIF)」を設立し、Anthropic の Model Context Protocol、OpenAI の AGENTS.md 規約、Block の Goose エージェントを中核に据えました。Google、Microsoft、AWS、Cloudflare、Bloomberg が支援しています。
同じく 12 月には、Sourcegraph がエージェント事業を独立させ Amp Inc. を設立。Mistral は Vibe と Devstral 2 モデルで参入し、Codebuff もマルチエージェントを支援するハッチをオープンソース化しました。
2026 年前半は、業界全体がやや混沌とした状況でした。2 月には GitHub Copilot CLI が一般提供(GA)に到達し、Cline は CLI 2.0 をリリースして独自 SDK をオープンソース化しました。また Kilo CLI もバージョン 1.0 に達しています。
6 月には JetBrains の Junie がベータ版から GA へ移行しました。一方、5 月の I/O で Google が注目を集めました。Antigravity CLI のローンチと、Gemini CLI の終了(無料およびコンシューマープラン利用者は 6 月 18 日)が発表されたのです。
同様に 5 月末には x.ai の Grok Build が登場し、6 月には Moonshot が従来の CLI から Kimi Code に切り替えました。
モデルの進化は止まらなかった。4 月に MIT ライセンスで公開された「DeepSeek V4」、6 月には「GLM-5.2」や「Kimi K2.7 Code」が相次いで登場し、オープンウェイトモデルの性能は最先端のターミナルスコアに僅差まで迫った。一方、無料枠の縮小もこの時期に加速した。「Qwen Code」のホスト版無料プランが 4 月に終了し、「Amp Free」は待機リストへ移行。さらに「Gemini」に至っては、夏までに一般消費者向けの利用経路が姿を消している。
モデル開発ラボ
モデル開発ラボから登場するエージェントは、モデルベンダー自身が提供するものだ。ハルネスとモデルが一体化したパッケージで、通常は既存のサブスクリプションプランに追加項目として含まれている。
| エージェント | 開発者(メンテナ) | デビュー | ソースコード | モデル | アクセス方法 | 特徴・評価点 |
|---|---|---|---|---|---|---|
| Claude Code | Anthropic | 2025 年 2 月 | クローズド | Claude のみ | Pro/Max プランまたは API | カテゴリのテンプレート;最も深いオーケストレーション機能 |
| Codex CLI | OpenAI | 2025 年 4 月 | Apache-2.0、95k stars | OpenAI Codex モデル | ChatGPT プランまたは API | Rust コア、OS レベルのサンドボックス化、クラウドへのハンドオフ |
| Gemini CLI | 2025 年 6 月 | Apache-2.0、106k stars | Gemini | API キー / 2026 年 6 月 18 日以降はエンタープライズのみ | 一般消費者向けアクセスは終了し、Antigravity CLI に移行 | |
| Antigravity CLI | 2026 年 5 月 | クローズド | Gemini 3.5 および選択されたサードパーティ製モデル | 無料のパブリックプレビュー | Antigravity 2.0 デスクトップアプリとハッチを共有 | |
| Grok Build | xAI | 2026 年 5 月 | クローズド | grok-build (コンテキスト 256K) | SuperGrok 30/月、XPremium+30/月、X Premium+ 40/月、または API | 隔離された git worktree で最大 8 つの並列サブエージェント |
| Mistral Vibe | Mistral AI | 2025 年 12 月 | Apache-2.0、4.6k stars | Mistral (Medium 3.5, Devstral 2); オープンウェイト | 無料 CLI; Mistral プラン/API | 欧州向けオプション;リモート非同期エージェント;ACP |
| Kimi Code CLI | Moonshot AI | 2026 年 6 月 | MIT、3k stars | Kimi K2.7 Code | Kimi プランまたは低コスト API | 組み込みの coder/explore/plan サブエージェント;ACP |
| Qwen Code | Alibaba (Qwen) | 2025 年 7 月 | Apache-2.0、26k stars | Qwen3-Coder または任意のエンドポイント | BYOK; 有料プラン(無料ホスト版は 2026 年 4 月に終了) | Qwen のオープンウェイトコーダー向けにチューニングされた Gemini CLI フォーク |
「Claude Code」はまだ、エージェントチームやフック、スキルといった用語の基準となっています。複数のセッションが互いに通信する実験的なエージェントチームは、依然として他をリードしています。その代償は、モデルへの完全なロックインです。2026 年 6 月現在、ヘッドレス環境や SDK の利用料金は、別のクレジットプールから請求されます。7 月初旬に公開された「Terminal-Bench 2.0」のリーダーボードでは、Anthropic の最新モデルが首位を維持しています。同社の強みは、共学習されたモデルとハッチ(実行基盤)の組み合わせであり、この実績はリーダーボードの結果によって裏付けられています。
一方、「Codex CLI」は最も強力な対抗馬として、春先から活発に動き続けています。オープンソースで OS サンドボックス環境に対応し、ChatGPT のサブスクライバーベースを活用しています。現在の最前線デフォルトには「GPT-5.5」を採用しており、2026 年前半にはトークン予算を伴う永続的なゴール機能や、サブエージェントへのスレッドレベルでの委任、プラグインマーケットプレイス、ブラウザ操作、暗号化されたリモート実行機能を追加しました。さらに、Claude Code の設定を一発でインポートするワンライナーも用意され、スイッチングコストが現実的なものとなったため、Codex 側はこれを自動化せざるを得ませんでした。
2025 年に Gemini CLI を標準ツールとして採用していた人にとって、6 月 18 日は悪夢のような一日でした。Google はこのプロジェクトを同カテゴリで最もスター数の多いリポジトリの一つに成長させましたが、無料プランや一般ユーザー向けには、クローズドソースの Antigravity CLI へと移行するよう案内しました。
Antigravity は Go で書き直されたもので、デスクトッププラットフォーム「Antigravity」と共通の基盤を共有しています。コアとなるのは Gemini 3.5 です。非同期で複数のエージェントが連携するワークフローを実行可能で、現在はパブリックプレビューとして無料で利用できます。メニューには Google 製以外のモデルも含まれています。
無料プランがあるからといって、それが基盤(ファウンデーション)であるわけではありません。
Grok Build、Kimi Code、Mistral Vibe は登場が遅れましたが、プランニングモードやサブエージェント、ヘッドレス CI への対応はすでに完了しています。Grok Build はベータ版リリース初日から、プランニングモード、ワークツリーを分離した並列サブエージェント、そしてヘッドレス CI サポートを実装して出荷されました。Kimi Code と Mistral Vibe はどちらも、非常に安価で一部がオープンなモデルに依存しています。Mistral の Devstral 2 は、123B パラメータのモデルからベンダー報告値として SWE-bench Verified で 72.2% を達成しました。また、6 月中旬に重み付きを公開した Moonshot の K2.7 Code は、最先端モデルの価格を桁違いに下回っています。両者とも ACP(Agent Communication Protocol)を採用しており、互換性のあるエディタであればどこでもホスト可能です。Qwen Code はフォークの事例です。Gemini CLI を派生させて Qwen のオープンウェイトモデル向けに再調整したもので、無料枠が終了してもツールが Apache ライセンスでありエンドポイント非依存であるため、現在も利用可能です。
プラットフォームおよびプロダクト CLI
プラットフォームベンダーは、エージェントを開発スタックの一部として提供しています。つまり、配布、統合、ガバナンスを含む形です。これらほぼすべてがマルチモデル対応となっています。
| エージェント | 維持管理者 | デビュー | ソースコード | モデル | アクセス方法 | 知られている点 |
|---|---|---|---|---|---|---|
| GitHub Copilot CLI | GitHub | 2025 年 9 月、GA は 2026 年 2 月 | クローズド | Anthropic, OpenAI, Google, オープンウェイトの Kimi K2.7 | 有料 Copilot プラン(月額 10 ドルから) | 自動委譲型専門エージェント;& で作業をクラウドエージェントに引き渡す |
| Cursor CLI | Anysphere | 2025 年 8 月 | クローズド | Cursor Composer + フロンティアモデル | Cursor プラン | Cursor IDE と同じエージェント、ルール、MCP 設定 |
| Amp | Amp Inc.(元 Sourcegraph) | 2025 年、2025 年 12 月にスピンアウト | クローズド | 厳選されたフロンティアモデルのミックス、モデル選択機能なし | 従量課金;広告掲載による無料枠(ウェイトリストあり) | あえて意見が明確;広告収益化の実験 |
| Auggie CLI | Augment Code | 2025 | クローズド | Augment を介したフロンティアモデル | Augment プラン | 作業開始前にリポジトリ全体をインデックス化するコンテキストエンジン |
| Droid | Factory AI | 2025 | クローズド | マルチモデル、BYOK 対応 | Factory プラン | オリジナルの Terminal-Bench で首位;CLI、IDE、Web、Slack、チケット管理を跨ぐ |
| Junie CLI | JetBrains | ベータ版 2026 年 3 月、GA は 2026 年 6 月 | クローズド | OpenAI, Anthropic, Google, xAI | JetBrains AI プラン | ACP ネイティブ;エージェント型デバッグ;IDE とデータベースの統合 |
| Qoder CLI | Alibaba | プラットフォーム 2025 年 8 月;CLI は 2026 年 | クローズド | Alibaba Model Studio | 従量課金 / コーディングプラン | 仕様駆動型の自律タスク向けのクエストモード |
| CodeBuddy Code | Tencent | 2025 年 9 月 | クローズド | DeepSeek + Hunyuan | CodeBuddy プラン | スキル、プランモード、ACP、サンドボックス実行;Tencent 社内 12,000 ユーザー |
「Copilot CLI」は、その圧倒的な普及度で頭一つ抜けています。2026 年 2 月に一般提供(GA)され、月額 10 ドルから始められるエントリープランを用意。GitHub の MCP サーバーが標準搭載されており、リポジトリからのプラグイン対応や、リポジトリごとのメモリ管理機能も備えています。さらに「探索」「計画」「レビュー」「構築」に特化した専門エージェントも用意されています。
プロンプトの先頭に & を付けるだけで、ジョブは GitHub のクラウド上のコーディングエージェントへ転送されます。GitHub 中心のチームにとっては、最も手軽なデフォルト選択肢であり、モデル選択機能のおかげで、大手クローズド型エージェントの中でも特にベンダー中立性が高いと言えます。7 月 1 日、Copilot は Azure 上で Moonshot のオープンウェイトモデル「Kimi K2.7 Code」を追加しました。これは主要なクローズドプラットフォームのエージェント内で採用された初のオープンウェイトモデルです。
Copilot に続く形で、市場の注目は分岐しました。
Cursor CLI は「一貫性」を訴求する製品です。IDE 内、ターミナル、CI 環境すべてで同じエージェントが、同じルールと MCP 設定で動作します。1 つの AGENTS.md ファイルがこれら 3 つの領域を統括することが、同社の製品全体の核となっています。
一方、Amp はモデル選択機能そのものを廃止し、チームの判断に基づいてモデルの組み合わせを切り替えるアプローチを取ります。広告収入で成り立つ「Amp Free」プラン(スポンサー支援により学習データへの使用保証なし)は、リストアップされたサービスの中で最も特異なビジネスモデルと言えるでしょう。ただし、現在は利用登録がクローズされています。
Auggie は、最初のプロンプトを送る前にリポジトリ全体をインデックス化することでコンテキストの初期設定に注力しています。このアプローチはレガシーなモノリスでは効果を発揮しますが、新規プロジェクト(グリーンフィールド)ではその恩恵は限定的です。
Droid はエンタープライズ向けのバンドル製品で、専門的なエージェントと高度な並列処理を特徴とし、Slack やチケット管理システムとの連携機能も備えています。Factory 社は 2025 年末に発表された結果において、あらゆるプロダクトハッチの中で最も高い Terminal-Bench のスコアを記録しました。
Junie は、ACP(Agent Communication Protocol)を通じて JetBrains のデバッガやデータベース接続機能をターミナル環境へ統合します。
Qoder と CodeBuddy は中国市場向けのスタック製品です。2026 年版の機能リストは他社と共通していますが、ファイアウォール内での流通経路が異なります。
オープンソースのハッチ
「Bring-your-own-key」型のエージェントが最大勢力を占めており、オープンウェイトモデルの登場で実行コストが大きく変化しました。GLM、DeepSeek、Qwen3-Coder、Devstral、Kimi などのトークンを利用する適切なハッチを使えば、サブスクリプション価格の数分の一で最先端の能力を手に入れることができます。
この格差は6月に急激に縮まりました。Z.ai が MIT ライセンスの下で GLM-5.2 をリリースしたのです。これは約400億パラメータが活性化する7440億パラメータ規模の混合専門家モデル(MoE)で、100万トークンのコンテキスト長を備えています。このモデルは、GPT-5.5 を複数の長期コーディングベンチマークで上回る性能を発揮し、そのコストは約6分の1です(Terminal-Bench 2.1 の分析)。また、最も積極的な量子化を施せば、最大で約245GBのメモリを持つ高スペックなハードウェア上で完全にオフラインでも動作可能です。
| エージェント | メンテナ | デビュー | ソース | モデル | 知られている点 |
|---|---|---|---|---|---|
| OpenCode | Anomaly (SST チーム) | 2025 | MIT, 182k stars | 75+ プロバイダ | GitHub で最もスター数の多いエージェント; TUI、デスクトップ、IDE; エージェントとスキル |
| Crush | Charmbracelet | 2025 年 7 月 | FSL-1.1-MIT (ソース利用可能), 26k stars | マルチプロバイダ対応 | カテゴリ内で最も精巧に作られた TUI; LSP コンテキスト; MCP |
| Goose | AAIF (元 Block) | 2025 年 1 月 | Apache-2.0, 51k stars | 任意、ローカル含む | MCP ネイティブ、ローカルファースト、財団ガバナンス型 |
| Aider | Aider-AI | 2023 年中盤 | Apache-2.0, 47k stars | LiteLLM を介して 100+ | Git ネイティブのペアプログラミング; リポジトリマップ; 開発ペースは鈍化 |
| Cline CLI | Cline | 2025 年末; 2.0 は 2026 年 2 月 | Apache-2.0, 64k stars | 任意のプロバイダ | Open @cline/sdk ランタイム; 並列エージェント; ヘッドレス CI |
| Kilo CLI | Kilo | 1.0 は 2026 年 2 月 | オープンソース, 26k stars | 500+ モデル | アーキテクト/コード/デバッグ/質問/オーケストレーターモード; メモリバンク |
| Continue CLI (cn) | Continue | 2025 年 9 月 | Apache-2.0, 35k stars | 任意のプロバイダ | ヘッドレス PR チェックおよび CI エージェント; 開発ペースは鈍化 |
| OpenHands CLI | OpenHands | CLI は 2026 年 5 月; プロジェクトは 2024 年 | オープンソース, 79k stars | 任意のプロバイダ | イベントソーシング再生機能を持つエージェント SDK; CLI は安定した軽量クライアントとして |
| DeepSeek-Reasonix | esengine | 2026 年 4 月 | MIT, 26k stars | DeepSeek V4 (デフォルト) または任意のエンドポイント | キャッシュファースト設計; 極端な予算効率性 |
| Every Code | JustEvery | 2025 年 8 月 | Apache-2.0, 3.8k stars | OpenAI, Anthropic, Google, 任意 | 複数のベンダーのモデルを同時にオーケストレートする Codex フォーク |
| ForgeCode | Antinomy | 2024 年 12 月 | Apache-2.0, 7.4k stars | 300+ モデル | Rust、シェルネイティブ、セマンティックコードベース検索 |
| Codebuff | Codebuff AI | 2024 年; オープンソース化は 2025 年末 | Apache-2.0, 7.1k stars | OpenRouter を介して任意 | 明示的なエージェント役割: ファインダー、プランナー、エディタ、レビュアー |
| Kode | shareAI-lab | 2025 年 7 月 | Apache-2.0, 5.2k stars | マルチモデルコラボレーション | タスク中に特定のモデルに @ask |
| Nanocoder | Nano Collective | 2025 年 7 月 | オープンソース, 2.2k stars | ローカルファースト + 任意のエンドポイント | コミュニティ所有、テレメトリなし、背後に企業なし |
OpenCode は、182 万の GitHub スターを記録し、Gemini CLI や Codex を引き離して同カテゴリで最多となっています。75 社以上のプロバイダーに対応し、クライアント・サーバーアーキテクチャを採用しているほか、洗練された TUI(テキストユーザーインターフェース)とデスクトップ、IDE 用のインターフェースも提供しています。
ただし、そのフォークの経緯は複雑です。2025 年半ばに原作者が Charm に合流し、その系譜は Crush(より洗練された TUI と、ソースコード公開ライセンスである FSL ライセンスを採用)へと発展しました。一方、SST チームが引き継いだ OpenCode も成長を続けています。一つのプロジェクトから生まれた混乱した分断の結果、2 つの健全なプロジェクトが誕生したのです。
DeepSeek-Reasonix は急上昇中のツールです。たった 1 つの静的 Go バイナリから、約 10 週間で 26 万スターを獲得しました。その設計思想は、推論コストをエンジニアリングの問題として捉える点にあります。セッション中は DeepSeek のバイト単位の安定したプレフィックスキャッシュをホット状態に保ちます。具体的には、起動時に環境サマリーを安定させ、圧縮前に古くなったツール出力を削除し、プランナーとエグゼキューターをそれぞれキャッシュが安定したスレッド上で動作させることで最適化を図っています。
ある報告によると、1 日あたり 4.35 億トークンの入力を処理し、キャッシュヒット率は 99.82% に達しました。未キャッシュ推定値は約 61 です。OpenAI 互換のエンドポイント、stdio および HTTP を介した MCP(Model Context Protocol)、チェックポイント機能、巻き戻し機能、そして Feishu、Lark、WeChat とのチャット操作ブリッジをサポートしています。同じく 4 月に MIT ライセンスで登場した DeepSeek V4 の存在も、この成功を阻害するものではありませんでした。
Goose は耐久性を高めるために異なるアプローチを取りました。ブロックは Goose を「Agentic AI Foundation」に寄付し、同基金が管理する主要なエージェントとして唯一の存在となりました。これは、MCP や AGENTS.md といった標準規格自体と同様に、中立な基盤の下で運営されることを意味します。
Cline の重要性はユーザー数だけではありません。2026 年 5 月に Apache ライセンスの SDK が抽出されたことで、人気製品が他社が構築できるインフラへと変貌したからです。GitHub(Copilot SDK)、Anthropic(Agent SDK)、そして OpenHands(決定論的な再生機能を持つ研究グレードの Python SDK)も同様の動きを見せました。
Kilo と Continue は IDE 拡張機能をターミナルへ拡張しています。Kilo はオーケストレーションモードと 500 以上のモデルに対応し、Continue は CI を重視したアプローチを採用しています。ただし、2026 年のコミット頻度から見て、Continue は要チェックリスト入りとなっています。
「Aider」は、他社が追随した Git の運用規範の半分を創出し、その多言語ベンチマークは 2 年間にわたりモデル評価の基準となりました。しかし、市場がエージェント指向へと移行する中、Aider はペアプログラミングツールとしての地位に留まり続け、リリース頻度は極端に低下しました。原子単位のコミットにおいては依然として強力ですが、カテゴリ全体は別の方向へ進んでしまいました。
長尾(ニッチ層)の規模は、表立っている主要ツールが示す以上に大きいです。Code は自社の上位リポジトリを上回る機能を備え、Codex ベースで OpenAI、Anthropic、Google のモデルをオーケストレーションしています。「Codebuff」は作業を「発見者」「プランナー」「エディタ」「レビュアー」という明確なエージェントに分解し、独自の実行 175 タスクの評価において Claude Code の 53% を上回る 61% の性能を達成したと主張しています。Kode は、セッション中に特定のモデル名を指定して参照できる機能を提供します。Nanocoder は、共同所有でテレメトリ(利用状況追跡)がなく、ローカルファーストな環境を求める人々向けです。背後に企業はありません。
最小コアと先駆者
| エージェント | メンテナー | デビュー | ソース | 知られている点 |
|---|---|---|---|---|
| Pi | Earendil (Mario Zechner) | 2025 年 8 月 | MIT, 67k stars | 4 つのツール、1,000 トークン未満のシステムプロンプト、TypeScript 拡張 SDK |
| oh-my-pi | can1357 | 2025 年 12 月 | MIT, 16k stars | 最大規模の Pi ディストリビューション;カテゴリ内で最大の機能範囲を有する |
| mini-SWE-agent | SWE-agent チーム | 2025 年 6 月 | MIT, 5.6k stars | 約 100 行の Python、bash のみ;研究における対照群 |
| gptme | gptme プロジェクト | 2023 年 3 月 | MIT, 4.4k stars | 本リストで最古のツール;git ベースのメモリを備えた永続的な自律エージェント |
| Open Interpreter | Open Interpreter | 2023 年 7 月 | オープンソース, 64k stars | LLM によるコード実行を先駆;より広範なコンピュータ制御;2026 年は静か |
「Pi」は反動の象徴です。Mario Zechner氏は、現代のモデルには「読み書き」「編集」「bash 実行」という 4 つのツールと、1,000 トークン未満のシステムプロンプトがあれば十分であり、それ以外はいらないと主張しています。MCP(Model Context Protocol)も TypeScript の拡張 SDK に含めるべきです。このプロジェクトは公開から 1 年足らずでスター数 67 万に達し、ツール集約への飽きが生じていることを示しています。
一方、「oh-my-pi」(略称:Omp)はその対極をなす賭けです。これは「Pi」のディストリビューション版ですが、機能を追加し続け、最終的に比較対象の中で最も豊富な機能セットを持つようになりました。同じコミュニティから生まれた両者ですが、方向性は正反対です。「Pi」は削ぎ落とすことで勝負し、「Omp」は積み上げることで勝負します。どちらもまだ支持層を持っています。
「mini-SWE-agent」はコントロールグループです。約 100 行の Python スクリプトで構成され、bash のみを使用していますが、最先端モデルを使えば SWE-bench の大きな割合をクリアできます。これに勝てないのに 50 もの機能を持つハーンネス(環境)を使う正当性は薄いです。
「gptme」と「Open Interpreter」は古参の枠組みです。「gptme」はまだ活発に開発が続けられており、永続的で自己改善型のエージェントに焦点を当てています。一方、「Open Interpreter」は歴史的意義は大きいものの、明らかに開発ペースが鈍化しています。
機能ごとの比較
2025 年までに、本格的なエージェントはすべて同じ骨格を持っていました。「編集 - 実行 - テスト」のループ、プロジェクト指示ファイル、MCP、プランニングモード、権限プロンプト、ヘッドレスモードです。もはやこれらの要素が誰かを区別することはありません。
マーケティングページでまだ隠されているのは、残りのツールたちです。Claude Code、Codex CLI、OpenCode、Omp、そして大規模展開プラットフォームの選定品である Copilot CLI の 5 つが、ここでは徹底的に扱われます。それ以外のツールは、各表の下の注釈欄に記載されています。
コンテキストとメモリ
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| プロジェクト指示 | CLAUDE.md | AGENTS.md | AGENTS.md + 指示 | AGENTS.md | AGENTS.md、および .claude, .cursor, .codex, .gemini, .cline 設定を直接読み取り |
| セッション間での学習済みメモリ | はい、永続的メモリファイルあり | 部分的:永続的なゴール | はい、リポジトリごとの Copilot Memory | なし、手動の AGENTS.md | はい、Hindsight:プロジェクト SQLite バンク上で保持・想起・反省 |
| 自動圧縮 | はい | はい | はい、95% 以上で自動または /compact で実行 | はい | はい、ビットマップフレーム履歴圧縮あり |
| チェックポイントと巻き戻し | はい、/rewind 対応 | なし | なし | 元に戻す/やり直す | はい、コンテキストの剪定を伴うチェックポイント/巻き戻し |
1 年前、プロジェクトの記憶とはチームの誰かが手動で更新する Markdown ファイルを指していました。しかし現在、リーダー格の 5 つのうち 3 つは自律的に学習しています。Claude Code はセッションを超えてメモリファイルを蓄積し、Copilot はリポジトリごとに理解を構築します。また Omp は実行中に事実を書き込み、必要に応じて後でそれらを統合します。Omp はさらに、他のエージェントがリポジトリに残したルールやスキル、MCP 登録情報も読み取ります。移行は設定への影響を最小限に抑えるだけで済みます。一方、Kilo の Memory Bank はエージェントの状態をリポジトリ内の構造化された Markdown に保存し、gptme は数ヶ月にわたって実行されるエージェント向けに Git を活用したメモリ管理で最も先を行っています。
編集とコードインテリジェンス
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| アンカー付きまたは AST 対応編集 | なし | なし | なし | なし | あり:ハッシュアンカー付きパッチ(出力トークンが約 61% 削減、プロジェクト報告)および 50 以上の文法に対する ast-grep 書き換え |
| LSP 統合 | あり、診断とナビゲーション | なし | なし | あり、内蔵 | あり、深層:診断、参照、コードアクション、再エクスポートを伝播するリネーム |
| デバッガ制御 | なし | なし | なし | なし | あり、DAP:lldb, dlv, debugpy;ブレークポイント、ステップ実行、変数検査 |
| grep を超えたコード検索 | エージェント型 grep | エージェント型 grep | grep + GitHub コード検索 | grep + LSP 記号 | インプロセス ripgrep、tree-sitter 構造的サマリー、ファジーマッチ |
| フォーマッタとリンタの認識 | フック経由 | なし | ドキュメントなし | あり、内蔵フォーマッタサポート | あり、LSP 診断とリンタが意思決定にフィードバック |
Omp(https://omp.sh/)が圧倒的なリードを誇ります。ハッシュアンカー編集は、誰もが遭遇するリファクタリングの失敗——空白文字の変化によってパッチが誤った行に適用される問題——を対象としています。AST を用いた書き換えでは「プレビュー後に承認」のワークフローを採用し、エージェントが操作できるデバッガーも備えています。2026 年 7 月現在、この機能を搭載したラボ製のエージェントは存在しません。Claude Code に比べてセットアップは重くなりますが、モデル自体はユーザーが選択可能です。パッチが常に誤った行に適用されてしまう場合、ここを参照すべきです。
残りのツールたちはコードインテリジェンスに対して異なるアプローチを取っています。Auggie(https://www.augmentcode.com/product/cli)のコンテキストエンジンは作業開始前にセマンティックなインデックスを構築し、Aider はプロジェクト構造を圧縮してプロンプトに埋め込みます。ForgeCode は同期ステップ後に埋め込みベースの検索を行い、Junie は JetBrains のインデクサーを活用します。おそらくエージェントが利用する静的解析の中で最も深掘りされたものですが、オープンソースではありません。
Orchestration
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| サブエージェント | はい、カスタムエージェント定義あり | はい、スレッドレベルの委任(2026 年 6 月) | はい、自動委任される専門エージェント | はい、カスタムエージェントあり | はい、スケーラブルな展開とスキーマ検証済み JSON 返却 |
| 並列作業と分離 | はい、ワークツリー;エージェントチーム(実験的) | クラウドタスク;スレッドごとのトークン予算 | クラウドエージェント経由 | 並列セッション | はい、ファイルシステムクローンによる分離(APFS/btrfs/overlayfs) |
| エージェント間の相互通信 | はい、チームメッセージング(実験的) | なし | なし | なし | はい、ライブエージェント間の IRC 風チャネル |
| 第 2 モデルによる監視 | なし | なし | なし | なし | はい、Advisor:別のモデルが各ターンをレビューし注釈を挿入 |
| バックグラウンドプロセス | はい | 一部対応 | クラウド経由 | なし | はい、ジョブ制御あり |
| クラウドへのハンドオフ | はい、Web およびリモートセッション | はい、Codex クラウド;暗号化されたリモート実行 | はい、& プレフィックス | セルフホストサーバーモード | なし;/collab は代わりにライブセッションを共有 |
| スケジュール実行および定期実行 | はい、スケジュールされたクラウドエージェント | 時間指定リマインダー(2026 年 6 月) | GitHub Actions 経由 | CI またはサーバーモード経由 | なし |
Claude Code と Omp は、エージェントを対等な存在として連携させる仕組みを提供します。片側では「エージェントチーム」が機能し、もう片側ではプロセス内のチャットチャンネルと常駐するアドバイザーモデルが担います。一方、Codex と Copilot は並列処理をクラウド側にシフトさせ、ローカルデバイスの負荷を減らす代わりに運用キューの負担を増やしています。
Grok Build は本表には含まれていませんが、リリース初日からワークツリー単位で隔離された並列サブエージェント(最大 8 つ)を実装しました。Codebuff は役割をハードコードするアプローチを採用し、Mistral Vibe はターミナルを閉じた後もリモートエージェントの稼働を維持します。Amp のオラクル(困難なステップに用いるより強力なモデル)は、Omp のアドバイザーに最も近い存在と言えます。
スケジューリング機能の登場はやや遅れましたが、Claude Code では再発するクラウドエージェントが、Codex では 6 月にタイムリマインダーが実装され、それ以外の環境では CI での対応が進んでいます。
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| MCP クライアント | あり | あり | あり、GitHub MCP 内蔵 | あり | あり |
| プラグインまたは拡張機能 API | あり、プラグインとマーケットプレイス | あり、プラグインマーケットプレイス(2026 年) | あり、リポジトリからの /plugin install コマンド対応 | あり、JS/TS プラグイン対応 | あり、ホットリロード対応の TypeScript 拡張機能 |
| スキル | あり、フォーマットの元祖 | なし | なし | あり | あり、チェーン処理用の入力/出力スキーマ対応 |
| ライフサイクルフック | あり、フォーマットの元祖 | 限定的 | なし | プラグインイベント経由 | あり、生成中のリアルタイムで中止・修正可能な中間ルール対応 |
| カスタムスラッシュコマンド | あり | あり、カスタムプロンプト対応 | ドキュメントなし | あり | あり、拡張機能がコマンドとホットキーを登録可能 |
| 他アプリ向けのサーバー/SDK として機能 | あり、Agent SDK | あり、MCP サーバーモードと SDK | あり、Copilot SDK(2026 年 6 月 GA) | あり、サーバー+SDK | あり、Node SDK と NDJSON RPC モード |
外部ツールには MCP、振る舞いにはプラグイン、指示にはスキル。このスタックが現在ではデファクトスタンダードとなっています。Claude Code はこの 3 つのフォーマットのうち 2 つを考案しました。Omp のストリームルールは特異な存在で、トークンストリームの正規表現が文の途中で停止し、修正を挿入して再開します。これはプロンプトに情報を詰め込むことなくプロジェクトルールを実装する手法です。Goose は MCP 主義者であり、Pi はコアに MCP を搭載せず、すべての機能を TypeScript 拡張 SDK に委ねています。67,000 スターという数字が、このアプローチの支持の高さを物語っています。
Git とレビューワークフロー
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| コミット支援 | はい、git 対応のコミットと PR | はい; クラウドタスクによるコミットと PR 作成 | はい、クラウドエージェントを介した PR ネイティブ対応 | Git ベースの元に戻す/やり直し; 要求時にコミット | omp commit は無関係な変更を依存順序に従った複数のコミットに分割します |
| 組み込みのコードレビュー | はい、/code-review コマンドによる難易度別対応 | はい、/review コマンド | はい、専用レビューエージェント | カスタムエージェントを介して | はい、/review: 並列レビュー、P0-P3 の優先度付け、マージ判定 |
| PR とイシューの統合 | gh コマンドおよび GitHub Actions を介して | GitHub Action および Codex クラウド | プラットフォームネイティブ対応 | GitHub および GitLab 統合 | pr:// および issue:// パスとしてアクセス可能な PR とイシュー; Actions の実行をリアルタイムで監視 |
| マージ競合対策ツール | エージェントのみ対応 | エージェントのみ対応 | エージェントのみ対応 | エージェントのみ対応 | 宣言型: conflict://N に対して @ours, @theirs, または @base を記述 |
Git はかつて、後回しにされがちな存在でした。2023 年、Aider が AI の変更ごとに 1 つの原子的なコミットを行うことでその状況を変えました。以来、本格的なツールはすべてこの流れに従っています。しかし 2026 年の現状では様相が異なります。Copilot CLI がプラットフォームを掌握し、レビュー、PR(プルリクエスト)、イシューがネイティブなオブジェクトとして扱われるようになりました。
Omp(https://omp.sh/)は、Git を単なるツールの一つとして扱います。omp commit コマンドでは、関連しない変更点を順序立てて複数のコミットに分割し、依存関係の循環を記述前に検出して拒否します。また、マージ操作も...
解決されたコンフリクトは、従来の方式ではなく、@ours / @theirs / @base を用いてファイルに反映されます。
テキストの改変は脆いものですが、組み込みレビューが今や基準となっています。比較ツールの 5 つのうち 4 つがこの機能を搭載しており、Continue も CI のアイデンティティを同じ考えを中心に構築しています。
セーフティと信頼性
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| 細粒度の権限管理 | はい、モードと許可リストあり | はい、承認およびサンドボックスポリシーあり | はい、ツールごとの承認あり | はい | はい、破壊的処理ではプレビュー後に承認が必要 |
| OSレベルのサンドボックス | はい | はい、Seatbelt/Landlock 対応 | はい、ローカルおよびクラウドサンドボックスあり | いいえ、権限のみ | ファイルシステムクローンによるワークスペース分離 |
| オープンソースのハーンネス | いいえ | はい、Apache-2.0 ライセンス | いいえ | はい、MIT ライセンス | はい、MIT ライセンス |
| ローカルまたはセルフホストモデル | いいえ | はい、--oss オプションあり | いいえ | はい | はい:Ollama, LM Studio, llama.cpp, vLLM |
明確な選択肢はありません。Codex CLI は、オープンソースでありながら OS サンドボックス化され、ローカルモデルも利用可能な唯一のラボエージェントです。Claude Code と Copilot はサンドボックス化に優れていますが、クローズドでクラウド依存となっています。OpenCode と Omp は完全にオープンでローカルモデルを実行できますが、カーネル分離ではなく権限管理に頼っています。CodeBuddy は実行をサンドボックス化し、OpenHands はすべてをコンテナ化します。Amp Free はコードの学習を行わないと約束し、OpenCode はサーバーサイドでコードを保存しないと明言しています。
企業と個人開発者が重視する点は異なります——これがプラットフォーム CLI が存在する理由の半分です。
自動化とインターフェース
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| ヘッドレス / CI | はい、-p | はい、exec | はい、-p | はい、サーバーモード | はい、-p および RPC |
| IDE 統合 | VS Code、JetBrains 拡張機能 | VS Code 拡張機能 | Copilot エコシステム | IDE 拡張機能 | ACP (Zed およびその他の ACP エディタ) |
| ターミナルおよび IDE 以外の表面 | デスクトップ、Web、モバイル | ChatGPT Web およびモバイル | github.com、モバイルアプリ | デスクトップアプリ | なし |
デスクトップ、Web、モバイル、スマホからラップトップへの引き継ぎなど、さまざまな画面で勝敗が決まるのはラボやプラットフォームです。一方、オープンツールはプロトコルに賭けます。ACP(Junie、Kimi Code、Mistral Vibe、CodeBuddy、Omp)を使えば、1 つのエージェントが互換性のあるあらゆるエディターを統括できます。
gptme は常駐型のバックグラウンドエージェントを実行します。Continue の cn は、それがトレンドになる前から CI 最優先で設計されていました。Omp の /collab コマンドは、クライアントサイド暗号化を施したライブセッションをリンク経由で共有できます。一見奇妙な機能ですが、実務では非常に役立ちます。
Web、メディア、入力
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Web 検索およびフェッチ | はい、組み込み | はい、サーバー承認済みインデックスモードを含む | MCP を経由 | フェッチツール | 18 の検索プロバイダーをチェーン化し、サイト認識型抽出を実行 |
| ブラウザ制御 | MCP を経由 | はい、組み込みのブラウザ使用(2026 年 4 月) | MCP を経由 | MCP を経由 | 組み込みのヘッドレス Chromium;同じ API が Electron アプリを駆動 |
| 画像入力 | はい | はい | 文書化なし | はい、ドラッグ&ドロップ | はい、inspect_image によるビジョン分析 |
| リッチドキュメントの読み込み | 画像、PDF、ノートブック | 画像 | 文書化なし | 画像 | ファイル、ディレクトリ、アーカイブ、SQLite、PDF、ノートブック、URL を 1 つの read ツールで処理 |
| 音声およびメディア生成 | いいえ | はい、リアルタイム音声制御 | いいえ | いいえ | プロバイダーモデルによる画像生成および TTS |
ドキュメントの取得、Issue の確認、ブラウザ操作ができないエージェントは、作業を人間に引き渡す必要があります。Codex は 4 月にブラウザ制御を主要なツールとして正式に採用しました。Omp は 18 種類の検索プロバイダーを連携させ、ローカルの Electron アプリを駆動できるヘッドレス Chromium を実行します。PDF、SQLite のテストデータ、ノートブックといったファイルはチェックリスト上では些細なことですが、仕様ページをコピーしてプロンプトに貼り付けるしかない場合の代償は甚大です。Aider はすでに数年前から音声入力をサポートしていますが、Codex も現在は音声対応となりました。一方、Open Interpreter は「マシン全体を制御する」ことを意味し続けています。
コストエンジニアリング
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| プロンプトキャッシュ戦略 | 自動(Anthropic API) | 自動 | プラン管理 | プロバイダ依存 | プロバイダ依存;トークン予算はリアルタイムでカウント |
| 低コストモデルルーティング | サブエージェントごとのモデル選択 | スレッドごとの予算とプロファイル | 部分的、エージェントごとのモデル | エージェントごとのモデル選択 | ロールベース:デフォルト/smol/slow/plan/commit、フォールバックチェーンとパスごとの固定あり |
| 使用状況の可視性 | あり | あり、予算追跡ビュー付き | プランメーター | あり、セッションごと | あり、ライブトークンカウント |
同じリファクタリングの指示を出して、忙殺される一日。一方は Claude Code のクレジット残高、もう一方は 99.8% のキャッシュヒット率と 12 ドルの請求書。この差こそが、「コストエンジニアリング」を機能として組み込む理由となった。
多くのハルネスではプロバイダー側のキャッシングを「運」と捉えるが、Reasonix はそれを設計の一部としている。DeepSeek のプレフィックスキャッシュが確実にヒットするようバイト単位で安定したコンテキストを提供し、繁忙日でもレポートによると 99.82% のヒット率を実現している。Claude のトークン数では採算が取れないような自動化処理も、この仕組みなら成り立つ。Omp は役割に応じてサブタスクを安価なモデルに振り分け、認証情報をローテーションする。サブスクリプション型のエージェントは請求書を隠蔽していたが、ヘッドレス利用分には専用のメーターが設けられるようになった。Claude Code も 6 月にこの機能を追加した。
Feature leaders
| 分野 | リーダー格 | 強力な候補 |
|---|---|---|
| 編集精度 | Omp (ハッシュアンカー、AST) | Aider (diff 形式) |
| Git ワークフロー | Omp (コミット分割、競合ツール)、Aider (アトミックコミット) | Copilot CLI (PR ネイティブ) |
| コードインテリジェンス | Omp、Auggie、Junie | OpenCode、Aider |
| デバッグ | Omp (DAP) | Junie (エージェント型デバッグ) |
| オーケストレーション | Claude Code (チーム)、Omp (IRC + アドバイザー) | Grok Build、Codebuff、Droid |
| クラウドへのハンドオフ | Copilot CLI、Codex、Claude Code | Mistral Vibe (リモートエージェント) |
| メモリ | Copilot CLI、Claude Code、Omp | Kilo (Memory Bank)、gptme |
| 拡張性 | Claude Code (スキル/フック)、Pi/Omp (TS 拡張) | Goose (MCP ネイティブ)、OpenCode |
| サンドボックス化 | Codex、Copilot CLI | CodeBuddy、OpenHands (コンテナ) |
| コストエンジニアリング | DeepSeek-Reasonix (キャッシュファースト) | Omp (ロールルーティング)、Amp Free |
| プロバイダーの網羅性 | OpenCode (75+)、Kilo (500+ モデル) | ForgeCode (300+)、Omp (40+ 種 plus コーディングプランログイン) |
| コラボレーション | Omp (/collab)、Amp (スレッド) | OpenCode (共有リンク) |
| Web、ブラウザ、メディア | Omp (Chromium、18 社検索、画像生成、TTS) | Codex (ブラウザ操作、音声)、gptme |
Omp は比較対象の中で最も機能豊富です。ただし、チーム規模が小さく(バイス・ファクターが低い)、複雑なタスクでは共同学習されたモデルの方が依然として優れているため、会社の命運をこれに賭けるのは危険です。しかし Claude Code が新機能をリリースした際は、まず Omp を確認しましょう。コミュニティによるフォーク版の方が先に実装されているケースが多いためです。
料金、ベンチマーク、信頼性
2026 年前半の料金体系は大きく二つに分裂しました。ラボ系エージェントは推論コストをサブスクリプションにバンドルしており、通常は最安値で最先端モデルへのアクセスを提供しますが、その代償としてベンダーロックインが発生します。一方、BYOK(Bring Your Own Key)方式は利便性を犠牲にして自由度を得る代わりに、オープンウェイトモデルの価格差を利用した裁定取引が可能になります。
原文を表示
The terminal was an unlikely winner. In 2024 the bet was on IDEs — Copilot already lived in the editor, Cursor was climbing, and "agent" still meant chat with shell access bolted on. By mid-2026 the serious usage ran from CI, SSH, and machines with no GUI at all. Scriptable, boring, and the model never fights something else for the window manager.
35 actively maintained CLI coding agents as of July 3, 2026. Crowded field.
Skip to feature by feature comparison · Skip to conclusions
Short history
The 1st wave arrived in 2023, before anyone called these things agents. gptme (March 2023) let a model run shell commands from the terminal. Aider (mid-2023) built AI pair programming around git, with atomic commits as the unit of change. Open Interpreter (July 2023) generalized the idea to controlling the whole computer. All 3 survive — gptme as a daemon, Aider as a pair programmer, Open Interpreter as a general computer controller.
Anthropic's Claude Code research preview in February 2025 set the shape: agentic loop, file and shell tools, project memory file, permission prompts, plan mode, hooks, subagents. Everyone else cloned it. OpenAI shipped Codex CLI in April 2025 (later rewritten in Rust). Google followed with Gemini CLI in June 2025 and an aggressive free tier. By the 2nd half of 2025 the releases piled up in one quarter — Cursor, Amp, Augment, Factory, Charm, and 12 open-source teams.
Teams that standardized on Gemini CLI in 2025 still have configs from last fall. December 2025 is when the standards people showed up: the Linux Foundation formed the Agentic AI Foundation (AAIF), anchored by Anthropic's Model Context Protocol, OpenAI's AGENTS.md convention, and Block's Goose agent, with backing from Google, Microsoft, AWS, Cloudflare, and Bloomberg. The same month, Sourcegraph spun its agent out as the independent Amp Inc., Mistral entered with Vibe and the Devstral 2 models, and Codebuff open-sourced its multi-agent harness.
The 1st half of 2026 was messier. Copilot CLI hit GA in February. Cline shipped CLI 2.0 and peeled off an open SDK. Kilo CLI reached 1.0. Junie went from beta to GA in June. Google stole the headline at I/O in May: Antigravity CLI launched, and Gemini CLI got a kill date — June 18 for free and consumer-plan users. Grok Build landed late May. Moonshot replaced its earlier CLI with Kimi Code in June.
Models kept pace. DeepSeek V4 under MIT in April. GLM-5.2 and Kimi K2.7 Code in June, pulling open weights to within a few points of frontier terminal scores. The free tiers started dying around the same time: Qwen Code's hosted tier in April, Amp Free waitlisted, Gemini's consumer path gone by summer.
Model labs
Lab agents come from the model vendors — harness and model in one box, usually on a subscription line item already there.
| Agent | Maintainer | Debut | Source | Models | Access | Known for |
|---|---|---|---|---|---|---|
| Claude Code | Anthropic | Feb 2025 | Closed | Claude only | Pro/Max plans or API | The category template; deepest orchestration features |
| Codex CLI | OpenAI | Apr 2025 | Apache-2.0, 95k stars | OpenAI Codex models | ChatGPT plans or API | Rust core, OS-level sandboxing, cloud handoff |
| Gemini CLI | Jun 2025 | Apache-2.0, 106k stars | Gemini | API key / enterprise only after Jun 18, 2026 | Consumer access retired in favor of Antigravity CLI | |
| Antigravity CLI | May 2026 | Closed | Gemini 3.5 plus selected third-party models | Free public preview | Shares one harness with the Antigravity 2.0 desktop app | |
| Grok Build | xAI | May 2026 | Closed | grok-build (256K context) | SuperGrok 30/mo,XPremium+30/mo, X Premium+ 40/mo, or API | Up to 8 parallel subagents in isolated git worktrees |
| Mistral Vibe | Mistral AI | Dec 2025 | Apache-2.0, 4.6k stars | Mistral (Medium 3.5, Devstral 2); open weights | Free CLI; Mistral plans/API | European option; remote async agents; ACP |
| Kimi Code CLI | Moonshot AI | Jun 2026 | MIT, 3k stars | Kimi K2.7 Code | Kimi plans or low-cost API | Built-in coder/explore/plan subagents; ACP |
| Qwen Code | Alibaba (Qwen) | Jul 2025 | Apache-2.0, 26k stars | Qwen3-Coder or any endpoint | BYOK; paid plans (free hosted tier ended Apr 2026) | Gemini CLI fork tuned for Qwen's open-weight coders |
Claude Code still sets the vocabulary — agent teams, hooks, skills, the lot. Experimental agent teams (multiple sessions messaging each other) remain ahead of the pack. The trade is total model lock-in. As of June 2026, headless and SDK usage bill from a separate credit pool. Terminal-Bench 2.0 in early July had Anthropic's newest models on top. The bundled pitch is a co-trained model plus harness, and the leaderboard backs it.
Codex CLI is the strongest counterweight, and it had a busy spring. It's open source, OS-sandboxed, rides ChatGPT's subscriber base, and runs GPT-5.5 as its current frontier default. Through the 1st half of 2026 it picked up persistent Goals with token budgets, thread-level delegation to subagents, a plugin marketplace, browser use, encrypted remote execution, and a one-liner to import Claude Code configuration — switching costs are real enough that Codex had to automate them.
If you standardized on Gemini CLI in 2025, June 18 was a bad day. Google grew it into one of the most-starred repos in the category, then pointed free and consumer-plan users at the closed-source Antigravity CLI instead. Antigravity is a Go rewrite that shares its harness with the Antigravity desktop platform (Gemini 3.5 at the core), runs async multi-agent workflows, and is free in public preview, with non-Google models on the menu. Free tiers are not foundations.
Grok Build, Kimi Code, and Mistral Vibe showed up late — plan mode, subagents, headless CI already checked off. Grok Build shipped in beta with plan mode, worktree-isolated parallel subagents, and headless CI support on day 1. Kimi Code and Mistral Vibe both lean on unusually cheap, partly open models: Mistral's Devstral 2 posts a vendor-reported 72.2 percent on SWE-bench Verified from a 123B model, and Moonshot's K2.7 Code, released mid-June with open weights, undercuts frontier pricing by an order of magnitude. Both adopted ACP so any compatible editor can host them. Qwen Code is the fork case: a Gemini CLI derivative retuned for Qwen's open-weight coders, still alive after its free tier died because the tool is Apache-licensed and endpoint-agnostic.
Platform and product CLIs
Platform vendors sell the agent as part of a development stack — distribution, integration, governance. Almost all of them are multi-model.
| Agent | Maintainer | Debut | Source | Models | Access | Known for |
|---|---|---|---|---|---|---|
| GitHub Copilot CLI | GitHub | Sep 2025, GA Feb 2026 | Closed | Anthropic, OpenAI, Google, open-weight Kimi K2.7 | Paid Copilot plans (from $10/mo) | Auto-delegating specialist agents; & hands work to a cloud agent |
| Cursor CLI | Anysphere | Aug 2025 | Closed | Cursor Composer + frontier models | Cursor plans | Same agent, rules, and MCP config as the Cursor IDE |
| Amp | Amp Inc. (ex-Sourcegraph) | 2025, spun out Dec 2025 | Closed | Curated frontier mix, no model picker | Pay-as-you-go; ad-funded free tier (waitlisted) | Deliberately opinionated; the ad-supported experiment |
| Auggie CLI | Augment Code | 2025 | Closed | Frontier models via Augment | Augment plans | Context engine that indexes the whole repo before work starts |
| Droid | Factory AI | 2025 | Closed | Multi-model, BYOK supported | Factory plans | Topped the original Terminal-Bench; spans CLI, IDE, web, Slack, ticketing |
| Junie CLI | JetBrains | Beta Mar 2026, GA Jun 2026 | Closed | OpenAI, Anthropic, Google, xAI | JetBrains AI plans | ACP-native; agentic debugging; IDE and database integration |
| Qoder CLI | Alibaba | Platform Aug 2025; CLI 2026 | Closed | Alibaba Model Studio | Pay-as-you-go / coding plans | Quest mode for spec-driven autonomous tasks |
| CodeBuddy Code | Tencent | Sep 2025 | Closed | DeepSeek + Hunyuan | CodeBuddy plans | Skills, plan mode, ACP, sandboxed execution; 12k internal Tencent users |
Copilot CLI wins on reach. GA in February 2026, $10/mo entry tier, GitHub's MCP server built in, plugins from repos, per-repo memory, specialist agents for explore/plan/review/build. Prefixing a prompt with & pushes the job to GitHub's cloud coding agent. For GitHub-centric teams it's the lazy default, and its model picker makes it one of the most vendor-neutral closed agents. On July 1 Copilot added Moonshot's open-weight Kimi K2.7 Code on Azure — the first open-weight model inside a major closed platform agent.
After Copilot, the bets split.
Cursor CLI is the consistency play — same agent, same rules, same MCP config in the IDE, the terminal, and CI. One AGENTS.md driving all three is the whole product.
Amp drops the model picker entirely and swaps models when its team decides the mix should change. The ad-funded Amp Free tier (sponsor-backed, no-training guarantee) is the strangest business model in the list; admission is closed for now.
Auggie front-loads context: index the whole repo before the first prompt. Pays off on legacy monoliths, less on greenfield repos.
Droid is the enterprise bundle — specialist agents, heavy parallelism, Slack and ticketing hooks. Factory posted the best Terminal-Bench numbers of any product harness in late 2025.
Junie drags JetBrains' debugger and database wiring to the terminal over ACP.
Qoder and CodeBuddy are the China stack plays — same 2026 feature checklist as everyone else, different distribution behind the firewall.
Open-source harnesses
Bring-your-own-key agents are the largest group, and open-weight pricing changed what they cost to run. A decent harness plus GLM, DeepSeek, Qwen3-Coder, Devstral, or Kimi tokens gets you frontier capability for a fraction of the subscription price. The gap closed sharply in June, when Z.ai released GLM-5.2 under MIT: a 744B mixture-of-experts model (about 40B active) with a 1 million-token context that beats GPT-5.5 on several long-horizon coding benchmarks at roughly 1/6 of the price (Terminal-Bench 2.1 analysis). It also runs fully offline on high-memory hardware, needing around 245GB of memory at the most aggressive usable quantization.
| Agent | Maintainer | Debut | Source | Models | Known for |
|---|---|---|---|---|---|
| OpenCode | Anomaly (SST team) | 2025 | MIT, 182k stars | 75+ providers | Most-starred agent on GitHub; TUI, desktop, IDE; agents and skills |
| Crush | Charmbracelet | Jul 2025 | FSL-1.1-MIT (source-available), 26k stars | Multi-provider | The best-crafted TUI in the category; LSP context; MCP |
| Goose | AAIF (ex-Block) | Jan 2025 | Apache-2.0, 51k stars | Any, incl. local | MCP-native, local-first, foundation-governed |
| Aider | Aider-AI | Mid-2023 | Apache-2.0, 47k stars | 100+ via LiteLLM | Git-native pair programming; repo map; pace has slowed |
| Cline CLI | Cline | Late 2025; 2.0 Feb 2026 | Apache-2.0, 64k stars | Any provider | Open @cline/sdk runtime; parallel agents; headless CI |
| Kilo CLI | Kilo | 1.0 Feb 2026 | Open source, 26k stars | 500+ models | Architect/Code/Debug/Ask/Orchestrator modes; Memory Bank |
| Continue CLI (cn) | Continue | Sep 2025 | Apache-2.0, 35k stars | Any provider | Headless PR checks and CI agents; pace has slowed |
| OpenHands CLI | OpenHands | CLI May 2026; project 2024 | Open source, 79k stars | Any provider | Agent SDK with event-sourced replay; CLI as stable thin client |
| DeepSeek-Reasonix | esengine | Apr 2026 | MIT, 26k stars | DeepSeek V4 (default) or any endpoint | Cache-first design; extreme budget efficiency |
| Every Code | JustEvery | Aug 2025 | Apache-2.0, 3.8k stars | OpenAI, Anthropic, Google, any | Codex fork orchestrating several vendors' models at once |
| ForgeCode | Antinomy | Dec 2024 | Apache-2.0, 7.4k stars | 300+ models | Rust, shell-native, semantic codebase search |
| Codebuff | Codebuff AI | 2024; open-sourced late 2025 | Apache-2.0, 7.1k stars | Any via OpenRouter | Explicit agent roles: finder, planner, editor, reviewer |
| Kode | shareAI-lab | Jul 2025 | Apache-2.0, 5.2k stars | Multi-model collaboration | @ask a specific model mid-task |
| Nanocoder | Nano Collective | Jul 2025 | Open source, 2.2k stars | Local-first + any endpoint | Community-owned, no telemetry, no company behind it |
OpenCode has the most GitHub stars in the category at 182k — ahead of Gemini CLI and Codex — with 75+ providers, a client-server architecture, and a polished TUI plus desktop and IDE surfaces. It also has the messiest fork story: the original author joined Charm in mid-2025, where that lineage became Crush (prettier TUI, source-available FSL license), while the SST team's OpenCode kept growing. Two healthy projects out of one messy split.
DeepSeek-Reasonix is the fastest riser: 26k stars in roughly 10 weeks from one static Go binary. The design treats inference cost as an engineering problem. Sessions keep DeepSeek's byte-stable prefix cache hot — stable env summaries at startup, stale tool output pruned before compaction, planner and executor on separate cache-stable threads.
One reported day: 435M input tokens, 99.82% cache hits, about 12.Uncachedestimate:12. Uncached estimate: 61. Any OpenAI-compatible endpoint, MCP over stdio and HTTP, checkpoints, rewind, and chat-ops bridges to Feishu, Lark, and WeChat. DeepSeek V4 landing the same April under MIT did not hurt.
Goose took a different route to durability: Block donated it to the Agentic AI Foundation, making it the only major agent under neutral foundation governance, alongside the MCP and AGENTS.md standards themselves. Cline matters beyond its user base because of its Apache-licensed SDK, extracted in May 2026, which turned a popular product into infrastructure others can build on; GitHub (Copilot SDK), Anthropic (Agent SDK), and OpenHands (a research-grade Python SDK with deterministic replay) made the same move. Kilo and Continue extend IDE-extension franchises into the terminal — Kilo with orchestration modes and 500+ models, Continue with a CI-first angle — though Continue's 2026 commit cadence puts it on the watch list.
Aider invented half the git discipline everyone else copied, and its polyglot benchmark shaped model evaluation for 2 years. It stayed a pair-programming tool while the market went agentic, though, and releases have slowed to a trickle. Still strong at atomic commits; the category moved elsewhere.
The long tail is larger than the headline tools suggest. Every Code out-features its own upstream, orchestrating OpenAI, Anthropic, and Google models on a Codex base. Codebuff decomposes work into explicit finder, planner, editor, and reviewer agents, and claims 61 percent versus 53 for Claude Code on its own 175-task evaluation. Kode lets a session consult a specific model by name mid-task. Nanocoder is for people who want collectively owned, telemetry-free, local-first — no company behind it.
Minimal cores and pioneers
| Agent | Maintainer | Debut | Source | Known for |
|---|---|---|---|---|
| Pi | Earendil (Mario Zechner) | Aug 2025 | MIT, 67k stars | 4 tools, sub-1,000-token system prompt, TypeScript extension SDK |
| oh-my-pi | can1357 | Dec 2025 | MIT, 16k stars | Maximalist Pi distribution; the largest feature surface in the category |
| mini-SWE-agent | SWE-agent team | Jun 2025 | MIT, 5.6k stars | About 100 lines of Python, bash-only; the research control group |
| gptme | gptme project | Mar 2023 | MIT, 4.4k stars | Oldest tool here; persistent autonomous agents with git-backed memory |
| Open Interpreter | Open Interpreter | Jul 2023 | Open source, 64k stars | Pioneered LLM code execution; broader computer control; quiet in 2026 |
Pi is the backlash. Mario Zechner argues a modern model needs 4 tools (read, write, edit, bash), a system prompt under 1,000 tokens, and nothing else in core — MCP included belongs in a TypeScript extension SDK. Personal project to 67k stars in under a year, which says the audience for harness bloat had thinned.
oh-my-pi — Omp for short — is the opposite bet: a batteries-included Pi distribution that kept adding features until it had the widest feature set in the comparison below. Same community, opposite directions. Pi bets on subtraction, Omp bets on accumulation, and both arguments still have an audience.
mini-SWE-agent is the control group: about 100 lines of Python, bash only, and strong frontier models still clear a large share of SWE-bench. Hard to justify a 50-feature harness without beating that. gptme and Open Interpreter round out the elders — gptme still in active development and focused on persistent, self-improving agents; Open Interpreter historically important and visibly slowing.
Feature by feature
By 2025 every serious agent had the same skeleton: edit-run-test loop, project instructions file, MCP, plan mode, permission prompts, headless mode. None of that separates anyone anymore.
What marketing pages still hide is the rest. Five harnesses get the full treatment — Claude Code, Codex CLI, OpenCode, Omp, plus Copilot CLI as the deployed-at-scale platform pick. Everyone else appears in the notes under each table.
Context and memory
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Project instructions | CLAUDE.md | AGENTS.md | AGENTS.md + instructions | AGENTS.md | AGENTS.md, plus reads .claude, .cursor, .codex, .gemini, .cline configs directly |
| Learned memory across sessions | Yes, persistent memory files | Partial: persistent Goals | Yes, Copilot Memory per repo | No, manual AGENTS.md | Yes, Hindsight: retain/recall/reflect over a project SQLite bank |
| Auto-compaction | Yes | Yes | Yes, automatic at 95% plus /compact | Yes | Yes, plus bitmap-frame history compression |
| Checkpoints and rewind | Yes, /rewind | No | No | Undo/redo | Yes, checkpoint/rewind with context pruning |
A year ago project memory meant a markdown file someone on the team updated by hand. Now 3 of the 5 leaders learn on their own: Claude Code accumulates memory files across sessions, Copilot builds a per-repository understanding, and Omp writes facts mid-run and synthesizes them later on request. Omp also reads the rules, skills, and MCP registrations other agents leave in a repo. Switching over barely touches config. Elsewhere, Kilo's Memory Bank stores agent state in structured markdown inside the repo, and gptme goes furthest of all with git-backed memory for agents that are meant to run for months.
Editing and code intelligence
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Anchored or AST-aware edits | No | No | No | No | Yes: hash-anchored patches (~61% fewer output tokens, project-reported) and ast-grep rewrites across 50+ grammars |
| LSP integration | Yes, diagnostics and navigation | No | No | Yes, built-in | Yes, deep: diagnostics, references, code actions, renames that propagate through re-exports |
| Debugger control | No | No | No | No | Yes, DAP: lldb, dlv, debugpy; breakpoints, stepping, variable inspection |
| Beyond-grep code search | Agentic grep | Agentic grep | Grep + GitHub code search | Grep + LSP symbols | In-process ripgrep, tree-sitter structural summaries, fuzzy match |
| Formatter and linter awareness | Via hooks | No | Not documented | Yes, built-in formatter support | Yes, LSP diagnostics and linters feed decisions |
Omp leads here by a lot. Hash-anchored editing is aimed at the refactor failure everyone hits — the patch lands on the wrong line because whitespace moved. AST rewrites with preview-then-accept, plus a debugger the agent can drive. No lab agent ships that as of July 2026. Heavier setup than Claude Code; you still pick the model. This is where to look when patches keep landing on the wrong line. The rest of these tools approach code intelligence differently: Auggie's context engine builds a semantic index before work begins, Aider's repo map compresses project structure into the prompt, ForgeCode does embedding-based search after a sync step, and Junie leans on the JetBrains indexer — probably the deepest static analysis any agent gets, just not an open one.
Orchestration
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Subagents | Yes, custom agent definitions | Yes, thread-level delegation (Jun 2026) | Yes, auto-delegating specialists | Yes, custom agents | Yes, fan-out with schema-validated JSON returns |
| Parallel work and isolation | Yes, worktrees; agent teams (experimental) | Cloud tasks; per-thread token budgets | Via cloud agent | Parallel sessions | Yes, filesystem-clone isolation (APFS/btrfs/overlayfs) |
| Agents talking to each other | Yes, teams messaging (experimental) | No | No | No | Yes, IRC-style channel between live agents |
| Second-model oversight | No | No | No | No | Yes, Advisor: a separate model reviews every turn and injects notes |
| Background processes | Yes | Partial | Via cloud | No | Yes, job control |
| Cloud handoff | Yes, web and remote sessions | Yes, Codex cloud; encrypted remote execution | Yes, & prefix | Self-hosted server mode | No; /collab shares a live session instead |
| Scheduled and recurring runs | Yes, scheduled cloud agents | Timed reminders (Jun 2026) | Via GitHub Actions | Via CI or server mode | No |
Claude Code and Omp let agents coordinate as peers — agent teams on one side, an in-process chat channel plus a standing Advisor model on the other. Codex and Copilot push parallelism to the cloud — less load on the laptop, more on the ops queue. Grok Build isn't in the table but shipped worktree-isolated parallel subagents on day 1 (up to 8). Codebuff hardcodes roles. Mistral Vibe keeps remote agents running after the terminal closes. Amp's oracle — a stronger model for hard steps — is the nearest cousin to Omp's Advisor. Scheduling showed up late: recurring cloud agents on Claude Code, timed reminders on Codex in June, CI everywhere else.
Extensibility
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| MCP client | Yes | Yes | Yes, GitHub MCP built in | Yes | Yes |
| Plugin or extension API | Yes, plugins and marketplaces | Yes, plugin marketplace (2026) | Yes, /plugin install from repos | Yes, JS/TS plugins | Yes, TypeScript extensions with hot reload |
| Skills | Yes, originated the format | No | No | Yes | Yes, with input/output schemas for chaining |
| Lifecycle hooks | Yes, originated the format | Limited | No | Via plugin events | Yes, plus mid-stream rules that can abort and correct generation live |
| Custom slash commands | Yes | Yes, custom prompts | Not documented | Yes | Yes, extensions register commands and hotkeys |
| Acts as a server/SDK for other apps | Yes, Agent SDK | Yes, MCP server mode and SDK | Yes, Copilot SDK (GA Jun 2026) | Yes, server + SDK | Yes, Node SDK and NDJSON RPC mode |
MCP for external tools, plugins for behavior, skills for instructions. That stack is the default now. Claude Code invented two of the three formats. Omp's stream rules are the oddball: a regex on the token stream aborts mid-sentence, injects a correction, resumes. Project rules without stuffing every prompt. Goose is MCP-purist. Pi ships no MCP in core and hands everything to a TypeScript extension SDK; 67k stars say that resonated.
Git and review workflows
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Commit assistance | Yes, git-aware commits and PRs | Yes; cloud tasks commit and open PRs | Yes, PR-native via the cloud agent | Git-backed undo/redo; commits on request | omp commit splits unrelated changes into dependency-ordered commits |
| Built-in code review | Yes, /code-review with effort levels | Yes, /review | Yes, dedicated review agent | Via custom agents | Yes, /review: parallel reviewers, P0-P3 ranking, ship verdict |
| PR and issue integration | Via gh and GitHub Actions | GitHub action and Codex cloud | Native to the platform | GitHub and GitLab integrations | PRs and issues addressable as pr:// and issue:// paths; watches Actions runs live |
| Merge-conflict tooling | Agentic only | Agentic only | Agentic only | Agentic only | Declarative: write @ours, @theirs, or @base to conflict://N |
Git used to be an afterthought. Aider fixed that in 2023 with one atomic commit per AI change; every serious tool followed. The 2026 wrinkles look different. Copilot CLI owns the platform — review, PRs, issues are native objects. Omp treats git as another tool surface: omp commit splits unrelated changes into ordered commits, rejects dependency cycles before writing, and turns merge conflicts into files resolved with @ours / @theirs / @base instead of fragile text surgery. Built-in review is baseline now: 4 of the 5 comparison tools ship it; Continue built its CI identity around the same idea.
Safety and trust
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Granular permissions | Yes, modes and allowlists | Yes, approval and sandbox policies | Yes, per-tool approval | Yes | Yes, preview-then-accept on destructive operations |
| OS-level sandbox | Yes | Yes, Seatbelt/Landlock | Yes, local and cloud sandboxes | No, permissions only | Filesystem-clone workspace isolation |
| Open-source harness | No | Yes, Apache-2.0 | No | Yes, MIT | Yes, MIT |
| Local or self-hosted models | No | Yes, --oss | No | Yes | Yes: Ollama, LM Studio, llama.cpp, vLLM |
There is no clean option. Codex CLI is the only lab agent that's open source, OS-sandboxed, and local-model-capable at once. Claude Code and Copilot sandbox well but stay closed and cloud-bound. OpenCode and Omp are fully open and run local models, but permissions instead of kernel isolation. CodeBuddy sandboxes execution, OpenHands containerizes everything, Amp Free promises no training on your code, OpenCode says it stores no code server-side. Enterprises weight different rows than individual developers — which is half the reason platform CLIs exist.
Automation and surfaces
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Headless / CI | Yes, -p | Yes, exec | Yes, -p | Yes, server mode | Yes, -p and RPC |
| IDE integration | VS Code, JetBrains extensions | VS Code extension | Copilot ecosystem | IDE extensions | ACP (Zed and other ACP editors) |
| Surfaces beyond terminal and IDE | Desktop, web, mobile | ChatGPT web and mobile | github.com, mobile app | Desktop app | No |
Labs and platforms win on surfaces — desktop, web, mobile, handoff from phone to laptop. Open tools bet on protocols instead. ACP (Junie, Kimi Code, Mistral Vibe, CodeBuddy, Omp) lets one agent serve any compatible editor. gptme runs persistent background agents. Continue's cn was CI-first before that was trendy. Omp's /collab shares a live session over a link with client-side encryption. Strange feature, useful in practice.
Web, media, and input
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Web search and fetch | Yes, built in | Yes, incl. a server-approved indexed mode | Via MCP | Fetch tool | Chains 18 search providers with site-aware extraction |
| Browser control | Via MCP | Yes, built-in browser use (Apr 2026) | Via MCP | Via MCP | Built-in headless Chromium; the same API drives Electron apps |
| Image input | Yes | Yes | Not documented | Yes, drag-and-drop | Yes, vision analysis via inspect_image |
| Rich-document reading | Images, PDFs, notebooks | Images | Not documented | Images | Files, directories, archives, SQLite, PDFs, notebooks, and URLs through one read tool |
| Voice and media generation | No | Yes, realtime speech controls | No | No | Image generation and TTS via provider models |
An agent that can't fetch docs, check an issue, or poke a browser hands the work back to a human. Codex made browser control a first-class tool in April. Omp chains 18 search providers and runs headless Chromium that can drive a local Electron app. PDFs, SQLite fixtures, notebooks — minor on a checklist, painful when the alternative is copy-pasting spec pages into the prompt. Aider had voice input years ago; Codex takes speech now; Open Interpreter still means "control the whole machine."
Cost engineering
| Claude Code | Codex CLI | Copilot CLI | OpenCode | Omp | |
|---|---|---|---|---|---|
| Prompt-cache strategy | Automatic (Anthropic API) | Automatic | Managed by plan | Provider-dependent | Provider-dependent; token budgets counted in real time |
| Cheap-model routing | Per-subagent model choice | Per-thread budgets and profiles | Partial, per-agent models | Per-agent model choice | Role-based: default/smol/slow/plan/commit, with fallback chains and per-path pinning |
| Usage visibility | Yes | Yes, budget tracking views | Plan meter | Yes, per session | Yes, live token counting |
Same refactor prompt, heavy day, Claude Code vs Reasonix: credit meter on one side, 99.8% cache hit rate and a $12 bill on the other. That gap is why cost engineering became a feature.
Most harnesses treat provider-side caching as luck. Reasonix treats it as design — byte-stable context so DeepSeek's prefix cache keeps hitting, 99.82% on heavy days per their reports. Automations too marginal for Claude tokens start to pencil out. Omp routes sub-tasks to cheap models by role and rotates credentials. Subscription agents hide the bill until headless usage gets its own meter, which Claude Code added in June.
Feature leaders
| Area | Leader | Also strong |
|---|---|---|
| Editing precision | Omp (hash-anchored, AST) | Aider (diff formats) |
| Git workflows | Omp (commit splitting, conflict tooling), Aider (atomic commits) | Copilot CLI (PR-native) |
| Code intelligence | Omp, Auggie, Junie | OpenCode, Aider |
| Debugging | Omp (DAP) | Junie (agentic debugging) |
| Orchestration | Claude Code (teams), Omp (IRC + Advisor) | Grok Build, Codebuff, Droid |
| Cloud handoff | Copilot CLI, Codex, Claude Code | Mistral Vibe (remote agents) |
| Memory | Copilot CLI, Claude Code, Omp | Kilo (Memory Bank), gptme |
| Extensibility | Claude Code (skills/hooks), Pi/Omp (TS extensions) | Goose (MCP-native), OpenCode |
| Sandboxing | Codex, Copilot CLI | CodeBuddy, OpenHands (containers) |
| Cost engineering | DeepSeek-Reasonix (cache-first) | Omp (role routing), Amp Free |
| Provider breadth | OpenCode (75+), Kilo (500+ models) | ForgeCode (300+), Omp (40+ plus coding-plan logins) |
| Collaboration | Omp (/collab), Amp (threads) | OpenCode (share links) |
| Web, browser, and media | Omp (Chromium, 18-provider search, image gen, TTS) | Codex (browser use, voice), gptme |
Omp has the widest feature set in the comparison. Tiny bus factor, and co-trained models still win on gnarly tasks, so don't bet the company on it. But when Claude Code ships something new, check Omp first; the community fork often had it earlier.
Money, benchmarks, and trust
The pricing story split in the 1st half of 2026. Lab agents bundle inference into subscriptions — usually the cheapest frontier access, with lock-in attached. BYOK harnesses trade convenience for freedom and can arbitrage open-weight pricing;
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み