Simon Willison、AI ツール選択の指針を公開しモデル進化を追跡
Simon Willison は、業務遂行に最適な AI モデルを選択するための意見指南記事を公開し、1 年前から現在に至るまでのチャット型 AI や最新モデル(o3、Claude 4 Opus など)の進化を追跡した。
AI深層分析を開く2026年7月28日 09:15
AI深層分析
キーポイント
AI ガイドの焦点変化
Ethan Mollick のガイドは過去 1 年間はチャットモデル中心だったが、現在は「数時間の人間労働を一度に遂行できる」アジェンシーシステムが主流となっている。
Gemini の評価低下
Google が Codex や ChatGPT Work 相当の明確なエントリーを持たないため、Ethan Mollick のリストから Gemini が除外された状態である。
アジェンシーモードの命名混乱
ChatGPT の「Work」と「Codex」、Claude の「Cowork」と「Code」は名称が対応しておらず、ユーザーを混乱させる要因となっている。
モバイルとデスクトップの違い
ChatGPT モバイルの Work モードはコードインタープリタのインターネットアクセス制限が解除される点で、デスクトップ版とは実質的に異なる挙動を示す。
重要な引用
Today it's much more about agentic systems - "where the AI is capable of doing the equivalent of many hours of real human work in one go".
Gemini has fallen off Ethan's list, since Google still doesn't have an established entry in the Codex/ChatGPT Work/Cowork category.
The names do not map onto each other in any way that will help you remember them.
編集コメントを表示
編集コメント
この記事は、AI ツールの進化が「何ができるか」から「どう自律的に動くか」へとパラダイムシフトしている現状を浮き彫りにしている。ユーザーは名称の混乱に惑わされず、環境ごとの機能差を理解してツールを選択する必要があるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI を使って何をするかを決めるための、主観的なガイド
イーサン・モリック氏のガイドが時とともにどう進化してきたかを眺めるのは興味深いものです。
[1 年前] (https://www.oneusefulthing.org/p/using-ai-right-now-a-quick-guide) の時点では、まだチャット中心でした。ChatGPT、Claude、Gemini が主力モデルで、o3、Claude 4 Opus、Gemini 2.5 Pro が挙げられ、Deep Research は有用な代替モードとして紹介されていました。
しかし今日では、話題は「エージェント型システム」へと移っています。「AI が一度の操作で人間が何時間もかかる作業を完遂できるような」領域です。
Google はまだ Codex や ChatGPT Work、Cowork といったカテゴリに確固たる地位を築けていないため、イーサンのリストから Gemini は外れています。Gemini Spark もまだ実力を証明しきれていません。
イーサンは、ChatGPT や Claude にコンピュータを使わせる方法について有用な解説を提供しています。
AI 企業が提供するコンピュータを使用する場合は、ChatGPT では「ChatGPT Work」、Claude では「Cowork」というモードを選択する必要があります(命名がこれほど混乱を招くとは、私としてもお詫び申し上げます)。 [...]
AI を最も強力に活用する方法は、コンピューターへのアクセス権を与えることです。ChatGPT や Claude のアプリをダウンロードし、使用するモードを選択すれば実現できます。
ChatGPT には「Work」と「Codex」の 2 つのエージェントモードがあり、Claude には「Cowork」と「Code」があります。これらの名称は互いに対応しておらず、記憶の助けになるような関係性はありません。また、前述した Work や Cowork モードと同じ名称が使われていますが、動作は異なり、コンピューターにアクセスできるため、より多くの機能と能力を備えています。
私は、モバイル端末上の「ChatGPT Work」と、デスクトップアプリ内の「ChatGPT Work」(これは実質的に Codex の少し親しみやすいインターフェース)との違いについて、あまりにも直感的でないと感じています。
簡単に言うと、ChatGPT モバイルで「Chat」モードから「Work」モードに切り替えると、コード実行コンテナがインターネットへのアクセス制限を解除されたバージョンになります。
原文を表示
An opinionated guide to which AI to use to do stuff
It's interesting watching the evolution of Ethan Mollick's guide over time.
A year ago it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode.
Today it's much more about agentic systems - "where the AI is capable of doing the equivalent of many hours of real human work in one go".
Gemini has fallen off Ethan's list, since Google still doesn’t have an established entry in the Codex/ChatGPT Work/Cowork category. Gemini Spark has yet to prove itself!
Ethan offers a useful explanation of the ways you can give ChatGPT or Claude a computer to use:
To use the computers provided by the AI companies, the mode you want is called ChatGPT Work in ChatGPT, and Cowork in Claude (the naming will not get less confusing, I am sorry to say). [...]
The most powerful way to use AI is to give it access to your computer. You do that by downloading the ChatGPT or Claude apps and picking a mode to use. ChatGPT's two agent modes are Work and Codex; Claude's are Cowork and Code. The names do not map onto each other in any way that will help you remember them. And yes, these use the same names as the Work and Cowork modes we discussed above, but operate differently, and have more features and capabilities because they can access your computer.
I think the difference between ChatGPT Work on a mobile device and ChatGPT Work inside the desktop app (where it's effectively a less intimidating skin on top of Codex) is spectacularly unintuitive.
Short version: if you flip ChatGPT mobile from "Chat" to "Work" mode you get a version where its Code Interpreter container is no longer restricted from accessing the internet!
Tags: ai, generative-ai, llms, ethan-mollick, code-interpreter, general-agents
AI算出
論評・提言ainew評価標準
記事は AI モデルの進化(o3, Claude 4 Opus など)と「エージェント型」へのシフトを論じており、AI が主題であるため関連性は最高値です。新規性については、既存のガイドの更新や特定の視点からの分析が含まれるが、世界初の重大発表ではないため中程度と判断しました。検索機会では具体的なモデル名(Claude 4, o3 など)が含まれているため高評価とし、日本固有の情報や企業事例がないため関連性は低めです。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 50
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み