チャットボットの黄昏時
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
One Useful Thing
米国の主要 AI ラボが従来以上に迅速に高性能モデルをリリースしているが、政府の介入により Claude Fable や GPT-5.6 といった強力なモデルへのアクセスは停止されている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI の進展が加速していると感じるなら、おそらくそれは正しい感覚です。米国の主要な AI ラボからより優れた AI モデルがこれまで以上に迅速にリリースされています(ただし、政府の介入により、最も強力な 2 つのモデルである Claude Fable と GPT-5.6 へのアクセスは停止されています)。

しかし、それは単にリリースのタイミングの問題だけではありません。証拠は、能力の向上も加速していることを示しています(ただし、最前線はまだ凹凸があり、AI は多くの点で依然として弱いです)。特に AI が実際の業務を遂行する能力を見ると、この傾向は顕著です。人間がどの程度の作業を AI に任せられるかを測定しようとするいくつかの良い評価があります。その中で最も有名な 2 つは、METR と英国の公式政府 AI セキュリティ研究所によるもので、1 つのプロンプトで AI が実行できる人間のプログラマーの労働時間の量を推定しています。GDPval は、多くの分野の人間専門家と AI のパフォーマンスを専門家の審査員を用いて比較します。これらすべてが、指数関数的な成長を上回るペースで増加しています。

同様の実験を行っている別の組織、Epoch は最近、Opus 4.7 が独自に 14 時間作業を行った結果、人間がエンジニアリングに 2〜17 週間を要するソフトウェアパッケージの構築に成功したことを発見しました(トークンコストは 251 ドル)。再び、AI システムがすべてのテストに合格するわけでも、常に低コストで運用できるわけでもないですが、確実に非常に急速な速度で進化しています。私の個人的な実験では、Fable が自律的に 9 時間作業を行い、チームであれば 1 週間以上を要する非常に複雑なソフトウェアプロジェクトを実行できたことがわかりました。

これまで私は、最も「知的」であるフロンティアモデルに焦点を当ててきました。これらは 3 つの米国の企業、Anthropic、OpenAI、Google(ただし Google が新しいモデルを発表して以来は少し時間が経っています)によって開発されています。しかし、フロンティアモデルから通常 6〜12 ヶ月遅れて登場するもう一つの AI モデル群が存在します。これらはすべて中国発のものであり、「オープンウェイト」モデルです。これは、リリース後に誰でも使用や改変が可能であることを意味し(フロンティアモデルがプロプライエタリであるのとは対照的です)。このため、運用コストは非常に安価になります。これらもまた指数関数的な改善曲線を描いて上昇していますが、米国のモデルにはまだ遅れをとっています。これは、複雑な数週間にわたるコンサルティング業務をシミュレートし、AI が多種多様な分析を行わなければならない「AA-Briefcase」と呼ばれるテストにおける私の AI パフォーマンスグラフから確認できます。オープンウェイトモデルは、クローズドな米国製モデルに続く独自の指数関数的曲線上にあります。

しかし、抽象的なグラフでは限界があり、フロンティアがいかに不規則であるか(また、オープンウェイトモデルは非常に印象的ではあるものの、ベンチマークが示すほどには常に性能を発揮するわけではないという事実)を隠してしまう可能性があります。真の洞察を得るためには、AI を異なるユースケースで実際に試し、あなたにとって重要な領域においてどれほど優れているかを厳密に評価する必要があります。面白い例として、私は AI に時間の経過とともに変化する港のインタラクティブなシミュレーションを構築させるテストを作成しました。すべての結果はここで試すことができます。このテストは、デザインやスタイルのアプローチ、さらには判断力といった分野でモデル同士がどれほど異なるかについて興味深い視点を与えてくれると思います。システムがより長いタスクを実行するにつれて、ベンチマーク化が難しいこれらの要素はますます重要になります。

AI の使い方が変化している
AI がより長いタスクを実行できるようになるにつれ、人々が AI を使う方法も変化しています。最近まで、AI を使う支配的な方法は「共知能(co-intelligence)」として利用することでした。AI に何かを依頼し、結果を確認した上で、次のステップを依頼するという形です。 careful なプロンプトと人間の注意によって、複雑で長期的なタスクを AI に実行させることができました。
この AI の利用方法は依然として一般的で有用ですが、価値ある仕事において AI が使われる方法としては次第に主流ではなくなっています。長期間稼働し、賢く、自己修正機能を持つ AI システムは、絶え間ない人間の介入を必要とせず、異なる働き方を要求します(これは私の近刊『Co-Existence』の主題でもありますので、こちらで予約を検討されるのも良いでしょう)。また、チャットボットとは異なり、エージェントには追加的な仕組みが備わっています。AI にツールや行動環境へのアクセスを与えるハルネス(harness)や、Claude Code や OpenAI の Codex のようにエージェント向けに構築されたアプリです。その結果、すでに向上しつつある AI モデルの能力は、優れたハルネスやアプリによってさらに高めることができます。
したがって、仕事の内容は次第にチャットボットと共同で作業を行うことよりも、エージェントに仕事を割り当てることに焦点を移しています。OpenAI と経済学者による共同研究では、この変化が組織内部でいかに急速に進んでいるかが示されています。重要なのは、単なるプログラマーだけがエージェントを利用しているわけではないことです。法務、人事、その他の非技術部門でも、ほぼ同じペースでエージェントの導入が進んでいます。OpenAI は、他の職場で何が起こるかを予兆する「炭鉱のカナリア」のような存在と言えるかもしれません。

次第に、OpenAI での仕事は AI の管理を行うもののように見えてきています。OpenAI の労働者の四分の一が、週に少なくとも一度、同時に 4 つ以上のエージェントを稼働させています。そして、コーディングが専用のハーンチとアプリ内で AI によって行われるようになると、他の役割も一種のコーダーへと変わり始めます。彼らはその点において非常に優れています。Claude Code ユーザーに関する別の研究では、ソフトウェアエンジニアが Claude Code を実際にコーディングタスクに使用した際、他の職業と同様の成功率を収めていることが示されました。

実際に重要だったのは、ユーザーの職業ではなく、その専門性でした。ある分野での経験が深いほど、その分野で Claude Code を使用して成功する確率は高まりました。さらに興味深いことに、各プロンプトから得られる有用な出力も多くなりました。

私たちは、非専門家がチャットボットを使って隙間を埋める世界から、専門家がエージェントを使って仕事を完了させる世界へと移行しています。そして、エージェントを最も効果的に使う方法は、自分自身をマネージャーと捉えることです。
ある瞬間
指数関数的成長にあるということは、一定の期間における変化は、その前の変化よりも大きくなることを意味します。もしあなたの組織が 2025 年の冬以前に AI 計画を策定していたとすれば、それは数時間の作業を高いエラー率で行うシステムを記述したものであったはずです。それから数ヶ月後には、単一のプロンプトから 16 時間以上の作業成果を得られるようになります。
これが、AI がグラフ上の曲線であるにもかかわらず、飛躍しているように感じられる理由です。私たちは能力の安定した倍増を一連の衝撃として経験し続けています。私たちは内部から指数関数的成長を実感するのが非常に苦手であり、現在まさにその中にいるのです。

これは、通常語られる「過熱」に関する物語よりも、AI を巡る混乱をよりよく説明していると思います。AI は、ある日突然現実的なサイバーセキュリティ上の脅威となり、政府の最高レベルで突発的かつ応急的な政策変更を引き起こすまで、実際にそのような脅威となる能力を持っていません。同様に、市場は AI がビジネスモデルを崩壊させる可能性について評価を控えていますが、ある日突然それが可能になると、株価が大幅に乱高下します。これらの揺れ動きは、最終的に安定した状態へと落ち着く未熟な分野の兆候として読み取られます。しかし、私はすぐに落ち着くと考えていません。不安定なのは、人間の速度(あるいは最悪の場合、委員会の速度)で動く機関が、人間とは全く異なる性質を持つ能力曲線を追跡しようとするときに起こることです。そして、ある種の指数関数的な成長が続く限り、またそれが続く限り、その格差は広がる一方です。
購読する
共有する

原文を表示
If you feel like things are accelerating in AI, you are probably right. Better AI models from the leading American AI labs have been releasing more quickly than ever (though government interventions have stopped access to two of the most powerful models, Claude Fable and GPT-5.6).

But it isn't just release timing. The evidence points to accelerating capability gains as well (though the frontier stays jagged, and AIs remain weak in many places). This is especially obvious when we look at the ability of AIs to do real work. There are a few good assessments that try to measure how much human work AIs can do. Two of the most famous, from METR and the UK’s official government AI Security Institute, estimate the amount of human programmer hours’ worth of effort the AI can do with a single prompt. GDPval compares human experts in many fields to AI performance using professional judges. They are all increasing at a better than exponential rate.

Another organization doing similar experiments, Epoch, recently found Opus 4.7, working on its own for 14 hours, was able to build a software package that would take 2-17 weeks of human engineering work (it cost $251 in tokens). Again, AI systems cannot pass every test, nor are they always cheap to run, but they are definitely improving at a very rapid rate. In my own experiments, I found Fable was able to work autonomously for 9 hours to execute on very complex software projects that would have taken a team well over a week to do.

So far, I have focused on the frontier models, those with the highest “intelligence.” They are made by three American companies — Anthropic, OpenAI, and Google (though it has been a while since Google has released a new model). But there is a second set of AI models that typically lag 6-12 months behind the frontier, all of which are from China. These are open weights models, which means that anyone can use or modify them after release (as opposed to the frontier models which are proprietary). That makes them quite cheap to operate. They, too, are climbing up an exponential improvement curve, though lagging the American models. You can see this in my graph of AI performance in a test called AA-Briefcase, which simulates a complex multi-week consulting engagement where AI has to do many kinds of analysis. The open-weights models are on their own exponential curve, behind closed US models

But abstract graphs only get you so far, and they can hide how jagged the frontier is (and also the fact that the open weights models, while very impressive, do not always perform as well as their benchmarks would indicate). To get real insight, you need to try using AI for different use cases and rigorously assess how good they are in the areas that matter to you. As a fun example, I created a test where AIs have to build an interactive simulation of a harbor evolving over time. You can play with all the result here. I think it gives an interesting perspective on how much models can differ from each other in areas like design, stylistic approach, and even judgement. As systems do ever longer tasks, these hard-to-benchmark factors become more important.

The way we use AI is changing
As AIs can do longer and longer tasks, the way people are using AI is changing. Until recently, the dominant way to use AI was as a co-intelligence. You would ask the AI to do something, check the results, and then ask for it to do the next step of your job. By careful prompting and human attention, you could guide AIs to do complex and long-term tasks.
This approach to using AI is still common and useful, but, increasingly, it is not the way AI is being used for valuable work. Long-running, smart, and self-correcting AI systems do not need constant human intervention, and they require a different way of working (this is also the subject of my upcoming book, Co-Existence, which you might want to pre-order here). And, as opposed to chatbots, agents come with extra machinery: harnesses that give the AI access to tools and an environment to act in, and apps built for agents like Claude Code or OpenAI's Codex. As a result, the already increasing ability of AI models can be improved still further by a good harness or app.
So work is increasingly about assigning work to agents, rather than working together with chatbots. A joint study by OpenAI and academic economists shows how quickly this is happening inside their own organization. Critically, it isn’t just coders who are using agents. Legal, HR, and other non-tech functions have adopted agents at nearly the same rate. OpenAI may be a sort of canary in the coal mine for what will happen elsewhere in work.

Increasingly, work at OpenAI looks like managing AI. A quarter of OpenAI workers have at least four agents running at one time every week. And, as coding is done by AIs in specialized harnesses and apps, other roles start to become coders of a sort. And they are good at it. A separate study of Claude Code users found that software engineers had a similar success rate to other professions when actually using Claude code on coding tasks.

What actually mattered was not the profession of the user, but their expertise. The more domain experience someone had, the more successful they were in using Claude Code in that domain. And, even more interestingly, the more useful output they got from Claude from each prompt.

We are moving from a world where non-experts use chatbots to fill in gaps to one in which experts use agents to get work done. And the best way to use agents is to think of yourself as a manager.
A moment in time
Being on an exponential means each change over a fixed window is larger than the one before it. If your organization wrote an AI plan any time before the winter of 2025, it described a system that could do a couple of hours of work with a fairly high error rate. A few months later, you can get sixteen hours or more of work from a single prompt. This is why AI keeps feeling like it is making leaps, even though it is a curve on a graph, we keep experiencing a steady doubling of capability as a series of shocks. We are very bad at feeling exponentials from the inside, and we are currently inside one.

I think this also explains the turbulence around AI better than the usual stories about hype. AI is not capable of being a real cybersecurity threat until suddenly it is, causing sudden and improvised policy changes at the highest level of government. Markets discount whether AI might threaten to undermine a business model until suddenly it can, leading to massive swings in stocks. These lurches these get read as signs of an immature field that will eventually settle into something stable. I don’t think it is going to settle anytime soon. The instability is what happens when institutions that move at the speed of people (or worse, committees) try to track a capability curve that is very much not human in nature. And as long as we are on some sort of exponential, and for as long as it lasts, the gap only widens.
Subscribe now
Share

関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み