AIの形状:不規則性、ボトルネック、顕著な特徴
筆者らは2023年、「ジャグドフロンティア」という用語を提唱し、AIが人間の直感とかけ離れた能力の偏り(特定のタスクは超人的に優れ、他は著しく劣る)を説明した。この不規則性はAIの主要な特徴であり、混乱の原因となっている。
2023 年という古くからの AI の時代に戻りましょう。私と共著者たちは、AI がタスクの難易度に対する人間の直感とはあまり一致しない方法で、ある作業は驚くほどよくこなす一方で別の作業は驚くほど苦手とするという奇妙な能力を説明するために、ある用語を発明しました。私たちはこれを AI 能力の「ジャグド・フロンティア(Jagged Frontier)」と呼びましたが、これは依然として AI の主要な特徴であり、混乱の絶え間ない源泉となっています。なぜ AI は、高度な医療診断や非常に難しい数学(はい、最近までこのフロンティアの外にありましたが、今では数学が本当に得意です)においては人間を超えた能力を持ちながら、比較的単純な視覚パズルや自動販売機の運転においてはまだ苦手なのでしょうか?AI の正確な能力はしばしば謎のままであるため、AI が見た目よりも使いにくいのは当然のことなのです。
私は、ジャグドネス(不均衡さ)は今後も AI において大きな部分を占め続けると思いますが、それが何を意味するかについては確信が持てません。トマス・プエヨ(Tomas Pueyo)氏は X で彼のビジョンを概説したこの viral な画像を投稿しました。彼の見解では、拡大するフロンティアはジャグドネスを凌駕していくでしょう。確かに AI はある分野が苦手であり、改善が進んでも相対的に苦手なままとなるかもしれませんが、集約された人間の能力のフロンティアは主に固定されており、AI の能力は急速に成長しています。もし AI が自動販売機の運転において相対的に苦手であっても、それでも人間よりも優れた存在になるのであれば、それは問題なのでしょうか?

未来は常に不確実ですが、この考え方には仕事と技術の性質に関するいくつかの重要な側面が欠けていると思います。第一に、フロンティア(最先端領域)は確かに非常にギザギザしており、そのギザギザさゆえに、人間のタスクと完全に重なり合わない超知能 AI が生まれる可能性があります。例えば、このギザギザさの主要な原因の一つは、大規模言語モデル(LLM: Large Language Models)が新しいタスクを記憶し、永続的にそこから学習できない点にあります。多くの AI 企業がこの問題に対する解決策を追求していますが、この問題は研究者たちが予想するよりも解決が難しい可能性があります。記憶機能がなければ、AI は他の分野では人間を超えた能力を持っていても、人間ができる多くのタスクを実行することに苦労することになります。コリン・フレイザーは、このような AI と人間の重なり合いがどのようなものになるかを示す 2 つの例を描きました。AI が一部の分野で確かに人間を超えている一方で、他の分野では人間レベルに遠く及ばないか、あるいは全く重ならない様子がわかります。これが真実であれば、AI は人間と補完し合って働く新たな機会を生み出すことになります。なぜなら、私たちそれぞれが異なる能力を備えているからです。
これらは概念的な図解ですが、科学者たちのグループが最近、AI の能力の形状をマッピングしようとし、それが不均一に成長していることを発見しました。これはまさに、不揃いな最前線が予測する通りです。読書、数学、一般知識、推論 — AI はこれらすべての分野で急速に改善しています。しかし、記憶については前述した通り、非常に改善の少ない弱点となっています。プロンプトの改善やより優れたモデル(GPT-5.2 は GPT-5 よりもはるかに優れています)によって最前線の形状が変わる可能性はありますが、「不揃いさ」そのものは残ります。
ボトルネック
たとえわずかな不整(ジャグネス)であっても、超知能を持つ AI がタスクを自動化できなくなる問題を引き起こす可能性があります。システムの機能性は、その最悪のコンポーネントによって決定されます。私たちはこれらの問題をボトルネックと呼びます。いくつかのボトルネックは、AI が特定のタスクにおいて頑固に人間未満の能力しか持たないことに起因します。LLM(大規模言語モデル)を備えたビジョンシステムは医療画像の読影が十分ではないため、まだ医師を代替することはできません;LLM は押し返すべき場面で過度に親切すぎるため、まだセラピストを代替することはできません;ハルシネーション(幻覚)は頻度が減ったとしても依然として存在するため、100% の精度が求められるタスクにはまだ対応できません。そして他にも多くの例があります。フロンティア(最先端技術の領域)がさらに拡大すれば、これらの問題の一部は消滅するかもしれませんが、弱点こそがボトルネックの唯一の形態ではありません。
いくつかのボトルネックは、能力とは無関係のプロセスに起因します。たとえ AI が従来の方法よりも劇的に迅速に有望な薬剤候補を特定できるようになったとしても、臨床試験では依然として実際の患者が必要であり、彼らの募集、投与、モニタリングには実際の日数がかかります。FDA(米国食品医薬品局)もなお申請に対する人間の審査を要求しています。たとえ AI が優れた薬のアイデアの生成率を10倍以上に引き上げたとしても、制約となるのは発見の速度ではなく承認の速度です。ボトルネックは知能から制度へと移行し、制度は「制度特有の速度」でしか動きません。
Google の Nano Banana Pro から提供された画像。詳細は後ほど!
AI がほぼ完全に人間を超えている領域であっても、エッジケースには人間の介入が必要となる場合があります。例として、Cochrane レビュー(多くの医学研究を統合して特定のトピックに関する科学的合意を導き出す、著名で徹底的に調査されたメタ分析)を AI を用いて再現した研究があります。研究者チームは、適切にプロンプトを与えられサポートされた GPT-4.1 が「Cochrane レビューの 1 号分(n=12)全体を 2 日間で再現・更新し、これは従来の体系的レビュー作業で約 12 人年分に相当する」と報告しました。AI は 146,000 件以上の引用文献をスクリーニングし、論文全文を読み込み、データを抽出し、統計分析を実行しました。実際には、精度において人間の審査員を上回る結果を示しています。奇妙なことに、関連する研究の発見や適切な数値の抽出、結果の統合といった多くの知的に困難な作業は、すでに最先端技術の領域内に確立されています。しかし、AI は補足ファイルへのアクセスができず、未発表データの請求のために著者にメールを送ることもできません。これらは人間が日常的に行う業務です。これらの欠陥はレビュー全体の誤りの 1% 未満を占めますが、このわずかなエラーが存在する限り、プロセスの完全な自動化は不可能となります。12 人年分の作業が 2 日間に短縮されるのは、科学の実践方法に精通した人間がエッジケースに対応する場合に限られます。
これはパターンです:不整(ジャグネス)がボトルネックを生み、そのボトルネックとは、非常に賢い AI でさえも人間を容易に代替できないことを意味します。少なくとも今はまだそうです。これはある面では良いことかもしれません(急速な失業を防ぐため)、別の面では苛立たしいことです(科学的研究のスピードアップを私たちが望むほどには進められないため)。また、ボトルネックは、AI 企業の仕事を、AI を阻んでいる要因に対して AI の能力を高めることに集中させることになります。数学的能力が明白な障壁となった後に急速に向上したようにです。
歴史学者トーマス・ヒューズはこの現象に対する用語を持っていました。電気システムの発展を研究する中で、彼は進歩がしばしば単一の技術的または社会的問題で停滞することに気づきました。彼はこれを「逆突出部(リバース・サリエンツ)」と呼びました。これはシステムが飛躍的に前進することを阻む、唯一の技術的または社会的な問題です。

逆突出部(リバース・サリエンツ)
ボトルネックは、AI が実際には決してあることができないかのような印象を与えることがあります。しかし現実には、進歩が単一の不整な弱点によって阻まれているのです。その弱点が「逆突出部」となり、AI 研究所が突如としてその問題を解決した瞬間、システム全体が一気に前進します。
先月におけるこの現象の最も強力な例は、Google の新しい画像生成 AI「Nano Banana Pro」です(はい、AI 企業はまだ名前を付けるのが下手なのです)。これは二つの進歩を組み合わせたものです:非常に優れた画像作成モデルと、そのモデルを指示するために活用できる非常に賢い AI です。必要に応じて情報を検索しながらモデルを導きます。例えば、私のおたくのテストにおける究極版として Nano Banana Pro にプロンプトを入力すると、「オッターである科学者たちがホワイトボードを使って、AI の WiFi テストにおけるイーサン・モリックのオッターが飛行機に乗っている様子について説明し、それが通過したことを証明している。壁一面にラップトップを使用する飛行機のオッターの写真が並んでいる」というものです。私はこれを得ます:

一貫性のある言葉、異なる角度、影、主要な誤字はありません。非常に素晴らしいものです。覚えておいてください、「WiFi を使用する飛行機に乗ったオッター」というプロンプトは、2021 年にこの画像を得ていました:

しかし、実は非常に優れた画像生成能力こそが、多くの新機能にとってボトルネックとなっていたのです。例えば PowerPoint のスライドデッキを作成するケースを考えてみましょう。主要な AI 企業は皆、自社の AI に PowerPoint を作成させることに注力してきました。その手段として、AI が得意とするコンピュータコードを記述させ、ゼロから PowerPoint を生成させています。これは困難なプロセスですが、Claude と ChatGPT の両方とも大幅に改善されており、スライドの内容がやや地味であるとしてもです。例えば、私の著書『Co-Intelligence』を Claude に読み込ませて、スライドデッキ要約を作成させました。モデルは非常に賢いのですが、その PowerPoint デッキはコードで記述しなければならないという制約によって限界があります。

次に、Google の NotebookLM アプリケーションで同じことを試した例を示します。ここでは賢い Gemini AI モデルと Nano Banana Pro を組み合わせて使用しています。こちらはコードを使用せず、各スライドを単一の画像として生成しています。画像の品質が低かった時代にはこれは不可能でしたが、今は突然それが可能になりました。

画像は非常に柔軟性があるため、スタイルやアプローチを自由に試すことができます。私は NotebookLM に学習の科学的根拠に基づく方法について深掘り調査レポートを作成させ、それをさまざまなスタイルで読みやすいように設計された密度の高いスライドデッキに変換してもらいました。一つは手書き風、もう一つは 1980 年代パンクにインスパイアされたもの、「非常にドラマチックでコントラストが強く、背景が鮮やかな黄色のスライド」というスタイルのもの、そしてもちろん、カワウソが飛行機に乗っているというテーマのスライドです。

多くの点で、Claude と Gemini の両方にとっての「難しい部分」はフロンティア(最先端領域)の中にあります。これらは単にソース資料、トピック、アイデアを受け取れば、それをスライド形式で要約することができます。ハルシネーション(幻覚・誤情報生成)は非常に稀であり、出典も正確です。カワウソの比喩を作成したり、パンク風の記述を考案したりすることも可能です。これは知的に要求される部分ですが、AI はすでに一年以上この能力を持っています。しかし、スライドやその他の視覚的プレゼンテーションを作成することは、テキストの壁を有用なものにするためのボトルネックでした。問題は完全に解決されたわけではありません:画像は完璧ではなく、編集もできません(ただし、これは間もなく修正されるとのことです)。それでも、これから何が起きるのかは見えてきます。
多くの転換点
たとえ AI が分析やパワーポイント作成において人間を超えた能力を獲得したとしても、それが必ずしもコンサルタントやデザイナーの仕事を AI に置き換えることを意味するとは考えません。これらの仕事には、AI が苦手とする一方で人間が卓越している「ジグザグな最前線」に沿った多様なタスクが含まれています。多くの関係者から情報を収集し合意形成を図れるか?人々が実際に必要としているものを決定づける暗黙のルールを理解できるか?AI の素材とは一線を画し、深い課題に独自に対応する何かを創出できるか?このジグザグな最前線には、人間の仕事にとって多くの機会が存在します。
しかし、逆突出(reverse salient)に焦点を当てることでボトルネックが突然解消され、飛躍的な前進が見られることも予想されます。かつては人間のみが行っていた業務の領域が、AI が実行可能なものへと変化していくのです。AI の行先を理解したいなら、ベンチマーク(評価指標)を見るのではなく、ボトルネックに注目すべきです。一つでも突破されれば、その背後にあったすべてのものが一斉に流れ込んでくるからです。画像生成はこれまでプレゼンテーションや文書作成、あらゆる種類の視覚コミュニケーションを阻害する要因となっていました。しかし今はもうそうではありません。次なるボトルネックは何でしょうか?記憶力?リアルタイム学習?物理世界での行動実行能力?
どこかの AI 研究所では、今まさにこれらのボトルネックを逆突出として扱っているはずです。突破される際に多くの警告があるとは考えられません。しかし、ジグザグな最前線は両刃の剣です。これまでに起こったすべての飛躍的前進が、人間が必要とされる新たなエッジ(境界)をさらに生み出してきました。今後にも多くの飛躍的前進が待っています。同時に、多くの機会も存在するでしょう。その両方に注意を向けるべきです。
購読する
共有する

私は Gemini 3 に、この投稿のための魅力的なタイトル画像を作成するよう依頼しました。これがその結果です。
原文を表示
Back in the ancient AI days of 2023, my co-authors and I invented a term to describe the weird ability of AI to do some work incredibly well and other work incredibly badly in ways that didn’t map very well to our human intuition of the difficulty of the task. We called this the “Jagged Frontier” of AI ability, and it remains a key feature of AI and an endless source of confusion. How can an AI be superhuman at differential medical diagnosis or good at very hard math (yes, they are really good at math now, famously outside the frontier until recently) and yet still be bad at relatively simple visual puzzles or running a vending machine? The exact abilities of AI are often a mystery, so it is no wonder AI is harder to use than it seems.
I think jaggedness is going to remain a big part of AIs going forward, but there is less certainty over what it means. Tomas Pueyo posted this viral image on X that outlined his vision. In his view, the growing frontier will outpace jaggedness. Sure, the AI is bad at some things and may still be relatively bad even as it improves, but the collective human ability frontier is mostly fixed, and AI ability is growing rapidly. What does it matter if AI is relatively bad at running a vending machine, if the AI still becomes better than any human?

While the future is always uncertain, I think this conception misses out on a few critical aspects about the nature of work and technology. First, the frontier is very jagged indeed, and it might be that, because of this jaggedness, we get supersmart AIs which never quite fully overlap with human tasks. For example, a major source of jaggedness is that LLMs do not remember new tasks and learn from them in a permanent way. A lot of AI companies are pursuing solutions to this issue, but it may be that this problem is harder to solve than researchers expect. Without memory, AIs will struggle to do many tasks humans can do, even while being superhuman in other areas. Colin Fraser drew two examples of what this sort of AI-human overlap might look like. You can see how AI is indeed superhuman in some areas, but in others it is either far below human level or not overlapping at all. If this is true, then AI will create new opportunities working in complement with human beings, since we both bring different abilities to the table.

These are conceptual drawings, but a group of scientists recently tried to map the shape of AI ability and found that it was growing unevenly, just as the jagged frontier would predict. Reading, math, general knowledge, reasoning — all were things that AI was improving on rapidly. But memory, as we discussed, is a weak spot with very little improvement. Better prompting or better models (and GPT-5.2 is much better than GPT-5) might change the shape of the frontier, but jaggedness remains.

Bottlenecks
And even small amounts of jaggedness can create issues that make super-smart AIs unable to automate a task. A system is only as functional as its worst components. We call these problems bottlenecks. Some bottlenecks are because the AI is stubbornly subhuman at some tasks. LLM vision systems aren’t good enough at reading medical imaging so they can’t yet replace doctors; LLMs are too helpful when they should push back so they can’t yet replace therapists; hallucinations persist even if they have become rarer which means they can’t yet do tasks where 100% accuracy is required; and so on. If the frontier continues to expand, some of these problems may disappear, but weaknesses are not the only form of bottleneck.
Some bottlenecks are because of processes that have nothing to do with ability. Even if AI can now identify promising drug candidates dramatically faster than traditional methods, clinical trials still need actual human patients who take actual time to recruit, dose, and monitor. The FDA still requires human review of applications. Even if AI increases the rate of good drug ideas by ten times or more, the constraint becomes the rate of approval, not the rate of discovery. The bottleneck migrates from intelligence to institutions, and institutions move at institution speed.

Image from Google’s Nano Banana Pro. More on that in a minute!
And even where the AI is almost completely superhuman, humans may be needed for edge cases. As an example, take a study that used AI to reproduce Cochrane reviews, the famous deeply researched meta-studies that synthesize many medical studies to figure out the scientific consensus on a topic. A team of researchers found that GPT-4.1, when properly prompted and supported, “reproduced and updated an entire issue of Cochrane reviews (n=12) in two days, representing approximately 12 work-years of traditional systematic review work.” The AI screened over 146,000 citations, read full papers, extracted data, and ran statistical analyses. It actually outperformed human reviewers on accuracy. Oddly, much of the hard intellectual work — finding relevant studies, pulling the right numbers, synthesizing results — is solidly inside the frontier. But the AI can't access supplementary files and it can't email authors to request unpublished data, things human reviewers do routinely. This makes up less than 1% of errors in the review, but those errors mean you can't fully automate the process. Twelve work-years become two days, but only if a human with expertise in how science is actually done handles the edge cases.
This is the pattern: jaggedness creates bottlenecks, and bottlenecks mean that even very smart AI cannot easily substitute for humans. At least not yet. This is likely good in some ways (preventing rapid job loss) but frustrating in others (making it hard to speed up scientific research as much as we might hope). Bottlenecks also concentrate the work of AI companies into making the AI better at things that are holding it back, the way math ability rapidly improved once it became an obvious barrier. The historian Thomas Hughes had a term for this. Studying how electrical systems developed, he noticed that progress often stalled on a single technical or social problem. He called these “reverse salients” - the one technical or social problem holding back the system from leaping ahead.

Reverse Salients
Bottlenecks can create the impression that AI will never be able to do something, when, in reality, progress is held back by a single jagged weakness. When that weakness becomes a reverse salient, and AI labs suddenly fix the problem, the entire system can jump forward.
The most powerful example of this from the last month is Google’s new image generation AI, Nano Banana Pro (yes, AI companies are still bad at naming things). It combines two advances: a very good image creation model and a very smart AI that can help direct the model, looking up information as needed. For example, if I prompt Nano Banana Pro for the ultimate version of my otter test: “Scientists who are otters are using a white board to explain ethan mollicks otter on a plane using WiFi test of AI (you must search for this) and demonstrating it has been passed with a wall full of photos of otters on planes using laptops.” I get this:

Coherent words, different angles, shadows, no major misspellings. Pretty amazing stuff. Remember, the prompt “otter on a plane using wifi” got this image in 2021:

But it turns out that really good image generation was the bottleneck for a lot of new capabilities. For example, take PowerPoint decks. Every major AI company has been trying to get their AI to make PowerPoint, and they have done this by having the AIs write computer code (which they are very good at) to create a PowerPoint from scratch. This is a hard process, but both Claude and ChatGPT have improved a lot, even if their slides are a little dull. For example, I took my book, Co-Intelligence, and threw it into Claude and asked for a slide deck summary. The model is very smart, but the PowerPoint deck is limited by the fact that it has to be written in code.

Now here is the same thing in Google’s NotebookLM application, using its smart Gemini AI model combined with Nano Banana Pro. It isn’t using code, it is creating each slide as a single image. When image quality was low, this would have been impossible. Suddenly, it isn’t.

And since images are very flexible, I can play with style and approach. I had NotebookLM do a deep research report on science-backed methods of learning and then turn that into dense slide decks meant for reading in a variety of styles: one that looked hand-drawn, one that was inspired by 1980s punk, one that was “very dramatic and high contrast slides with a bright yellow background,” and, of course, one with an otter-on-a-plane theme.

In many ways, the hard stuff is inside the frontier for both Claude and Gemini, they can just take source materials, a topic, and an idea and summarize it in a slide. Hallucinations are very rare, and the sources are correct. It can create otter analogies or come up with a punk-themed description. This is the intellectually demanding part, and AIs have been capable of it for over a year. But making slides or other visual presentations was a bottleneck to making walls of text useful. The problem isn’t completely solved: images are not perfect, and you can’t edit them (apparently this will be fixed soon), but you can see where things are going.
Many lurches
Even if AI becomes superhuman at analysis and PowerPoint, I don’t think that means AI necessarily replaces the jobs of consultants and designers. Those jobs consist of many different tasks along the jagged frontier that AI is bad at and which humans excel: can you collect information and get buy-in from the many parties involved? Can you understand the unwritten rules that determine what people actually need? Can you come up with something unique to address a deep issue, that stands out from AI material? The jagged frontier offers many opportunities for human work.
Yet, we should expect to see lurches forward, where focusing on reverse salients leads to sudden removals of bottlenecks. Areas of work that used to be only human become something that AI can do. If you want to understand where AI is headed, don’t watch the benchmarks. Watch the bottlenecks. When one breaks, everything behind it comes flooding through. Image generation was holding back presentations, documents, visual communication of all kinds. Now it isn’t. What’s the next bottleneck? Memory? Real-time learning? The ability to take actions in the physical world?
Somewhere, right now, an AI lab is treating each of these bottlenecks as a reverse salient. We won’t get much warning when they break through. But a jagged frontier cuts both ways. So far, every lurch forward leaves yet more edges in which humans are needed. There will be many lurches ahead. There will also be many opportunities. Pay attention to both.
Subscribe now
Share

I asked Gemini 3 to come up with a compelling title image for this post, this is what it made.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み