Scale Labs、エージェント評価ベンチマーク「TERMINAL-BENCH 3.0」を公開
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Scale Labs Blog
Scale Labs は、AI エージェントの能力をより厳密に評価するため、Terminal-Bench 3.0 を公開し、特に科学・金融・システム分野で難易度を大幅に引き上げている。
AI深層分析を開く2026年8月1日 14:11
AI深層分析
キーポイント
難易度の大幅な引き上げ
既存のベンチマークがモデルによって半分程度クリアされる状態になったため、次世代 Frontier モデルでも 40% 未満しか解決できないレベルまでタスクを強化した。
Scale の大規模な貢献
専門家のネットワークとデータ運用体制を活用し、単一組織として最も多くのタスクを提供し、各タスクの検証プロセスも厳格化された。
実環境に近い評価手法
Docker サンドボックス内で実際のシェルを駆使させ、エージェントの自己申告ではなくコードのコンパイルやサーバー起動などの最終結果で採点する。
信頼性の高いタスク作成プロセス
すべてのタスクは自動化されたチェックと人間のレビューを経て採用され、失敗事例の分析を通じて真の能力不足を評価対象とする。
実行ではなく判断の欠陥が課題
モデルは指示に従う実行自体は得意だが、問題構造の理解や制約に合った方法選択といった判断段階で頻繁に失敗する。
重要な引用
"The best benchmarks measure the edge of what models can do. TERMINAL-BENCH 3.0 provides a huge diversity of real, hard, economically valuable tasks..."
"Once a benchmark gets that close to solved, it stops telling you much."
These aren't execution failures. They're judgment failures, and they're the kind of error that keeps agents out of real professional work.
Understanding these failure modes at a granular level is what lets us build tasks and verifiers that hold up against increasingly capable systems.
編集コメントを表示
編集コメント
ベンチマークの難易度がモデルの性能向上に追いつかなくなると評価価値が低下するという指摘は、開発現場における評価基準の見直しを促す重要な示唆である。特に Docker 環境での実機テストによる厳格な採点は、AI エージェントの実用性を判断する上で不可欠な要素と言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
技術的な作業の多くはターミナル上で行われます。コードのコンパイル、モデルのトレーニング、失敗したパイプラインの追跡、そしてシステム全体の稼働維持などです。
TERMINAL-BENCH 3.0 が ライブ になりました。より広範なカバレッジと難易度の高いタスク、そして最先端の追跡機能を備えています。このオープンソースでコミュニティ主導のベンチマークは、コマンドライン環境における複雑なタスクを通じて、AI エージェントがどれほどよく機能するかを測定するものです。
各タスクは、実機から隔離されたクリーンで自己完結型の Docker サンドボックス内で実行されます。タスクは自然言語による指示から始まり、エージェントが実際のシェルを操作し、出力を読み取り、人間と同じように問題を解決します。採点では、エージェント自身が行ったことの説明に頼るのではなく、最終結果のみをチェックします。コードはコンパイルされたか、サーバーは起動したか、パイプラインは正しい回答を返したか——それが問われます。
Scale Contributions
Scale は、ローンチセットにおいて単一組織として最も多くのタスクを提供しました。これを実現するために、専門家チームとエンジニアリングチームを拡大し、各タスクの作成と検証を行いました。この大量のタスクは、2 つの強みによって支えられています。一つは非常に規模が大きく熟練した専門家のネットワークであり、もう一つはその専門性をスケールして検証済みのタスクに変換するためのデータ運用体制です。
他のプロバイダーが対応可能な有資格者の数に制限される中、私たちは困難なドメインに対して深くリソースを投入し、品質を犠牲にすることなく複数の作業ストリームを並行して実行することが可能です。
大半のタスクは科学分野に特化したドメインワークフローです。これらはエンドツーエンドでツール依存度が高く、仕様を明確に定義するのが難しく、検証が緩い場合は手抜きで済ませてしまうリスクがあります。残りのタスクは金融やシステム関連をカバーしています。
これらのタスクを適切に設計するには、正確なタスク仕様の策定、抜け道を見つけられないような検証プロセスの構築、そして「正解」の厳密な定義が必要です。これらは各分野で熟練した貢献者たちのネットワークから生まれました。彼らは自らの専門領域において何が正解かを理解しており、それに耐えうるタスクと検証器を構築できます。
"最高のベンチマークは、モデルが到達できる限界を測るものです。TERMINAL-BENCH 3.0 は、プログラムによって厳密に検証される、多様で実用的かつ困難な経済価値の高いタスクを提供します。最前線のモデルでも、これらのタスクの40%未満しか解決できません。この厳格さが、Scale AI が最大の貢献者の一人となっている理由であり、私がこのパートナーシップに大きな期待を寄せる所以です。」— Scale AI 工学担当バイスプレジデント アーカシュ・サバルワル氏
なぜ TERMINAL-BENCH 3.0 なのか
Terminal-Bench 2.0 では、最前線のモデルがすでにタスクの約半分をクリアしており、予想よりも早く達成されていました。ベンチマークがこれほどまでに解決可能な状態になると、もはやモデルの能力について多くを語れなくなります。TERMINAL-BENCH 3.0 はタスクの難易度をさらに引き上げ、ベンチマークが最前線の能力を追うのではなく、それを測定し続けるように設計されています。
Harbor はその基盤となるオープンソースフレームワークです。コンテナ環境内でエージェントやモデルの評価・チューニングを行うための仕組みであり、TERMINAL-BENCH 3.0 も Harbor ネイティブのベンチマークとして構築されています。
実務的には、TERMINAL-BENCH 3.0 にタスクを登録する際、個別に作成されたスクリプトではなく、共有かつ再利用可能な評価セットに直接接続されることになります。対象となるタスクは約 16 のカテゴリにまたがり、ソフトウェアエンジニアリング、セキュリティ、科学計算からデータサイエンス、機械学習、システム管理、Web 設定、ハードウェアおよびカーネル開発、ゲーム、デバッグまで多岐にわたります。難易度はイージー、ミディアム、ハードの 3 レベルで構成されています。
タスクの信頼性を支えるもの
ベンチマークの質は、そこに含まれるタスクの質そのものです。研究者たちはタスクの選定プロセスに細心の注意を払っています。このプロセスが私たちの作業スタイルを形作ったため、少し詳しく見てみましょう。
- 自動化チェックの優先:人間によるレビューに入る前に、すべてのタスク PR は一連の自動チェックを通過する必要があります。
- 全工程での人的関与:タスク作成から品質管理に至るまで、あらゆる段階でレビュアーが関与し、マージされる前に厳格な審査が行われます。
- 失敗モードの分析:エージェントの実行結果は慎重に読み込まれ、なぜ実行が失敗したのかを解明します。タスクは、指示の不備や環境の破綻ではなく、真の意味での能力不足によってエージェントが達成できなかった場合にのみ採用されます。
失敗こそが目的
困難なタスクが最も有用に果たす役割は、エージェントのどこで破綻するかを明確に示すことです。TERMINAL-BENCH の各タスクを通じて、単なる知識不足を超えた一貫した失敗パターンを確認できます。
モデルは指示された方向へ進めば実行能力は高いものの、最初のステップ——問題文を読み込み、その構造を理解し、制約条件に実際に適合する手法を選択すること——でつまずきます。便利だからといって安易な方法を選ぼうとする傾向が見られます。
エージェントは、物理的な推論やドメイン固有の思考を避け、汎用的なアプローチを採用しようとするケースがあります。また、分析過程で正しくニュアンスのある解決策を導き出しながら、なぜかそれを却下し、「わずかに劣る」単純な形式に依存してしまう様子も観察されます。さらに、一見妥当で構造的にも整っているように見える回答を生成するものの、その根拠となる前提が誤っているケースもあります。
これらは実行能力の欠陥ではありません。判断力の欠如であり、これがエージェントを実際の専門業務から遠ざける要因となっています。
ベンチマークが捉えているのはまさにこの点です。そして、私たちのタスク設計はこうした失敗様式を明らかにすることを目的としています。微細なレベルでこれらの失敗モードを理解することが、より高度化するシステムに対しても耐えうるタスクと検証器を構築する鍵となります。
今後の展望
TERMINAL-BENCH 3.0 は、ターミナルエージェントが急速に進化しているタイミングで登場しました。この分野には、進化のスピードに追いつくためのより優れた評価指標が必要です。
ターミナルエージェントの構築や評価に取り組んでいる方へ、これはオープンプロジェクトです。タスクの提案、PR のレビュー、あるいは単にプロセスを追うことも可能です。議論は Discord で行われており、コードとタスクリポジトリは GitHub にあります。
このようなデータが必要ですか?
私たちが貢献したタスクは、科学や専門分野で Scale が構築する評価データのほんの一部です。難易度の高い問題、精密な仕様、そして専門家による厳密な検証に耐えるバリデーターが特徴です。エージェントの訓練や評価を行っており、この水準のデータがどのようなものか知りたい場合は、サンプルタスクのリクエスト をご検討ください。
原文を表示
A lot of technical work happens in the terminal: compiling code, training models, tracing failing pipelines, and keeping entire systems running. TERMINAL-BENCH 3.0 is now live with broader coverage, harder tasks, and frontier tracking. This open-source, community effort benchmark measures how well AI agents can perform by testing them on involved tasks in a command-line environment.
Each task runs inside an isolated Docker sandbox, a clean, self-contained environment sealed off from the real computer, and starts with a plain-language instruction. The agent drives a real shell, reads the output, and works the problem the way a person would. Scoring ignores the agent's own account of what it did and checks the end result instead: did the code compile, did the server come up, did the pipeline return the right answer?
Scale Contributions
Scale contributed the most tasks of any single organization in the launch set, scaling up our expert and engineering teams to create and verify each one. That volume comes from two strengths working together: a very large and highly skilled expert network, and a data operation built to turn that expertise into verified tasks at scale. Where other providers are limited by how many qualified contributors they can reach, we can staff difficult domains deeply and run many workstreams in parallel without giving up quality.
The majority are science-focused domain workflows: end-to-end, tooling-heavy problems that are hard to specify and easy to fake your way through if the checks are loose. The rest cover finance and systems.
Getting the tasks right took precise task specs, verification that catches shortcuts, and a strict definition of what "correct" means. These tasks came out of a network of skilled contributors across domains who know what a right answer looks like in their own field and can build a task and a verifier that holds up to it.
*"The best benchmarks measure the edge of what models can do. TERMINAL-BENCH 3.0 provides a huge diversity of real, hard, economically valuable tasks, each programmatically verified, where the best frontier models can solve less than 40% of the tasks. That rigor is why Scale is one of its largest contributors, and why I am so excited about this partnership." — Aakash Sabharwal, VP of Engineering at Scale AI*
Why TERMINAL-BENCH 3.0
On Terminal-Bench 2.0, frontier models were already clearing around half the tasks, and they got there faster than expected. Once a benchmark gets that close to solved, it stops telling you much. TERMINAL-BENCH 3.0 pushes task difficulty higher so the benchmark keeps measuring frontier capabilities instead of trailing it.
Harbor sits underneath all of this. It's an open-source framework for evaluating and tuning agents and models inside container environments, and TERMINAL-BENCH 3.0 is built as a Harbor-native benchmarks. In practice that means a task contributed to TERMINAL-BENCH 3.0 plugs into a shared, reusable evaluation setup rather than a one-off script. The tasks span roughly 16 categories, from software engineering, security, and scientific computing to data science, machine learning, systems administration, web configuration, hardware and kernel work, games, and debugging, at easy, medium, and hard levels.
What Makes the Tasks Trustworthy
A benchmark is only as good as its tasks, and researchers are careful about how they get in. It's worth walking through, because the process shaped how we worked.
- Automated checks first. Every task PR has to pass a set of automated checks before a human looks at it.
- Human in the loop throughout. Reviewers are involved at every stage of task creation and quality control before anything merges.
- Failure-mode analysis. Agent rollouts are read closely to understand why a run failed. A task is kept only when the agent falls short because of a genuine capability gap, not because the instructions were unclear or the environment was broken.
Failures are the Point
The most useful thing a difficult task does is demonstrate exactly where an agent breaks. Across our tasks, we see consistent failure patterns that go beyond simple knowledge gaps. Models are strong at execution once pointed in the right direction, but they struggle with the step that comes first: reading the problem, understanding its structure, and choosing a method that actually fits the constraints rather than reaching for the most convenient one.
We see agents skip over physical or domain-specific reasoning in favor of generic approaches. We see them derive a correct, nuanced solution during their own analysis and then argue themselves out of it, defaulting to a simpler form because it fits only slightly worse. We see them produce answers that look plausible and well-structured but rest on the wrong assumptions. These aren't execution failures. They're judgment failures, and they're the kind of error that keeps agents out of real professional work.
This is what the benchmark catches, and it's what our task design is built to expose. Understanding these failure modes at a granular level is what lets us build tasks and verifiers that hold up against increasingly capable systems.
What's next
TERMINAL-BENCH 3.0 arrives while terminal agents are improving quickly, and the field needs better measurement to keep up.
If you build or evaluate terminal agents, this is an open project where you can propose a task, review a PR, or just follow along. The conversation is on Discord, and the code and task repositories are on GitHub.
Want data like this?
The tasks we contributed are a small window into the evaluation data Scale builds across scientific and professional domains: hard problems, precise specifications, and verifiers that hold up to expert scrutiny. If you're training or evaluating agents and want to see what data at this bar looks like, request sample tasks.
AI算出
主要ニュースainew評価標準
記事は AI エージェントの能力を測定する新しいベンチマークの公開を報じており、AI テスト・評価の核心テーマであるため関連性は最高です。既存の 2.0 から難易度とカテゴリが大幅に拡張された具体的な新事実があるため新規性も高く評価されます。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み