ブラウザ使用評価基準の設計思想と、評価者(Judge)の一貫性確保について
本文の状態
日本語全文を表示中
詳細モードで約17分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Browser Use Blog
Browser Use はベンチマークの信頼性を高めるため、評価者(Judge)の一貫性確保とノイズ要因の分離に注力し、モデル性能だけでなく評価プロセス自体の質を向上させる新しい基準を提示した。
AI深層分析を開く2026年8月6日 16:31
AI深層分析
キーポイント
ベンチマーク構成要素の再定義
ベンチマークはタスクとそれを判定する「ジャッジ」の二要素から成り立ち、多くの場合後者が軽視されているが、これが数値を決定し、弱点があれば報酬ハックが可能になると指摘している。
評価者(Judge)の一貫性要件
整合性の取れた評価者には、同じ実行結果でのスコア再現性、タスク間での基準の統一、異なるモデルやハッチングに対する同等の評価、そしてランダムな順位変動の排除という4つの条件が求められる。
ノイズの二重構造と対策
スコア変動要因は「エージェントノイズ(Web 環境やモデルの不確実性)」と「ジャッジノイズ(同一実行結果での判定不一致)」に分類され、後者は排除可能であると主張している。
実験によるタスクの不安定性
GPT-5.5 を同一設定で 5 回実行した結果、106 タスクのうち 40 タスクがランダムに合格・不合格となり、モデル自体の性能だけでなく評価プロセスの不安定さが浮き彫りになった。
単一実行の限界
1回の試行ではモデルが何ができるかを測るのではなく、実際に成功する頻度しか測定できない。
重要な引用
A benchmark is two things: the tasks, and the judge that decides whether you succeeded.
An aligned judge has to be consistent in four ways
Agent noise is the web and the model. A site blocks you. A wrong turn on step 12.
So a single run doesn't measure what a model can do. It measures how often it does it.
編集コメントを表示
編集コメント
AI ベンチマークの信頼性は、タスクの難易度だけでなく評価アルゴリズムの安定性に大きく依存する。この分析は、スコア比較を行う際に背後にあるプロセスの質を再考させる重要な視点を提供している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

ベンチマークには2 つの要素が必要です。それは「タスク」と、「成功したかどうかを判定する審判」です。多くの人は審判を後回しにしがちですが、実は公開されるすべての数値を決めるのはこの審判であり、質が低いと結果は報酬ハック(ごまかし)されやすくなります。
そのため私たちは、自社のベンチマーク構築に長い時間をかけました。信頼できる審判とは、以下の 4 つの側面で一貫性を持っている必要があります。
- 同じ実行を二度行う。一度完了したランを再度評価し、同じスコアが得られること。
- タスク間での一貫性。タスク 90 とタスク 3 が、同じ基準で評価されること。
- モデルやハッチング(環境)間での一貫性。Browser Use と Codex が同じ方法で採点されること。
- ランレベルでの一貫性。モデル A がモデル B に勝つ場合、再度評価してもその順位が逆転しないこと。
これらを実現するにはコストがかかります。まずは、スコアのどこまでが「実態」なのかを問う必要があります。
スコアに影響する要因
影響するのは 2 つの要素だけです。そのうちモデルに関係するのは片方だけ。
エージェントノイズとは、ウェブ環境とモデル自体に起因するものです。サイトがアクセスをブロックしたり、12 ステップ目で間違った選択をしたりします。これは測定可能ですが、完全に排除することはできません。
一方、審判ノイズは、同じ実行結果に対して二度評価を行い、異なる判定が下されることです。これなら解消可能です。
エージェントノイズについて
私たちは GPT-5.5 を、同じ 106 のタスクに対して 5 回実行しました。モデルもハッチングも設定もすべて同一です。
各ランで解決できたのは約 89 タスクでしたが、その 89 は毎回異なりました。

各列が一つのタスク、各行が一つのランです。緑色は成功を示します。
64 のタスクは毎回動作しますが、2 つのタスクは一切動作しません。 問題の本質は真ん中の 40 タスクにあります。同じモデルで同じタスクなのに、なぜかランダムに成功したり失敗したりするのです。フィルターの読み間違いや誤ったクリック、あるいは一歩手前で諦めてしまうといった要因が絡んでいます。
1 回だけ実行すればスコアは 85% ですが、5 回試せば 106 タスクのうち 104 が少なくとも一度は成功します。
つまり、単一のランではモデルの「できること」を測ることはできません。それはあくまで「頻度」を測っているに過ぎないのです。
ジェッジ(判定者)のノイズ
ここでエージェントを固定し、そのセットから 104 の完了したラン(人間によるラベル付けが済んでいるもの)を取り出し、判定を行うモデルだけを変更しました。

考え方は同じです。各列がタスク、各行がジェッジで、全員が全く同じ作業結果を読み取っています。
ベンチマークの 45% は、どのジェッジに評価を任せるかによって変わります。 最も寛容なジェッジは 83.7% と報告し、最も厳しいジェッジは 62.5% です。この 21 ポイントの開きは、通常比較されるモデル間の差よりも広いです。
スコアを動かす要因がモデルの変化よりもジェッジの動きの方が大きければ、その変更が本当に役立ったのかどうかを判断することはできません。
ジェッジの構築
トレーシング全体を含む 1 つのプロンプト
最初のジェッジは、単一の LLM 呼び出しで構成されていました。トレーシング全体を入力し、判定結果を出力する形式です。
1 つの検証枠組みでは不十分でした。そこで複数の枠組みを追加したところ、トレースはまるで別物のように見え、数百ステップにわたって実行され、圧縮によって必要な部分が失われてしまいます。
証拠は埋もれてしまう

タスク:2,000 ドル以下の物件を抽出すること。34 ステップでエージェントはフィルターの金額を 20,000 ドルに設定してしまいます。その後のすべての処理は、一貫性があり自信を持って行われているのに、結果は誤っています。
これをプロンプトに盛り込むことはできません。また、1 回の試行では、この問題を見つけるのは一度きりのチャンスです。
そこで、審査役もエージェントになりました。トレースとワークスペースを受け取り、Codex や Claude Code(読み取り権限付き)のように調査を開始します。CSV ファイルを開き、34 ステップまで戻り、表示されたページの数値と比較して確認できます。
審査役が見られるもの
エージェント型の審査役が検証できるのは、アクセス可能な情報のみです。ただし、これは簡単に破綻する仕組みでもあります。スクリーンショットは容量を圧迫します。ページのテキストも同様です。予算に合わせて何かが切り捨てられてしまいます。

誤った部分を切り取ると、真実の主張も検証不能になります。審査役にとっては、それが捏造と見分けがつきません。結果として、正しい答えを「ハルシネーション(幻覚)」だと見なすようになります。プロンプトに触れる前に、証拠に関する契約条件を正しく設定しておく必要があります。
合格・不合格の問いをやめた理由
判定者は何かしらの出力を返す必要がありますが、単に「合格・不合格」で判断するのは誤りです。なぜなら、ほとんどのタスクは二値(バイナリ)ではないからです。
実際のプロンプトは曖昧なものです。「上位貢献者の一人をチェックする」「トレンドの puzzel を開く」といった指示に対し、どれが正解だったかを判定者自身を含め誰にも証明できません。また、エージェントの失敗も完全なものとは限りません。JSON 形式を要求したのに、正しいデータを含む CSV が返ってきた場合でも、400 ステップ連続で正答していたなら、それは「ゼロ」ではありません。
当社のセットから一例を紹介します。
スペースに移動し、推奨スペースのいずれかにアクセスして確認してください。そのスペース内の上位貢献者のプロフィールを開き、フォロワー数が何人か返してください。
エージェントは227 人のフォロワーと判定しました。しかし、推奨スペース一覧ページはログインが必要なゲートがかかっており、代わりに公開された別のスペースを使用していました。
11 人中5人の判定者がこれを合格とし、6人が不合格と判断しました。 同じスクリーンショット、同じフォロワー数です。問題は「代替のスペースがカウントされるか」の有無で意見が割れました。
つまり、結果は 0 でも 1 でもありません。ここでは 60 と評価します。45 点とする議論も成り立ちますが、その議論こそが本質です。スコアは判断を排除するものではなく、判断の根拠を可視化するものです。

そこで私たちは数値を求めます。「要求された成果のうち、正しい根拠に基づいて達成された割合はどれか」という問いです。同じ判定者が、同じ実行ログに対して 4 回評価を行いました。答えに影響するはずのない設定のみを変更しての比較です。

20ポイントと半ポイントの違い。 二値判定モデルがより間違っているわけではありません。単に誤差が増幅されているだけです。境界線上のタスクはコイン投げと同じ確率で評価され、その結果が最終スコアでは丸ごと1つのタスクとしてカウントされてしまいます。
アンサンブル手法
もう一つの重要な要素は、判定モデルを複数回実行し、その回答を組み合わせて評価することです。

これは、判定モデルの誤差がランダムであるため機能します。各実行は「真のスコア+誤差」で構成されており、これらの誤差はお互いに一致しないため、複数の結果を組み合わせることで大部分の誤差が相殺されます。
独立した3人の評議員(juror)がいる場合、分散は1/3に減少し、5人いれば1/5になります。
実際にはこれよりわずかに少ない効果しか得られません。繰り返し実行した測定では、全体の実行スコアは約1ポイント程度変動します。
組み合わせ方が重要
単純な平均化が正解に見えるかもしれませんが、実はそうではありません。隠れた欠陥を見つける作業は探索に近いため、ある評議員が正しいファイルを見つけて「10点」と評価する一方、他の2人が「96点」を付けることがあります。そのたった1つの外れ値が平均点を30ポイントも引き下げてしまいます。

そのメジアン(中央値)は、その状況を生き残ります。しかし、たった一人の陪審員が正しかったケースを除外してしまい、さらにその二つのケースを見分けることもできません。
私たちはこれを改善しようと試みました。三つのスコアが大きく乖離している場合、それらすべての監査結果を第四のモデルに送り、どちらが正しいかを判断させることにしました。

行き止まりでした。 仲裁役となったのは別の LLM なので、実行するたびに回答が異なります。安定していたメジアンを不安定なタイブレークに置き換えた結果、最終的な数値は以前よりも大きく揺らぐようになりました。
そこで、たった一人の針(正解)を見つけるケースは諦めることにしました。いくつかの実在するバグを見逃すことになりますが、数値はそれ以上変動しなくなります。
サイドクエスト:大学での採点方法 大量の答案用紙に採点をつけるのも同じ問題です。複数の TA(ティーチング・アシスタント)が同じ答案に対して wildly different な点数を付けるのを防ぐため、各自がゼロから判断するのを禁止します。代わりに採点基準(marking scheme)を用意します。「完全な証明」は 50 点満点ですが、その中で本当に難しい一歩を正しく踏めば、たとえその先へ進めなくても 40 点が与えられる、といったルールです。
私たちも失敗モードに対して同じアプローチをとります。異なるモデルでタスクを 20 回実行し、起こりうるすべてのエラーを収集して、判定者にリストとして渡します。現在、5,692 のパターンがあります。ただし、これは既知の事象にしか対応できません。エージェントがブラウザを完全にスキップしてサイトの API を逆解析するようなケースでは、該当する項目が存在せず、判定者は再び推測することになります。
モデル比較
各タスクにはスコアが付与されます。これを「A が B に勝つ」という形式に変換する方法は三つあり、これらは順に精度が高くなります。
閾値とカウント
閾値を設定し、それより上の結果を「合格」、以下を「不合格」とみなして合計します。これがほぼすべてのベンチマークで採用されている手法です。

この方法では、スコアが示していた情報がすべて失われます。実質的には 52 と 48 は同じ結果ですが、閾値の境目によって反対側のグループに分類されてしまいます。これでは以前の問題である「増幅効果」を解消できません。
タスクごとの比較
二つのモデルを同じタスクで並べ、そのタスクでの勝者を決定し、それをカウントします。スコアについては中央に 5 ポイントの幅を持つ「引き分け帯」を設定し、わずかな差が結果としてカウントされないようにします。

各行は両モデルが挑戦した一つのタスクを表しています。このペアリングは連続的なジャッジ方式よりも前に作られたもので、判定結果は合格・不合格の二値であり、「引き分け」は両者が同じ結果を出したことを意味します。
「84 対 78」という数値比較よりも有用です。モデル間の評価は概ね一致しており、Opus が 14 の特定のタスクで先行していることがわかります(詳細は該当箇所をご参照ください)。
ペアリングによりノイズの多くも排除されます。なぜなら両モデルが同じタスクに直面するからです。
勝敗の数を数えるだけでは、誤差も切り捨ててしまいます。これは閾値の問題であり、私はこの閾値に焦点を当てました。違いは、タスクレベルでは規模がシグナルとなる一方、ここでは通常はアーティファクト(ノイズ)として扱われる点です。
これが答えを変えます。Opus は 44 のタスクで Grok を 12.3 ポイント差で下回りましたが、Grok は 28 のタスクで Opus を 26.7 ポイント差で上回りました。Grok が勝つのは特定の条件においてですが、勝敗の数を数える限りでは、Opus が 56.6対43.4 で勝利します。私たちは、ある一つのタスクで90ポイントもの差がつくことは、ほぼ例外なく評価者が異常な振る舞いをしていることを示唆すると考えています。
Elo
一度に2つのモデルしか比較できないため、20個のモデルを並列で扱うのは現実的ではありません。また、すべてのモデルがすべてのデータセットを実行したわけではありません。
そこで、各タスクを「ゲーム」として扱います。あるモデル A が B よりも5ポイント高いスコアを出せば勝ちとし、それ以内であれば引き分けとします。5という数字は、同じ評価者が同じタスクで得点を変える際の典型的な変動幅に相当するため、これより小さい差はノイズとして扱うのが妥当です。6つの設定における106のタスクからなる組み合わせは、合計1,590試合分に相当します。

上位4つのモデルは互いに56ポイント以内の差しかありません。これは非常に近いため、明確な順位付けを行うには至らない状況です。
しかし、実際に重要なのは別のギャップです。Opus 5 の「低負荷」設定と「高負荷」設定では88ポイントもの差があります。推論能力を強化することは、ベンダーを変更するよりもモデルの性能を大きく向上させるのです。
Elo は既存の評価結果を集約するものであり、評価そのものを修復するものではありません。
今後の課題
エージェントのノイズを報告してください。複数回実行し、ベスト・オブ・N(Best-of-N)の結果を示すようにしてください。
実際に評価できるノイズを排除する。評価者にはすべての証拠を与え、判決ではなくスコアを求め、3 回実行して中央値を採用する。これにより「正しさ」は保証されないが、「一貫性」は確保され、測定可能な部分となる。
そして、その評価者をデータセットの隣に公開するのだ。
これで安定した状態になれば、以下のようなチャートが得られる。

同じ 106 のタスクで、すべてのモデルを同一の評価基準で判定した結果だ。Luna xhigh は、Opus 5 よりも 2 つ少ないタスクで合格したが、そのコストはわずか 16 分の 1 である。 この比較が意味を持つのは、評価者が両者の間で動いていないからだ。
このチャート上のすべてのモデルは Browser Use Cloud で動作しており、必要に応じて自分のタスクに指し示すことも可能だ。
TL;DR
評価者だけを切り替えるだけで、同じ実行でスコアが 62.5% または 83.7% に変わる。評価者にはワークスペース全体を与え、判決ではなくスコアを求め、3 回実行して中央値を採用する。
原文を表示

A benchmark is two things: the tasks, and the judge that decides whether you succeeded. Everyone treats the judge as an afterthought, but it decides every number you publish, and a weak one makes the whole thing reward-hackable.
So we spent a long time building ours. An aligned judge has to be consistent in four ways:
- The same trace twice. Judge one finished run again and get the same score.
- Across tasks. Task 90 is held to the same bar as task 3.
- Across models and harnesses. Browser Use and Codex get graded the same way.
- At the run level. If model A beats model B, that shouldn't flip when you judge it again.
None of that is free. Start with how much of the score is even real.
What moves the score
Two things. Only one of them is the model.
Agent noise is the web and the model. A site blocks you. A wrong turn on step 12. You measure it. You can't remove it.
Judge noise is the same finished run, graded twice, coming back with two different verdicts. That one we can kill.
Agent noise
We ran GPT-5.5 five times on the same 106 tasks. Same model. Same harness. Same config.
Each run solved about 89. Never the same 89.

Every column is one task. Every row is one run. Green means it worked.
64 tasks work every time. 2 never work. The 40 in the middle are the whole problem. Same model, same task, passing and failing at random. A misread filter. A wrong click. Giving up one step early.
Run once and you score 85%. Give it five tries and 104 of the 106 tasks work at least once.
So a single run doesn't measure what a model *can* do. It measures how often it does it.
Judge noise
Now freeze the agent. We took 104 finished runs from that set, the ones we have human labels for, and changed only the model doing the grading.

Same idea: every column a task, every row a judge, all reading identical work.
45% of the benchmark depends on which judge you asked. The most lenient reports 83.7%, the strictest 62.5%. Twenty-one points, wider than the gap between most models people compare.
If the judge moves the score more than the model does, you can't tell whether a change helped.
Building the judge
One prompt with the whole trace
The first judge was a single LLM call. Whole trace in, verdict out.
Fine with one harness. Then we added more. Traces look nothing alike, run hundreds of steps, and compaction eats the part you need.
The evidence is one thing, buried

Task: extract listings under $2,000. On step 34 the agent sets the filter to $20,000. Everything after that is clean, confident, and wrong.
You can't fit that in a prompt. And one pass gives one shot at finding it.
So the judge became an agent. It gets the trace and the workspace and goes hunting, as Codex or Claude Code with read access. It can open the CSV, scroll back to step 34, check the number against the page it came from.
What the judge can see
An agent judge can only verify what it can open. Easy to break, though. Screenshots are heavy. Page text is heavy. Something truncates to fit a budget.

Cut the wrong thing and a true claim becomes unverifiable, which to the judge looks identical to an invented one. It starts calling correct answers hallucinations. Get the evidence contract right before you touch the prompt.
Why we stopped asking for pass/fail
The judge has to output something, and pass or fail is the obvious choice. It's also the wrong one, because most tasks aren't binary.
Real prompts are loose. *Check one of the top contributors.* *Open the trending puzzle.* Nobody can verify which one was right, including the judge. And agents fail partially: ask for JSON, get a CSV with all the right data in it. After 400 correct steps that isn't a zero.
One from our set:
Go to spaces and navigate to one of the recommended spaces to view. Check the profile of one of the top contributors in this space and return how many followers they have.
The agent verified 227 followers. But the recommended-spaces page was login-gated, so it used a publicly reachable Space instead.
Five of our eleven judges passed it. Six failed it. Same screenshots, same follower count. They split on whether a substituted Space counts.
So it isn't 0 or 1. Call it a 60. You can argue for 45, and that argument is the point: the score doesn't remove the judgment call, it puts it somewhere you can see.

So we ask for a number instead: what percentage of the requested outcome was delivered, correct and backed by evidence. Here is the same judge on the same traces, run four times, changing only a setting that shouldn't move the answer at all.

20 points versus half a point. The binary judge isn't more wrong, it just amplifies. Every borderline task is a coin flip, and every coin flip becomes a whole task in the final score.
Ensembles
The other lever: run the judge several times and combine the answers.

It works because judge noise is random. Every run is the real score plus some error, and those errors don't agree with each other, so putting several together cancels most of them out. For independent jurors each with variance :
Three jurors cut the spread by , five by .
In practice you get a bit less than that. Measured across repeat runs, our whole-run number moves by about a point.
How you combine them matters
Averaging seems right and isn't. Finding a buried defect is a search, so sometimes one juror digs into the right file and comes back with a 10 while the other two sit at 96. That single outlier drags the average down 30 points.

The median survives that. It also throws away the case where the lone juror was *right*, and it can't tell the two apart.
We tried to fix that. When the three scores are far apart, send all three audits to a fourth model and let it decide who's right.

Dead end. The arbitrator is another LLM, so it answers differently on different runs. We had swapped a stable median for an unstable tiebreak, and the final number moved around more than before.
So we let the lone-needle case go. We miss a few real bugs and the number stops moving.
Side quest: how universities do it
Marking a stack of exams is the same problem. Nobody wants two TAs handing the same paper wildly different grades, so they don't let each TA decide from scratch. They get a marking scheme: the full proof is 50 points, and the one genuinely hard step is worth 40 even if you only got a tenth of the way through it.
We do the same thing with failure modes. Run a task twenty times across different models, collect every way it goes wrong, and hand the judge that list. We have 5,692 of them.
It only covers what you've already seen, though. The first time an agent skips the browser entirely and reverse-engineers the site's API, there's no line for it and the judge is guessing again.
Comparing models
You have a score for every task. Three ways to turn that into "A beats B", and they get better in that order.
Threshold and count
Pick a cut, call everything above it a pass, add them up. This is what almost every benchmark reports.

It throws away everything the score just told you. A 52 and a 48 are the same run for any practical purpose, and they end up on opposite sides. You are back to the amplification problem from before.
Task by task
Line the two models up on the *same* task, decide who won that one task, then count. On scores you leave a tie band in the middle, five points wide, so small gaps don't count as a result.

Every row is one task both models attempted. This pair predates the continuous judge, so the verdicts are pass/fail and a tie means both did the same thing.
More useful than "84 versus 78": the models mostly agree, and Opus is ahead on fourteen specific tasks you can go read.
Pairing also kills most of the noise, because both models faced the same task:
Counting wins throws away the margins too. That's a threshold, and I just attacked thresholds. The difference: at the task level magnitude is signal, here it's usually artifact.
It changes answers. Opus beat Grok on 44 tasks by 12.3 points, Grok beat Opus on 28 by 26.7. On , Grok wins. On win counts, Opus takes it 56.6 to 43.4. We believe Opus: a 90-point gap on one task is nearly always the judge being strange.
Elo
Two models at a time doesn't scale to twenty, and not every model has run every dataset.
So treat each task as a game. A beats B if it scores five points higher, draw inside that. Five is roughly how much the same judge moves on the same task, so anything smaller is noise. 106 tasks across six configurations gives 1,590 matches.

The top four sit within 56 points of each other. That's close enough that I wouldn't call it a ranking.
The gap that does matter: Opus 5 at low effort is 88 points below Opus 5 at high. Turning up reasoning moves a model further than switching vendors does.
Elo aggregates the judgments you already have. It can't repair them.
Where this leaves us
Report agent noise. Run more than once and show best-of-N.
Judge noise you can actually kill. Give the judge every piece of evidence, ask for a score instead of a verdict, run it three times, take the median. That makes it consistent, not correct, and consistent is the part you can measure.
And publish the judge next to the dataset.
Once it holds still, you get charts like this:

Same 106 tasks, every model judged the same way. Luna xhigh lands two tasks behind Opus 5 at a sixteenth of the cost. That comparison only means something because the judge didn't move between them.
Every model on that chart runs on Browser Use Cloud, if you want to point them at your own tasks.
TL;DR
Swap only the judge and the same runs score 62.5% or 83.7%. Give it the whole workspace, ask for a score instead of a verdict, run it three times, take the median.
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み