Arena、モデル評価に自動評価「AutoEval」を導入し即日スコア表示へ
本文の状態
日本語全文を表示中
詳細モードで約8分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Arena Blog
Arena はモデル発表直後に即座にスコアを提供する「AutoEval」機能をリーダーボードへ導入し、人間の投票を待つ従来の評価プロセスを報酬モデルによる自動投票で補完することで、数時間以内の迅速なフィードバックを実現した。
AI深層分析を開く2026年8月18日 06:17
AI深層分析
キーポイント
即時信号の提供機能
AutoEval はモデルがリリースされた直後にスコアを提供し、スコアは「AutoEval」として明確にラベル付けされる。十分な人間の投票が集まり検証が行われるまでこの状態が続く。
報酬モデルによる自動投票
人間の嗜好を捉えた報酬モデル(RM)を訓練し、これを用いて自動的に投票を行うことで、人間の投票に依存しない評価ルートを確立した。
迅速なフィードバックと柔軟性
報酬モデルに基づく投票によりランキングが数時間以内(1 時間未満)で得られ、また RM が対象とするプロンプトを制御することで特定ドメインへの評価も可能になる。
マルチモーダル対応の拡大
このシステムはテキスト、ビジョン、画像生成、コーディングの各領域に適用され、人間の投票だけでは迅速なスケーリングが困難な領域を補完する役割を果たす。
迅速かつ柔軟なフィードバックの実現
RMベースの投票により評価結果が1時間以内に得られ、特定ドメインへのターゲット評価も可能になる。
重要な引用
AutoEval provides a Day-1 signal as soon as a model launches.
RM-based voting gives a ranking in under an hour, not days.
Arena AutoEval works across text, vision, image generation, and coding domains.
Our text reward model predicts human preferences 8–10% more accurately than frontier LLM judges (Gemini-3-flash/pro, GPT-5).
編集コメントを表示
編集コメント
評価のスピードと信頼性のバランスを改善する実用的なアプローチであり、特に新モデルが頻繁に登場する現在の市場環境においてその価値は大きい。人間の判断と機械的予測の融合という手法は、今後の AI ベンチマークの標準的な形の一つとして定着する可能性がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AutoEval は、モデルがリリースされた瞬間に最初の評価結果(Day-1 シグナル)を提供します。スコアは明確に「AutoEval」としてラベル付けされ、十分な数の人間による投票が集まり検証されるまで更新されます。これにより、コミュニティは各タスクに適した最良のモデルをより早く特定できるようになります。

##
投票を待つ代わりに、人間の好みを捉えた報酬モデル(Reward Model: RM)を訓練し、それを使って自動的に投票を行います。ライブの人間による投票と、この RM が生成する代理投票を組み合わせることで、以下のような利点を実現します。
- 迅速なフィードバック: RM に基づく投票では、ランキング結果が数日ではなく1 時間以内に得られます。
- 高い柔軟性: RM が評価対象とするプロンプトを制御できるため、特定のドメインに特化した評価が可能です。
Arena の AutoEval は、テキスト、ビジョン、画像生成、コーディングの各領域で機能し、人間による投票だけでは迅速なスケーリングが難しい分野において、迅速かつターゲットを絞った自動化によって人間の評価を補完します。
##
AutoEval はデータから始まります。Arena の大規模な人間好みのデータセット(ライブの投票からの数百万件のペア比較から構成される)を用いて報酬モデルを訓練し、人間の判断を近似できるように学習させます。このモデルの中核となるのはポイントワイズ報酬モデルで、個々のプロンプトとレスポンスの組み合わせをスカラー値のスコアに変換します。ペア比較データを用いて訓練を行うことで、得られるスコアの順位付けが人間の評価基準と一貫して整合するように設計されています。
膨大な実世界の選好データを学習することは、大きな成果をもたらします。当社のテキスト報酬モデルは、最先端の LLM 判定器(Gemini-3-flash/pro, GPT-5)と比較して、人間の選好を8–10% より正確に予測できます。Arena で扱われるプロンプトには難易度が高く複雑なものも多く、競合するモデルも強力なため、すべての側面で明確に優位な回答が一つだけ存在しないケースが頻繁にあります。そのような場合、LLM 判定器は最終的な判断に苦慮する一方、当社の報酬モデルは人間が実際に何を好むかを捉え続けています。
ライブサービス向けの報酬モデルの訓練は、一度きりの学習問題ではありません。対象となるタスク、データ、そして評価されるモデル自体が絶えず変化し続けます。私たちはその過程でいくつかの課題に直面しました:
人間の嗜好は常に変化します。ユーザーが「良い回答」と考える基準は、期待値や対話スタイル、利用ケースの変化に伴って移り変わります。過去のフィードバックには広範なデータと安定性がありますが、最新のフィードバックこそがこうした変化を反映しています。効果的な学習には、膨大な過去のフィードバック信号に基づきつつ、時間とともに drifting する最新トレンドに適応できる仕組みが必要です。
モデルの最前線は常に進化しています。モデル性能が向上するにつれ、明らかな失敗事例は減り、品質の違いも微妙なものになっていきます。リワードモデルは新たな能力に対応し、より細かな違いを敏感に捉え続ける必要があります。
嗜好ラベルには本質的な曖昧さが伴います。人間の投票結果が常に明確な決着をつけるわけではありません。ユーザーによって重視する要素が異なり、同点(ties)の場合も、「どちらも同等に優れた回答」である場合と「どちらも同等に劣る回答」である場合があります。単純な勝ち負けとして扱うとこの情報が失われるため、同点や較正(calibration)は明示的にモデル化する必要があります。
多ターンでのフィードバックは文脈依存性が高いです。最終的な投票結果が、過去の判断や文脈管理、中間結果の影響を受けている場合があり、どの行動がその嗜好を生んだのかを特定するのは困難です。
複雑なタスクにはハイブリッドな信号が必要です。コーディングのような難易度の高い領域では、単一の指標で回答を評価することはできません。コードの正しさや技術的な品質、最終的なインターフェースの外観、そして全体的なユーザー体験など、あらゆる側面を考慮する必要があります。
これらの課題を慎重に解決した結果、当社の報酬モデルは実用において信頼できる精度を備えています。
##
AutoEval は人間の投票の代わりにこの報酬モデルを用い、Arena のランキング計算手法をそのまま適用して最終スコアを算出します。具体的には、対象となるすべてのモデルに対して以下の 4 つの手順を実行します。
- ライブ評価サンプリング — ライブ評価から実際のプロンプトと回答のペア (p, r_A, r_B) を抽出します。
- モデル生成 — 同じプロンプトに対し、対象モデルに回答 B を再生成させます。
- 報酬モデルによるペア投票 — 報酬モデルが各回答を採点し、スコアの差に対してソフトマックス関数を適用して、確率的な投票結果を取得します。
- AutoEval スコア算出 — 自動評価器による対象モデルの投票と、他のモデルに対する人間の投票を組み合わせて、Arena スコアを生成します。
では、AutoEval は実際の評価にどれほど近いのでしょうか?これを検証するため、2026 年 4 月末までのデータを用いて報酬モデルを訓練し、その後にリリースされたすべての新モデルについて AutoEval スコアを計算しました。これをライブ評価スコアと比較した結果(以下の図参照)から、AutoEval のスコアがライブ評価と高い相関を持つことが明らかになりました。具体的にはランキングの相関係数は 0.98 を超え、ほとんどの AutoEval スコアはライブ評価の信頼区間内に収まっています。

また、40 以上のモデルを対象とした直接比較において、AutoEval がより優れたモデルを正しく選択できる能力も検証しました。真の性能差が 10 ポイントを超える場合、AutoEval の精度は 90% を超え、15 ポイントを超えると 100% に達します。一方、モデル間の差が 5 ポイント以内の場合は、信頼区間が重なるため、ライブ評価では統計的に有意な差を判定できません。

この高い相関関係は、AutoEval が重要なモデル選定を担うのに十分な精度を持っていることを示しています。例えば、リソースを投入して完全なライブ評価を行う前に、有望なモデルを大規模な候補群から絞り込む際にも活用できます。
##
AutoEval はテキスト専用ではありません。同様の手法を、テキスト、ビジョン、画像生成、コードの各アレーナに適用しています。
画像生成を例に挙げると、300 万組以上の選好ペアを用いてテキストから画像への報酬モデルを訓練しました。この巨大な選好データセットは、実際のユーザーが求めるリアルな画像生成プロンプトに基づいて構築されたものであり、これにより訓練された報酬モデルは、自社のデータ分布において優れた性能を発揮するだけでなく、公開ベンチマークでもトップクラスの成果を収めています。
Meta が 2026 年に発表した多モーダル報酬ベンチマーク「MMRB2」での評価では、Arena-RM は他のポイントワイズテキストから画像への報酬モデルと比較して最先端の性能を示し、2 位(HPSV3、UnifiedReward、PickScore、VQAScore など)を 9 ポイント以上引き離しました。
![図注:Arena leaderboards のスコア比較グラフ]

##
Arena リーダーボードは公共財であり、AutoEval はその価値をモデルリリースの直ちに提供します。すでに Text、Vision、Image、Code Arena 全体で動作しており、新しいモデルを数日ではなく数時間でランク付けできる能力を提供しています。また、コミュニティが関心を持つ分野に焦点を当てています。
報酬モデリングにおけるオープンな課題に取り組み続ける中で、このアプローチはさらに多くのモダリティに拡張され、人間の選好をより正確に捉えるための追加信号も取り込まれていきます。
原文を表示
AutoEval provides a Day-1 signal as soon as a model launches. Scores are clearly labeled “AutoEval,” then updated once enough human votes arrive to validate them. This helps the community identify the best models for their tasks, sooner.

##
Instead of waiting on votes, we train a Reward Model (RM) that captures human preference, then use it to cast votes automatically. By integrating live human votes with these RM-generated proxy votes, we unlock:
- Rapid feedback: RM-based voting gives a ranking in under an hour, not days.
- High flexibility: Because we control the prompts the RM votes on, we can target evaluation on specific domains.
Arena AutoEval works across text, vision, image generation, and coding domains, supplementing our human evaluation with rapid, targeted automation where just human voting cannot scale quickly enough.
##
AutoEval starts with data. We train a reward model on Arena’s large-scale human preference dataset, consisting of millions of pairwise comparisons from live votes, so that it learns to approximate human judgments. At its core, the model is a pointwise reward model: it maps each individual prompt-response pair to a scalar score. We train it on pairwise comparison data so that the resulting score rankings consistently align with human preferences.
Learning from massive real preference data pays off. Our text reward model predicts human preferences 8–10% more accurately than frontier LLM judges (Gemini-3-flash/pro, GPT-5). Many prompts on Arena are hard and complex, and the competing models are strong, so there’s often no single response that clearly dominates in all aspects. In those cases, we found that LLM judges struggle to make the final call, while our reward model still captures what humans actually prefer.
Training a reward model for a live service is not a one-time learning problem. The target, the data, and the models being evaluated all continue to change. We encountered several challenges:
- Human preferences are a moving target. What users consider a good response changes as expectations, interaction styles, and use cases evolve. Historical feedback provides breadth and stability, while recent feedback reflects these shifts. Effective training must grounded in the vast historical feedback signal while adapting to the recent trend drifting over time.
- The model frontier keeps moving. As models improve, obvious failures become less common and quality differences become more subtle. Reward models must keep pace with emerging capabilities and remain sensitive to increasingly fine-grained distinctions.
- Preference labels contain genuine ambiguity. Human votes are not always decisive: different users value different qualities, and ties may represent either two similarly strong responses or two similarly weak ones. Treating every comparison as a simple win or loss loses this information, so ties and calibration need to be modeled explicitly.
- Multi-turn feedback is context-dependent. A final vote may reflect earlier decisions, context management, or intermediate results, making it difficult to isolate which behavior drove the preference.
- Complex tasks require hybrid signals. For difficult domains like coding, an answer can’t be judged on a single metric. Every aspect—including code correctness, technical quality, the look of the final interface, and the overall user experience—has to be taken into account.
With these challenges carefully handled, our reward model is accurate and reliable enough to be useful in practice.
##
AutoEval uses the reward model in place of human votes, then follows the exact same Arena ranking methodology to compute the final arena score. Specifically, for every target model it runs four steps:
- Live Eval Sampling — sample real prompt-and-response pairs (p,rA,rB)(p, r_A, r_B) from live evaluation.
- Model Generation — regenerate response B from the target model on the same prompt.
- RM Pairwise Vote — the reward model scores each response, applying a softmax to the score differences to obtain a soft probability vote.
- AutoEval Score — combine the autorater’s votes for the target model with the live human votes for other models to produce the Arena Score.
So, how close is AutoEval to the real thing? To test this, we trained a reward model using data up to the end of April 2026 and evaluated AutoEval scores for all new models released after that date. When comparing these to live evaluation scores, the results (shown in the figure below) demonstrate that AutoEval scores are highly correlated with live evals. Specifically, the ranking correlation exceeds 0.98, and most AutoEval scores fall well within the confidence intervals of the live scores.

We also tested AutoEval’s ability to correctly choose the better model in head-to-head comparisons across >40 test models. AutoEval achieves >90% accuracy when the true performance gap exceeds 10 points, and 100% accuracy beyond 15 points. When models are within 5 points of each other, they are statistically indistinguishable under live evaluation due to overlapping confidence intervals.

This high correlation demonstrates that AutoEval is accurate enough to drive critical model selection decisions. For example, it can be used to shortlist the most promising models from a large pool before committing resources to a full live evaluation.
##
AutoEval isn’t text-only. We apply the same methodology across modalities — text, vision, image generation, and code arenas.
Taking image generation as an example, we trained a text-to-image reward model on over 3 million preference pairs. By training on this massive preference dataset—built from realistic image generation prompts requested by real-world users—our reward model not only excels on our own distribution but also achieves top-tier performance on public benchmarks. Evaluated on MMRB2, Meta’s 2026 multimodal reward benchmark, Arena-RM achieves state-of-the-art compared with other pointwise text-to-image reward models, more than 9 points ahead of the runner-up (HPSV3, UnifiedReward, PickScore, VQAScore and others).

##
Arena leaderboards remain a public good and AutoEval makes that value available sooner: right at a model's release. It already works across Text, Vision, Image, and Code Arena, providing the ability to rank new models in hours rather than days, while focusing on the domains the community cares about. As we continue working through open challenges in reward modeling, we’ll be extending this approach to even more modalities and incorporating more signals to better capture human preferences.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み