Arena、ファクトアレーナにおける事実性の検証と評価手法を公開
本文の状態
日本語全文を表示中
詳細モードで約14分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Arena Blog
Arena AI は、LLM の出力が事実と一致する度合いを評価する「Factuality in the Arena」の指標を発表し、ハルシネーション対策におけるベンチマークの重要性を強調している。
AI深層分析を開く2026年8月1日 14:15
AI深層分析
キーポイント
Factuality in the Arena の発表
Arena AI が運営する評価プラットフォームにおいて、LLM の出力が事実と一致する度合いを測る新たな指標「Factuality in the Arena」を導入したことを発表した。
ハルシネーション対策の重要性
同社は、生成 AI が誤った情報を生成するハルシネーション現象が実用化における最大の課題の一つであり、これを定量的に評価する必要性を指摘している。
ベンチマークとしての役割
この指標は、開発者がモデルの信頼性を比較・検証するための新たな基準として機能し、業界全体の品質向上に寄与すると期待されている。
事実性の評価指標の導入
ユーザーが直面する永続的な課題である「事実性」に基づき、モデルの回答の正しさを評価する新ランキングを開始した。
検索エージェントによる自動検証
ウェブで検証可能な原子論を抽出し、検索エージェントシステムが補正された真偽確率を与えることで事実性を判定する。
重要な引用
Factuality in the Arena
Factuality remains one of the most persistent questions users face when using AI models.
We are adding factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality, two complementary signals that tell only a partial story in isolation.
Most top models (by human preference score) have decreasing score as factuality weight is increased.
編集コメントを表示
編集コメント
生成 AI の信頼性を高めるためには、能力だけでなく事実の正確さを評価する指標が不可欠である。この発表は、業界がハルシネーション問題に真剣に取り組むための重要な一歩となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
「Arena における事実性」

研究
製品

Arena チーム
- 公開日:2026 年 7 月 14 日
- 最終更新日:2026 年 7 月 31 日
- 読了時間:8 分
AI モデルの利用において、ユーザーが直面する最も根深い疑問の一つに「事実性」があります。本日、私たちは人間による評価だけでなく、回答の事実的な正確さにも基づいてモデルをランク付けするリーダーボードを発表します。
人間の嗜好は、Arena のモデルランキングの根幹として当初から重視されてきました。しかし、モデル品質の一部の側面は人間が評価するのが困難です。その一つが「事実性」、つまりモデルの回答に含まれる主張の正しさです。主張を手動で検証するのは時間がかかり手間がかかるため、事実性は必ずしも人間の投票結果に反映されるとは限りません。そこで本日、Arena に「事実性」を追加します。これは、人間による評価と事実性の両方を加重結合した新しいランキングです。これら二つの指標は、単独では物語の一部しか語らない相補的な信号なのです。
当初は Text Arena と Search Arena に事実性を導入します。これはリーダーボード UI で選択できる、デフォルトではないトグルとして機能します。
まず、事実性監査のためにランダムにバトル事例をサンプリングします。次に、各モデルの回答から原子論(アトミック・クレーム)を抽出し、ウェブ上で検証可能なものだけをフィルタリングします。その後、検索エージェントによる検証システムが、各原子論に対して較正済みの真偽確率を算出します。この較正は、高品質な事実性注釈データセットに基づいて生確率を後処理することで実現しています。
最後に、各回答の原子論の平均正解確率を比較します。平均的な真実性がより高いモデルがバトルに勝利し、その差(マージン)が大きいほど、勝敗は明確になります。
ラベル付けされた「事実性バトル」の結果を用いて、人間の選好と事実性の両方を反映した統合型リーダーボードを構築しました。これには複合ブラッドリー・テリーモデルを採用しています。このモデルにより、事実性ラベルと純粋な人間選好ラベルの重み付けを自由に調整できます。デフォルトの設定では、事実性の重みを 25% に設定しています。つまり、事実性に欠けるモデルはリーダーボードの上位に上がることはほぼ不可能です。逆に、誤った主張を一切出さないものの、人間にとって実用的でないモデルも同様に上位に食い込むことはできません。
##
これらのランキングを算出するために、私たちは実世界での会話において LLM が行った 200 万件以上の主張にラベル付けを行いました。その内訳は、Text Arena から 130 万件以上、Search Arena から 70 万件超です。これらは約 13 万 5,000 回の Text Arena の対戦と、4 万回の Search Arena の対戦にまたがっています。
Text Arena の対戦では、回答の少なくとも片方に主張が含まれているケースが 76% で見られ、Search Arena ではその割合は 88% に達しました。1 回の回答あたりの平均主張数は、Text Arena で 5 件、Search Arena では約 10 件です。真実である主張の割合(マージナル・トゥルークレームレート)は、Text Arena で 87%、Search Arena で 89% でした。
このデータ収集を支援しているのが、Arena の AutoModality クラスファイヤーです。これは、ウェブアクセスが必要とされる可能性が高いプロンプトを自動的に Search Arena に振り分ける機能を持っています。
各モデルのスコアと事実性重み付け(ファクトアリティ・ウェイト)の関係を図示すると興味深い傾向が見えてきます。重み付けが 100% の場合はスコアが事実性のみで決まり、0% の場合は人間の好みに基づいたスコアとなります。
モデルによって傾きは異なりますが、全体的なトレンドは共通しています。人間による評価スコアが高い上位モデルの多くは、事実性の重み付けを高めるにつれてスコアが低下する傾向にあります。
Text Arena と Search Arena の両方で同様の結果が見られました。OpenAI の GPT モデルと SpaceXAI の Grok モデルは、事実性の重み付けが増えるにつれてスコアが上がったり、横ばいを維持したりしています。一方、Anthropic の Claude や Google の Gemini は、重み付けの増加とともにスコアが低下するか、横ばいになります。
最後に、Meta の muse-spark-1.1 と Baidu の ernie-5.1 は、事実性の重み付けが高まるにつれてスコアが大きく下落しました。
上記のグラフは、4 つのプロバイダーごとにモデルを分類し、「事実性の重み」と「スコア」の関係を示したものです。これにより、各プロバイダーのモデル群における傾向が可視化されています。
OpenAI のモデルは概ねスコアが上昇する一方、Anthropic のモデルは低下するか横ばいとなっています。特に注目すべきは、最新バージョン 4.8 がリリースされる前の Claude Opus で最も事実性が高かったのは 4.5 バージョン(2025 年 11 月リリース)だったという点です。人間による評価では Anthropic のモデルが全体的に好まれる傾向がありますが、OpenAI の最新モデルは事実性がより高いことがわかります。
SpaceXAI の Grok-4-0709 は堅実な事実性を持つモデルでしたが、その後の Grok-4.1 シリーズでは事実性が大きく後退しました。その後、Grok-4.20、Grok-4.3、Grok-4.5 といった新バージョンで事実性は改善されましたが、Grok-4.3 は人間からの評価が大幅に低下しています。
Google のモデルは、OpenAI や SpaceXAI のモデルと比べて事実性スコアのばらつきが小さいのが特徴です。特に Gemini-2.5 シリーズは現在でも最も事実性が高いモデル群ですが、時とともに事実性が低下する傾向が見られます。
検索ツールを利用できる「Search Arena」でのテストでも、これらの傾向は概ね同様でした。
一方、オープンソースモデルの多くは、事実性の重みが高くなるにつれてスコアが低下します。ただし、mistral-medium-3.5 と Tencent の huyuan-hy3-preview は例外です。
ほぼすべてのプロバイダーで人間の評価は時間とともに上昇していますが、純粋な事実性スコアを主力モデルのリリース日と対比してプロットすると、OpenAI が唯一、同時に着実に事実性を向上させていることがわかります。SpaceX AI も最近のリリースで改善しており、Search Arena においても同様の傾向が見られます。特筆すべきは、Google と Anthropic のモデルが過去 1 年間を通じて比較的同程度の事実性レベルを維持している点です。特に検索ツールへのアクセス権限がある場合、その傾向は顕著です。Meta は「muse-spark」から「muse-spark-1.1」にかけて事実性に大きな改善を見せましたが、依然として主要なクローズドおよびオープンソースのラボには及んでいません。
また、人間の評価スコアと事実性スコアの関係を観察することもできます。両者には弱い正の相関が見られますが、本質的にはほぼ直交する指標です。このため、事実性を示す信号は評価における不可欠なサブコンポーネントとして追加する必要があります。例えば、モデルが情報提供を拒否した場合、その回答は完全に事実正しいものとなりますが、人間の評価という観点では低く評価されがちです。一方で、多くの事実に基づいた長い回答は一見包括的で、人間の投票で勝つように見えるかもしれませんが、そこに事実誤認が含まれている可能性もあります。
さまざまな業界カテゴリーにおける主張の正しさの平均値を分析すると、モデルの事実性の業界間にはわずかながらも直感的に理解できる差異があることがわかりました。数学的なタスクではモデルは概ね正確であり、ソフトウェア、医療、科学分野のタスクでも妥当な事実性が保たれています。一方、最も事実性に欠けるのは法律および政府関連のタスクでした。
以下では、抽出された主張のサンプルを確認できるビューアをご紹介します。
スタイルを重視すること自体が誤りではありません。チャットにおいて人間と効果的にコミュニケーションを取ることは不可欠であり、書式や文章構成は評価における重要な要素です。より懸念されるのは、「スタイリッシュ」な回答が、モデルの明らかな誤りを隠蔽してしまう点です。
今回の新しい事実性リーダーボードでは、この懸念に正面から取り組んでいます。望ましくない行動をスタイルという代理指標で推測するのではなく、そのような行動そのものを直接測定・検出することを目指しました。もちろん、事実性は回答の正しさを表す概念の一つに過ぎませんが、以下の理由から特に重要かつ優先すべき指標です。(1) 人間は専門外の質問にもモデルに頼るケースが増えていること、(2) モデル自体が事実確認において人間よりも優れているケースが多くなっていることです。
以下では、主張を公平に監査し、原則に基づいたリーダーボード計算に取り込むために用いられた手法についてより詳細に説明します。**
複合ブラッドリー・テリーモデル**
人間による評価や事実性のみを頼るのではなく、両方の目的を統合するために合成された Bradley-Terry モデルを採用しています。事実性は回答が事実に基づいているかを測る指標ですが、それだけではユーザーの問いに本当に答えているかが判断できません。そこで人間の嗜好性を組み込むことで、その信号を保持し、事実性に基づきつつもユーザーの意図に応える回答を促進します。
合成損失関数
リーダーボード上のモデル数 M を持つ rating ベクトル θ∈R^M を、2 つの成分と小さなリッジ正則化項からなる合成損失関数を最小化することで推定します。
- L_human(θ) は、各対戦における人間による投票結果に基づいて計算される標準的な Bradley-Terry の負の対数尤度です。
- L_fact(θ) は、各対戦で導出された事実性に基づくソフトな評価結果を用いた同様の負の対数尤度です。
- 注記:w_human + w_fact = 1 です。ここでは事実性の重みを単に w と表記し、人間の嗜好性の重みは 1 - w とします。
推定された rating θ は、人間による投票と事実性評価の同時分布を最もよく説明する値であり、その際 w の重みが適用されます。
各対戦ごとの事実性評価結果
事実性項が適用される対戦においては、事実性評価結果は (0,1) 範囲のソフトな値として得られます。これは、各側ごとの真実確率の差に対してシグモイド関数を適用することで算出されます。
avg_truth_prob_x は、戦いの片側 x から出力されたすべての検証済み主張について集約された真偽確率の平均値です。
TT は固定温度パラメータであり、事実性投票の分布の分散が人間の投票の分布の分散と一致するように設定されています。
どの戦いが事実性の項に組み込まれるか
両側のいずれかがウェブ上で検証可能な主張を出力した場合、その戦いは L_fact に寄与する資格があります。Text Arena の戦いの約 76% が関連する検証可能な主張を含んでいることが分かっています。
各ケースごとの処理方針は以下の通りです。
- 両側が検証可能な主張を出力する場合
標準的なソフトな結果として、σ((avg_truth_prob_a - avg_truth_prob_b) / T) を適用します。
- 片側のみが検証可能な主張を出力し、その主張の集計結果が真実である場合
L_fact において引き分け(0.5)とみなします。主張を出さなかった側は、主張をしなかったことに対してペナルティを受けません。
- 片側のみが検証可能な主張を出力し、その主張の集計結果が偽りである場合
誤った主張を出力した側の負けとなります。主張を控えたことは、誤った回答に対する信頼できる代替手段として扱われます。
- 両側とも検証可能な主張を出力しない場合
その戦いは L_fact から完全に除外されますが、L_human には通常通り寄与し続けます。
誤った主張を発信することは、たとえ対立側が沈黙していたとしても罰せられます。一方、沈黙(Abstention)自体は決してペナルティの対象にはなりません。我々は LhumanL_ ext{human} がこのシグナルを適切に重み付けして補正すると期待しています。複合システムは本質的に、純粋な沈黙を LhumanL_ ext{human} によって規律づける仕組みを持っています。事実性(Factuality)の軸は、自信満々に行われる誤った記述がもたらす害を捉えるためにのみ設けられています。
注意すべきは、我々が関心を持つのは「検証可能な主張」のみである点です。これはウェブ上で検証可能で、曖昧さがなく、客観的な主張を指します。これに該当しない主張は抽出されず、スキップされます。
具体例:
主張 | 事実確認可能か?
現在の大統領は良い仕事をしている。 | いいえ
米国の現大統領は82歳である。 | はい
PyTorch において、BCEWithLogitLoss のデフォルトの reduction は 'none' です。 | はい
Verilog において、disable ステートメントを使って関数を無効化することはできません。 | はい
校正データが与えられた場合の適合的カバレッジの条件付き分布は Beta(l, n+1−l) に従い、ここで l = ⌊(n+1)α⌋ です。 | はい
適合的予測(Conformal prediction)は、新しくクールな統計学の研究分野である。 | いいえ
カタラはアングではなくズコといるべきだった。 | いいえ
ボブは UFO に強い関心を持つ 8 歳の少年で、カンザス州のどこかの町に生まれた。 | いいえ
信頼区間(Confidence intervals)
モデルごとの評価における信頼区間は、過去のリーダーボードと同様にサンドイッチ分散推定量を用いて計算されます。具体的には、θ^=argminθL(θ とおくと:
ここで、H = ∑k wk・Hk は合成ヘッセ行列(Hk は人間の選好に基づく対戦結果のヘッセ行列、あるいは事実性の対戦結果のヘッセ行列のいずれか)を、B = ∑k (wk²/Nk)・Σk は合成勾配共分散行列を表します。また、Nk は各対戦における総バトル数を示します。
リーダーボードに報告されている信頼区間は、この推定量から導出された 95% 区間であり、モデル m に対しては (θ̂m ± 1.96√Var(θ̂)mm) という形で計算されます。これは標準的な Elo ランキング空間に変換された値です。
信頼区間の幅は重み付け w の変化に応じて変動します。特に、事実性のデータが人間の投票データよりも少ない場合、w を大きくすると区間が広がる傾向があります。
Share this Article
原文を表示

Research
Product

Arena
Team
- Published on 14 Jul 2026
- Last updated on 31 Jul 2026
- Reading Time: 8 minutes
Factuality remains one of the most persistent questions users face when using AI models. Today, we are launching a leaderboard that ranks models not only by human preference, but also by the factual accuracy of their responses.
Human preference has been at the core of Arena’s model rankings from the beginning. But some aspects of model quality are difficult for humans to evaluate. One of them is factuality—the correctness of claims in a model’s response. Manually fact-checking claims is slow and labor-intensive, so factuality may not always be reflected in human preference votes. That is why today we are adding factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality, two complementary signals that tell only a partial story in isolation.
We are initially adding factuality into our Text and Search Arenas. This will start as a non-default toggle which can be selected in our leaderboard UI.
##
First, we randomly sample battles for a factuality audit. Then, we extract the atomic claims from each model’s response. We filter for claims that are reasonably web-verifiable. Each atomic claim is then verified by a system of search agents that gives calibrated truth probabilities. Calibration is achieved by post-processing raw probabilities based on a high-quality dataset of verified factuality annotations. We then compare the average probability of claim correctness between each response. The model with higher average claim truthfulness wins the battle; the higher the margin the stronger the win.
Using the labeled “factuality battles” we construct a unified leaderboard that reflects both human preference and factuality with our composite Bradley-Terry model. The model gives us control over the weight between factuality labels and pure human preference labels. Our default factuality weight is 25%. This means models which are not factual will be essentially unable to top the leaderboard. Likewise, models that never produce false claims, but are otherwise unhelpful to humans, will also be unable to top the leaderboard.
##
To power these rankings, we’ve labeled over 2 million claims made by LLMs in real-world conversations, 1.3+ million from Text Arena, and 700k+ from Search Arena. These span roughly 130k Text Arena battles and 40k Search Arena battles. In Text Arena battles, a claim was found in at least one response 76% of the time; in Search Arena 88% of the time. In Text Arena, models averaged 5 claims per response; in Search Arena models averaged nearly 10. The marginal true claim rate in Text Arena was 87%, and 89% in Search Arena. This is assisted by Arena’s AutoModality classifier, which automatically routes prompts that likely need web access to Search Arena.
We can plot each model’s score against the factuality weight. Weight 100% means that the score is purely factuality-based, 0% means it is purely human-preference-based. Different models have different slope. Most top models (by human preference score) have decreasing score as factuality weight is increased. On both Text and Search Arena we see similar trends: OpenAI's GPT and SpaceXAI's Grok models go up or stay flat as factuality weight increases. Anthropic's Claude and Google's Gemini models go down or stay flat. Finally, Meta's muse-spark-1.1 and Baidu's ernie-5.1 drop heavily.
Above we show factuality weight vs. score for all models broken down by four providers to visualize patterns in each provider's model suite. We see most OpenAI models move up in score while Anthropic models decrease or stay the same. In particular, the previous most factual Claude Opus version before the newest 4.8 release was 4.5—released back in November 2025. We find that while Anthropic models are generally more preferred by humans, OpenAI's newest models are generally more factual.
SpaceXAI's Grok-4-0709 was a solidly factual model; however, the subsequent Grok-4.1 series heavily regressed in factuality. The newer Grok-4.20, Grok-4.3, and Grok-4.5 versions have improved their factuality, though Grok-4.3 is far less preferred by humans.
Google models have less variance in factuality scores OpenAI and SpaceXAI models. Notably, the Gemini-2.5 model series still remain the most factual to this day—Gemini seems to be getting less factual over time.
These trends remain similar when models have access to search tools in the Search Arena.
We find most open source models decline in score under higher factuality weight, with the exceptions of mistral-medium-3.5 and Tencent's huyuan-hy3-preview.
While human preference across nearly all providers has been climbing over time, plotting pure factuality score against flagship model release date shows OpenAI is the only provider steadily improving factuality at the same time. SpacexAI has also been improving with their recent releases; this is similarly true in Search Arena. Notably, Google and Anthropic models have stayed at relatively similar levels of factuality across the past year, especially when given access to search tools. Meta showed a large improvement in factuality between muse-spark and muse-spark-1.1 but still remains behind other top closed and open source labs.
We can also observe the relationship between human preference score and factuality score. We notice a weak positive correlation. Ultimately, the signals are largely orthogonal—this is why adding a factuality signal is a necessary subcomponent of evaluation. For example, a model response is perfectly factual if the model refuses to provide an informative response; such responses perform poorly in terms of human preference. At the same time, a long response with many facts may seem comprehensive and win a human preference vote, but it may also introduce factual errors.
Looking at average claim correctness for various industry categories, we find slight, but intuitive, variance in model factuality in different industries. Models are largely factual in mathematical tasks. They are also reasonably factual in Software, Medical, and Scientific tasks. The least factual industry area was Legal and Government tasks.
##
Below we provide a viewer to show a sample of extracted claims.
##
Rewarding style is not an incorrect thing to do. Communicating effectively with humans is essential in chat, so formatting and prose are important factors in evaluation. The deeper worry is that a more “stylish” response will obfuscate undeniable errors in a model's response.
With the new factuality leaderboard, we address this concern head-on: rather than using style as a proxy for such undesirable behavior, we attempt to measure and detect the existence of such behavior directly. Factuality is, of course, only one notion of response correctness, but it is a particularly important and appropriate one to prioritize because: (1) humans are increasingly trusting models to answer questions outside their area of knowledge; (2) increasingly models themselves are better at checking facts than humans.
##
Below we provide a more detailed explanation of the methods used to fairly audit claims and ingest them into a principled leaderboard calculation.**
Composite Bradley-Terry model**
Rather than relying on either human preference or factuality alone, we use the composite Bradley-Terry model to combine both objectives. While factuality measures whether a response is factually sound, it alone does not capture whether it addresses what the user actually asked for. The human preference component preserves that signal, promoting responses that are both factually grounded and responsive to user intent.
Composite loss
We fit one rating vector θ∈RM(M\theta \in \mathbb{R}^M (M = number of models on the leaderboard) by minimizing a composite loss with two components plus a small ridge penalty:
- Lhuman(θ)L_\text{human}(\theta) is a standard Bradley-Terry negative log-likelihood evaluated on per-battle human-vote outcomes.
- Lfact(θ)L_\text{fact}(\theta) is a standard BT negative log-likelihood evaluated on per-battle factuality-derived soft outcomes.
- Note: whuman+wfact=1w_\text{human} + w_\text{fact} = 1. We will most often write the factuality weight as just w;whuman=1−ww; w_\text{human} = 1 - w.
The fitted rating θ\theta is the rating that best explains the joint distribution of human votes and factuality outcomes simultaneously, weighted by ww.
Factuality outcome per battle
For battles eligible for the factuality term, the factuality outcome is a soft outcome in (0,1)(0,1) produced by applying a sigmoid to the per-side aggregate truth-probability gap:
- avg_truth_probx\text{avg\_truth\_prob}_x is the mean of the per-claim aggregated truth probabilities across all verified claims emitted by side xx of the battle.
- TT is a fixed temperature chosen so that the variance of the factuality-vote distribution matches the variance of the human-vote distribution.
**
Which battles enter the factuality term**
A battle is eligible to contribute to LfactL_\text{fact} whenever at least one side emits a web verifiable claim. We find that roughly ~76% of Text Arena battles have relevant verifiable claims. The handling per case:
Case
Treatment
Both sides emit verifiable claims
Standard soft outcome via: σ(avg_truth_proba−avg_truth_probbT)\sigma\left(\frac{\text{avg\_truth\_prob}_a - \text{avg\_truth\_prob}_b}{T}\right)
Exactly one side emits verifiable claims, and that side's claims are truthful in aggregate
Tie (0.5) under LfactL_{\text{fact}}. The abstaining side is not penalized for not making claims.
Exactly one side emits verifiable claims, and that side's claims are false in aggregate
Loss for the side that emitted false claims. Abstention is treated as a credible alternative to a false response.
Neither side emits verifiable claims
Battle is dropped from LfactL_{\text{fact}} entirely; it still contributes normally to LhumanL_{\text{human}}.
Abstention is never penalized, but emitting false claims is — even when the opposing side abstained. We expect LhumanL_\text{human} to counterweight this signal correctly; the composite system inherently disciplines pure abstention via LhumanL_\text{human}. The factuality axis is purely meant to capture the harm of confident misstatements.
Notice we only care about *verifiable claims*. These are claims that are web verifiable, unambiguous, and objective. Claims that do not fall under this are not extracted and are skipped.
Examples:
Claim
Fact Checkable?
The current president is doing a good job.
No.
The current president of the US is 82 years old.
Yes.
In PyTorch, BCEWithLogitLoss’s default reduction is ‘none’.
Yes.
In Verilog, a function cannot be disabled using the disable statement.
Yes.
The conditional distribution of conformal coverage given the calibration data follows Beta(l, n+1−l), where l = ⌊(n+1)α⌋.
Yes.
Conformal prediction is a new and cool statistics research area.
No.
Katara should have been with Zuko, not Aang.
No.
Bob was a 8 year old boy with a fascination with UFOs, born in a random Kansas town.
No.
Confidence intervals
Per-model rating confidence intervals are computed via the sandwich variance estimator, as they are on the previous leaderboards. In particular, letting θ^=argminθL(θ)\hat{\theta} = \operatorname*{argmin}_{\theta} L(\theta), we have:
where H=∑kwk⋅HkH = \sum_k w_k \cdot H_k is the composite Hessian (HkH_k denotes either the human preference or factuality BT Hessian), B=∑k(wk2/Nk)⋅ΣkB = \sum_k \left( w_k^2 / N_k \right) \cdot \Sigma_k is the composite gradient covariance, and NkN_k denotes the corresponding total number of battles. Reported intervals on the leaderboard are 95% intervals derived from this estimator, namely (θ^m±1.96Var^(θ^)mm)(\hat{\theta}_m \pm 1.96\sqrt{\widehat{\text{Var}}(\hat{\theta})_{mm}}) for model mm, mapped into the standard Elo-like rating space.
Confidence interval widths shift as ww changes—in particular, intervals tend to widen at higher ww when factuality data is sparser than human-vote data.
Share this Article
AI算出
主要ニュースainew評価標準
Arena チームが事実性の検証メカニズム(原子論抽出、検索エージェントによる確率算出)と複合ブラッドリー・テリーモデルを用いた統合ランキングを発表しており、AI ベンチマーク分野における具体的な技術的進歩を報じている。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み