LLM評価パネルにおける相関誤差が信頼性を損なう:9人の判事、実効投票は2票のみ
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Apple Machine Learning
Apple Machine Learningチームは、複数の大規模言語モデル(LLM)で構成される評価パネルの信頼性について調査した。その結果、9つの最先端モデルからなるパネルでも、相関する誤差により実質的な有効投票数は約2票に過ぎないことが判明した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
LLM-as-a-judge パネルは、複数のモデルからの投票を集約し、多様なモデルがより信頼性の高い評価をもたらすと期待されています。私たちは、このようなパネルの真の情報価値を測定し、その信頼性が独立した投票という理想からどれだけ乖離しているかを定量化するフレームワークを開発しました。7 つのモデルファミリーに属する 9 つの最先端 LLM からなるパネルを、3 つの自然言語推論データセット(各項目につき人間による注釈が 100 件ずつ)でテストした結果、9 人の判事は実質的に独立した投票約 2 票分の情報しか提供していないことが分かりました。パネルの形式的な独立性の約四分之三…
原文を表示
LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal. Testing a panel of 9 frontier LLMs from 7 model families on three natural language inference datasets (each with 100 human annotations per item), we find that the 9 judges effectively provide only about 2 independent votes’ worth of information. Roughly three-quarters of the panel’s nominal independence…
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み