LLM は A/B テストで人間の結果を代替できるか
本文の状態
日本語全文を表示中
詳細モードで約13分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Spotify AI Engineering
Spotify AI Engineering は、LLM を用いた A/B テストの実現可能性を Upworthy データセットで検証し、生データのバイアスや仮定条件の限界を明らかにした。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月14日 04:11
AI深層分析
キーポイント
LLM による代替は設計ではなく仮定に依存する
ランダム化実験が因果関係を特定できるのは設計によるものだが、LLM 予測への置き換えはその保証を失い、結果の妥当性は特定の仮定にのみ依存することになる。
生データのバイアスと効果量の過小評価
gpt-4o-mini を用いた検証では、生の LLM 予測値をそのまま使用すると実際の人間の処理効果の 39% しか回復できず、治療効果が実際よりも半分以下であると誤って結論付ける結果となった。
バイアスの性質と条件の検証困難性
LLM の出力は単なるノイズではなくゼロ方向への系統的なバイアス(減衰)を示し、新しい実験が過去のデータから遠ざかるほどこの手法の有効性を裏付ける条件は成立しにくくなる。
LLMによるバイアスの影響
LLMの出力は系統的な方向性を持つバイアスにより治療効果をゼロに近づけ、結果として treatments の有効性を過小評価する。このため組織はユーザーへの価値を誤って見積もり、不適切な出荷判断を下すリスクがある。
代理変数としての妥当性条件:代理性
LLM出力が治療効果と人間の結果の間の完全な媒介変数である必要がある。つまり、LLMの予測と共変量を考慮すれば、ユーザーの割り当ては追加的な情報を持たない状態を指す。
重要な引用
LLM predictions can stand in for human outcomes in A/B tests, but only by assumption, not by design.
Using those raw predictions in a standard experimental analysis recovered only 39% of the observed human treatment effect.
The bias is systematic and directional: LLM outcomes attenuate treatment effects toward zero.
The bias is systematic and directional: LLM outcomes attenuate treatment effects toward zero, making treatments look less effective than they are.
編集コメントを表示
編集コメント
Spotify のエンジニアリングチームが、LLM を活用した実験手法の限界をデータに基づき厳しく指摘している点は非常に示唆に富む。効率化を追求する現場において、統計的根拠を軽視しないよう警鐘を鳴らす重要な知見と言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
TL;DR: LLM の予測は、設計ではなく仮定に基づいて A/B テストの人間による結果を代替できます。Upworthy データセット(数千件の A/B テストを含む)への適用では、LLM の出力を人間のデータで較正することで処置効果が回復しましたが、これは特定の手法セットに限定された結果でした。この手法が機能する条件は新しい処置に対して検証できず、過去の実験から新しい処置が離れるほどその妥当性は低下します。つまり、最も恩恵が期待される局面においてこそ、この約束の根拠は最も薄弱なのです。
この約束は明快です。ユーザーではなくモデルで実験を行い、数週間の代わりに数時間で結果を得て、トラフィックの割り当てを完全に省略すること。しかし、LLM を用いて人間を代替しようとする提案の多くは、実験が有効となるための統計的な問い——「どのような条件下で、実験が関心のある処置効果を特定できるのか」という根本的な問い——を無視しています。
ランダム化実験は、設計上治療効果を因果的に特定できるため「ゴールドスタンダード」とされています。しかし、実際のユーザー反応を LLM が生成した予測に置き換えると、この保証が失われます。その場合の特定可能性は、あくまで仮定に依存することになります。
私たちは、生物統計学の代理エンドポイント理論を用いて、その仮定が何を指すのかを形式化した論文 a paper を発表しました。考え方はシンプルです。「治療の効果が結果に与える影響をすべて捉えているなら、その代理指標は有効な代用となり、それに対する実験も正しい結果をもたらす」というものです。
臨床試験では、迅速かつ低コストで臨床結果を予測するものとして、ラボで得られるバイオマーカーが一般的に代理指標として使われています。一方、デジタル環境における A/B テストでは、ユーザー反応の代理指標として LLM の予測が有望な候補となっています。その魅力は、A/B テストに伴う手間や時間、機会費用を回避できる点にあります。
大規模言語モデルの予測はノイズではなくバイアスである
私たちは、現在利用可能な最も大規模なオープンアクセスデータセットである Upworthy Research Archive を用いて、LLM による A/B テストの可能性を実証的に評価しました。このデータセットには、数千件の A/B テストにわたるニュース見出しのバリエーションごとのクリック率が含まれています。私たちは gpt-4o-mini に、各見出しに対して典型的なユーザーのクリック率を予測させました。これは治療群と対照群のバリエーションそれぞれで個別に行います。
これらの生データ(raw predictions)を標準的な実験分析に用いた場合、観測された人間の処理効果(treatment effect)のうち 39% しか再現できませんでした。もしこれらの予測値を実際の人間データであるかのように扱えば、「治療の効果が実際よりも半分以下しかない」という誤った結論に至ってしまいます。
これは単なるランダムなノイズではありません。バイアスは体系的かつ方向性を持っており、LLM の出力は処理効果をゼロに近づけるように減衰させます。その結果、治療の効果を実際よりも低く見せてしまうのです。A/B テストの目的は、ユーザーが何を価値あると感じているかを学び、より良い製品判断を下すことにあります。もし LLM による実験を多数の機能や製品領域で広く用いた場合、この減衰効果により組織はユーザーにもたらされる価値を見誤り、間違ったリリース判断を下すリスクが高まります。
LLM の出力が有効な代替指標となるための二つの条件
私たちの論文では、LLM の予測を用いて人間の平均処理効果を特定できる、2 つの条件を形式化して示しています。
代理変数仮説。この仮定では、LLM の出力が人間への介入効果の媒介を完全に担う必要があります。LLM の予測値と、介入とは無関係なベースライン特性を表す共変量を考慮した上で、ユーザーが介入群または対照群に割り当てられたことが、そのユーザーが次に何をするかについて追加的な情報を一切与えないことを意味します。平易な言葉で言えば、「人間への反応に関係する介入のすべての側面を LLM が捉えている」ということです。この仮定は LLM を用いた A/B テストにおいて暗黙的に前提とされることが多いものの、明確に記述されることは稀であり、実証検証が行われることもほとんどありません。

*図 1. 代理変数仮説。共変量を条件とした場合、LLM の予測値が人間への介入効果を完全に媒介する。
比較可能性。この仮定では、LLM の予測値と人間の結果の間の関係、すなわち較正関数が、新しい実験においても、その推定に用いられた歴史的データと同じでなければならないことを要求します。もし介入の変化に伴って LLM の予測値から人間行動へのマッピングが変わってしまうと、較正は崩れてしまいます。この仮定はより一般的には、「介入前の特性と LLM 予測の全分布が実験間で安定していること」を要求するものとして拡張できます。これが満たされれば、平均的な介入効果だけでなく、効果分布に関する他の指標も特定することが可能になります。

*図 2. 比較可能性の仮定。共変量と LLM の予測値を条件とした場合、人間の結果の分布は人間サンプルと LLM サンプルの間で安定していることを示す。
この 2 つの条件がともに満たされれば、LLM の出力をユーザーを対象とした A/B テストデータ上でキャリブレーションすることで、人間の介入効果を回復させることができます。いずれかの条件が崩れると、推定値はバイアスを含んでしまいます。このバイアスの原因はデータの不足ではなく、手順自体が別のものを測定しているためです。したがって、LLM の予測を無限に生成したとしても、ユーザーへの影響は回復されず、あくまで LLM 自身への影響しか得られません。ユーザー実験における無作為割り当てとは異なり、これらの仮定は設計上自動的に保証されるものではありません。
補正手法は効果を取り戻すが、適切な方法が必須
すべての補正手法が等しく機能するわけではありません。Upworthy のデータを用いて2つの手法を検証した結果、線形補正(OLS:最小二乗法による)は偽装検定に失敗しました。これは、学習に用いなかった実験データに対して、LLM 応答から測定された補正後の効果と人間の実験結果との間に差があるかどうかを確認する統計的テストです。
もし「差がない」という帰無仮説が棄却される場合、その補正推定量は過去のデータにおいて信頼できないと判断され、新しいデータに対しても信用すべきではありません。今回のケースでは、OLS でフィットさせた線形補正は、LLM の予測が人間の行動にどのようにマッピングされているかを捉えるには硬すぎました。その結果、人間を基準とした値から3.8標準誤差も離れてしまいました。
一方、機械学習モデル(ランダムフォレストと勾配ブースティング木)はより良い結果を示しました。これらの手法では、補正後の推定量が人間の効果の標本誤差の範囲内に収まり、統計的に有意な差は見られませんでした。つまり、これらの機械学習モデルは柔軟性に富み、LLM の予測と人間の結果との間の非線形関係を学習できるため、LLM の予測を適切に補正することができたのです。
LLM の予測がノイズに起因する別の問題があります。これはトークン生成におけるランダム性、つまりサンプリング温度によるものです。これを考慮しないと、LLM に本質的に備わったこのランダム性が効果推定値をゼロに偏らせ、その分散を増大させる傾向があります。
当論文では測定誤差理論に基づき、実験単位ごとに LLM から複数の出力を生成し、その平均を新しい LLM 予測として用いることで、この問題を緩和できることを示しました。直感的には、LLM の予測に含まれるノイズ成分が平均化されるため、残った部分が真のシグナルに近づくからです。
本当の限界は将来の介入に関するもの
代理性と比較可能性の条件は、過去のデータに基づいて部分的に評価できます。しかし、これまで一度もテストしたことがない介入に対して、これらの条件が常に成立することを証明することは不可能です。
これが中心的な制約となります。新しい介入がこれまでのテスト内容からどれだけ離れているかによって、LLM の出力を人間の反応の妥当な代替として信頼する根拠は弱まります。UI パラダイムの変更や新価格モデル、これまでリリースしたことがないような機能など、真に新しい介入においては、前提条件そのものが検証不可能です。つまり、LLM が A/B テストにおいて最も有益となるのは、まさに LLM が最も失敗しやすい状況なのです。したがって、真の製品革新には人間による実験が不可欠であり続けます。
この背景を踏まえると、Upworthy のデータセットは LLM を代理として用いたテストのほぼ理想的なケースと言えます。結果が「クリックされたか否か」という二値であり、介入手段もテキストベースで言語的な類似性を持つ(見出しのバリエーション)ためです。また、LLM はどのような要素が見出しを魅力的にするかを学ぶために膨大な量のテキストデータで訓練されています。
一方、レイアウトやアルゴリズム、価格を変更する介入については、LLM を用いた A/B テストを行うための必要条件を満たすことが難しく、正当化が困難です。例えば、企業が保有するイノベーションのポートフォリオ全体にわたってこれらの条件が一般化して成立するという実証的な証拠はなく、大規模な運用には不十分です。
収集したくないデータこそ、キャリブレーションには必要
本研究では、LLM を用いた A/B テストは理論上は機能し得るものの、強い仮定が成り立つこと、手法を文脈に応じて慎重に適用すること、そして徹底的な検証が必要であることを示しています。Upworthy のデータセットには同種の過去実験が数千件含まれていますが、多くの製品チームが扱うのは異なる画面や製品領域からの多様な実験のコレクションであり、ログ形式も一様ではない場合があります。
このフレームワークはユーザー実験そのものを不要にするものではありません。新しい実験が既に実施されたものに近い場合にのみ、外挿によってギャップを埋めることで必要なユーザー実験の数を減らせることを示しているに過ぎません。LLM を用いた A/B テストを信頼できるものとするためには、*実際の* ユーザー反応データを収集するための初期投資は必須であり、省略可能なものではありません。
LLM の進化はさらに複雑さを増しています。どの校正関数も、特定のモデルの特定時点に最適化されたものです。LLM プロバイダーはモデルを継続的に更新・置き換えます。今日学習した校正関数が、同じモデルであっても 6 ヶ月後には有効でなくなる可能性があります。また、新しい校正関数は理想的には新たなユーザー実験データに基づいて作成すべきです。そうしないと、時間的なバイアスが生じる恐れがあります。
たとえ LLM の性能向上やプロンプトの改善、ファインチューニングを通じて、代理指標としての妥当性や比較可能性が時間とともに高まったとしても、LLM の出力を校正したり、実際にユーザー成果と対応しているかを確認するためには、依然としてユーザー実験を実行する必要があります。LLM 予測技術の進歩は、人間による検証から逃れる手段ではありません。
おわりに:ユーザー実験は設計上機能し、LLM ベースの実験は仮定に依存する
代理指標フレームワークは、LLM の予測が人間の行動に対する有効な プロキシ指標 となり得るのかを検討します。他のすべてのプロキシ指標と同様に、LLM を用いた代理指標も、プロキシとアウトカムの間の関係性が変化するまで機能します。このフレームワークは、その関係を明確にし、検証可能にするとともに、関係が崩壊した際の帰結を浮き彫りにします。ただし、最も重要な実験、つまり新しい何かを検証する実験において、その関係性が常に成り立つことを保証するものではありません。
ただし、LLM の予測は人間による実験をより良くする可能性があります。例えば、実験の枠を消費する前に弱いアイデアをフィルタリングしたり、分散削減のための共変量として活用したりできます。これにより、過去のデータが豊富で、新しい処置が過去の事例と類似している場合、実験の選定や効率性を向上させることができます。
しかし、人間による結果を LM に置き換えるのは別の話です。これは「設計による同定」を「仮定による同定」に差し替える行為であり、決して軽視すべきではありません。
原文を表示
TL;DR: LLM predictions can stand in for human outcomes in A/B tests, but only by assumption, not by design. In an application to the Upworthy dataset with thousands of A/B tests, calibrating LLM outputs against human data recovered the treatment effect, but only using a specific set of methods. The conditions that make this work cannot be verified for new treatments, and they become less plausible the further the new treatment is from past experiments. The promise is least justified precisely when it offers the most benefit.
The promise is straightforward: run the experiment on a model instead of on users, get results in hours instead of weeks, and skip the traffic allocation entirely. But most proposals for replacing humans with LLM in A/B tests skip the statistical question that makes experiments valid in the first place: under what conditions does the experiment identify the treatment effect of interest?
Randomized experiments are considered the gold-standard because they causally identify the treatment effect by design. Replacing real user responses with LLM-generated predictions removes that guarantee. Identification then holds only by assumption. We wrote a paper that formalizes what those assumptions are, using surrogate endpoint theory from biostatistics. Our idea is simple: If a surrogate outcome captures everything about a treatment that matters for the outcome, then, clearly, the surrogate is a valid proxy and experimenting on it will give the correct result. In clinical trials, biomarkers from labs are commonly used as fast and cheap surrogates for the clinical result. For A/B tests in digital environments, LLM predictions have become attractive candidates as surrogates for user responses. The promise is that LLM predictions sidesteps the effort, time, and opportunity costs of A/B testing.
Raw LLM predictions are biased, not just noisy
We empirically evaluated the promise of LLM-based A/B testing using the Upworthy Research Archive, the largest open-access dataset on A/B tests currently available. The dataset contains click-through rates for variants of news headlines across thousands of A/B tests. We prompted gpt-4o-mini to predict the click-through rate of a typical user for each headline, separately for the treatment and control variants. Using those raw predictions in a standard experimental analysis recovered only 39% of the observed human treatment effect. If you used those predictions as if they were human data, you would conclude that treatments are less than half as effective as they actually are.
This is not just random noise. The bias is systematic and directional: LLM outcomes attenuate treatment effects toward zero, making treatments look less effective than they are. A/B tests are used to learn what users value for making better product decisions. If LLM-based experiments are used across many features and product areas, that attenuation causes the organization to underestimate the value they bring to users, and potentially make incorrect shipping decisions.
Two conditions make LLM outputs valid surrogates
Our paper formalizes two conditions under which LLM predictions can be used to identify a human average treatment effect.
Surrogacy. This assumption requires that the LLM output fully mediates the treatment effect on the human outcome. After accounting for the LLM's prediction and any covariates capturing baseline characteristics independent of treatment, the assignment of a user to the treatment or control condition tells you nothing additional about what the user would do. In plain language: the LLM captures everything about the treatment that matters for the human response. This assumption is often implicitly assumed but rarely spelled out in LLM-based A/B testing, nor is it commonly validated.

Comparability. This assumption requires that the relationship between LLM predictions and human outcomes, the calibration function, remains the same in the new experiment as in the historical data used to estimate it. If the way LLM predictions map to human behavior shifts when the treatment changes, the calibration breaks. This assumption can be generalized to require that the full distribution over pre-treatment characteristics and LLM predictions is stable across experiments. If satisfied, this allows for identifying not just the average treatment effect but other quantities of the effect distribution.

When both conditions hold, you can calibrate LLM outputs on A/B test data on users to recover the human treatment effect. When either fails, the estimate will be biased. This bias is not because of a lack of data, but because the procedure identifies something else. Thus, even if you generated an infinite number of LLM predictions, the effect on users would not be recovered, but just the effect on the LLM. Unlike random assignment in a user experiment, neither assumption is guaranteed by design.
Calibration recovers the effect, but only with the right method
Not all calibration methods work equally well. We tested two on the Upworthy data. Linear calibration using ordinary least squares (OLS) failed the falsification test: a statistical test that checks whether the calibrated effects measured on the LLM responses are different from the human effects on experiments that were held out from training. If the null hypothesis of no difference is rejected, we conclude that the calibrated estimator is unreliable for past data and should not be trusted on new data either. In this case, we found that linear calibration fitted with OLS was too rigid to capture how LLM predictions map to human behavior, landing *3.8 standard errors* from the human benchmark.
Machine learning models (random forest and gradient-boosted trees) worked better. For these methods, the calibrated estimate fell within the sampling error of the human effect and was thus not statistically significant. These machine learning models were therefore flexible enough to learn the nonlinear relationship between LLM predictions and human outcomes, thereby calibrating the LLM predictions appropriately.
A separate problem is that a single LLM prediction is noisy because of sampling temperature (the randomness in token generation). Unless accounted for, this randomness inherent to LLMs will tend to bias the effect estimate to zero and increase its variance. In our paper, we draw upon measurement error theory and show that simply drawing several outputs from the LLM per experimental unit and using the average as the new LLM prediction mitigates this problem. Intuitively, this stems from the noise component in the LLM predictions is then averaged out, leaving what's left to be closer to the true signal.
The real limitation is about future interventions
The surrogacy and comparability conditions can be partially assessed on historical data. They can never be proved to hold for a treatment you have never tested before.
This is a central constraint. The further a new treatment departs from what you have tested before, the weaker the basis for trusting the LLM output as a valid stand-in for human responses. For genuinely new interventions, such as a different UI paradigm, a new pricing model, a feature unlike anything you have shipped, the assumptions are inherently untestable. That is, the setting in which LLMs offer the most benefit for A/B testing is precisely where they are least likely to work. Human experiments therefore remain indispensable for true product innovation.
Against this background, it should be noted that the Upworthy dataset is a near-ideal test case for LLM surrogacy. The outcome is binary (click or not), the treatments are text-based and linguistically similar (headline variants), and LLMs are trained on vast amounts of text about what makes headlines engaging. For treatments that change layouts, algorithms, or pricing, the necessary conditions for LLM-based A/B testing are harder to justify. There is no empirical evidence that the conditions hold in general, for instance, across a company's portfolio of innovations, as would be required to use it at scale.
Calibration needs the data you are trying to avoid collecting
Our work shows that LLM-based A/B testing can, in theory, work, but requires strong assumptions to hold, careful and context-dependent application of methods, and thorough validation. The Upworthy dataset includes thousands of past experiments of a single type. Most product teams instead have a diverse collection of experiments from different surfaces and product areas, which may even come with different logging. The framework does not eliminate the need for user experiments; it shows that you can only reduce how many user experiments you need when the new experiments resemble ones already run, by letting extrapolation fill in the gap. The upfront investment in collecting *actual* user responses is not optional, but is what makes LLM-based A/B testing trustworthy.
Changes in LLMs add further complications. Any calibration function is fit to a specific model at a specific point in time. Providers of LLMs update and replace models. A calibration function learned today may not be valid six months from now, even for the same model. Moreover, any new calibration function should ideally be fitted on new user experiments, as otherwise it may be temporally biased. Even if surrogacy and comparability can be made more realistic over time, for instance through better LLMs, prompting or fine-tuning, one still needs to run user experiments to calibrate the LLM outputs, or check that they indeed do map to user outcomes. Advances in LLM predictions is not a way out of human validation.
Final words: User experiments work by design, LLM-based experiments by assumption
The surrogacy framework considers when an LLM prediction is a valid proxy metric for human behavior. Like all proxy metrics, LLM surrogates work until the relationship between proxy and outcome shifts. The framework makes that relationship explicit, testable, and highlights its consequences when it fails. It cannot guarantee the relationship holds for the experiment you care about most: the one testing something new.
That said, LLM predictions can make human experiments better. They can be used to filter out weak ideas before they consume an experiment slot, or serve as covariates for variance reduction. They thereby offer improvements to experimental selection and efficiency when you have strong historical data and the new treatment resembles past ones. Substituting them for human outcomes is a different move. It trades identification by design for identification by assumption. Do not give that up.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み