金融タスクにおける専門家判断の模倣学習(14 分読了)
本文の状態
日本語全文を表示中
詳細モードで約14分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
TLDR AI は、最先端モデルが単純な金融タスクで苦戦する一方、専門投資家がラベル付けした独自データで微調整されたカスタムモデルの方が性能が高く安価であると報告し、今後は組織ごとに最適化されたモデルが主流になると予測している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
情報の評価#
市場に勝つことは容易ではありません。すべての投資家が同じ公開情報源を利用できる状況では、優位性(アルファ)を生み出すのは、経験と勘に基づく独自の洞察しかありません。優れた投資家の判断力は、人間であれ AI であれ、他者に明確に説明したり直接教えたりするのが極めて困難です。それは経験から生まれるものです。
投資家の業務を最も単純な構成要素に分解してみても、そのタスクは LLM(大規模言語モデル)にとって驚くほど難しいことがわかります。本稿では、投資判断に関連する情報を浮き彫りにするために金融文書をフィルタリングし処理するという、ごく限られた特殊ケースを取り上げます。
投資家は毎日、ニュース記事、調査レポート、企業資料、メール、社内メモなど、無数の情報にさらされています。読むこと自体は容易です。真の難所は、その上で繰り返される微細な判断作業——フィルタリング、解釈、セグメンテーション、そして有用なシグナルがどこにあるかを見極めること——にあります。これらの判断は投資家の日常業務全体に埋め込まれており、膨大な時間を要します。
私たちは、情報の優先順位付け(トライアージ)を自動化できないかと考えました。つまり、「何を優先して読むべきか」を特定するタスクです。これだけで投資家の生産性を大幅に向上させ、浮いた注意力を高次な統合や意思決定に集中させることが可能になります。
LLM が単純な金融タスクで振る舞いが悪いという現状を踏まえ、私たちは「LLM に金融の判断力を教えることは可能か?」と問いかけました。その結果、高品質な人間の注釈を用いることで、LLM は専門家レベルの品味と判断力を持ってテキストを解釈できるようになることが分かりました。私たちの独自モデルは、コストが数分の1であるにもかかわらず、情報精度と再現性においてテストしたすべての最先端モデルを上回っています。
ここでは、公開が許可されたデータのサブセットに対するトレーニングプロセスとその結果について説明します。また、これらの結果に基づき、組織の特定のニーズに合わせて調整されたモデルを目指す「差別化知能(differentiated intelligence)」というビジョンの萌芽についても触れます。
最先端モデルのパフォーマンス#
私たちは、投資家の日常業務から抽出した 6 つの情報フィルタリングタスクを用いてモデルを評価しました。これら 6 つのタスク以外にも、同様の傾向を示す多くの内部タスクが存在しており、最先端モデルはこれらのタスクでも、私たちが内部でトレーニングしたモデルに劣る結果となりました。
精度(Accuracy)とは、投資家の判断に従って正しくラベル付けされた文書の割合を指します。分類タスクについては、F1 スコアも算出しました。F-score (Wikipedia)。
01
金融記事の関連性
金融記事が与えられた場合、それが C 級投資専門職にとって関連があるかどうかを分類します。
評価指標
F1 スコア、精度
02
中央銀行文書の関連性判定
与えられた中央銀行の文書が、将来の金利変動の方向を示唆しているかどうかを分類します。
評価指標
F1 スコア、精度
03
一般的な文書の関連性判定
投資家の質問と調査文書が与えられた場合、その文書が質問への回答に役立つかどうかを分類します。
評価指標
F1 スコア、精度
04
アドホックなコンテンツのラベル付け
調査文書は「反復的なもの(定型文の繰り返し)」か、「混合型(定型文に個別の問題分析が加わったもの)」のどちらかに分類し、個別分析が含まれる最後のページを特定します。
評価指標
精度
05
文書の切り捨て箇所の特定
文書内で定型文が始まる箇所を特定します。
評価指標
完全一致精度
06
メールの切り捨て箇所の特定
メール内で定型文が始まる箇所を特定します。
評価指標
完全一致精度
*本ブログ記事で評価する 6 つの金融タスクは、いずれも投資家が日常的に行う業務から抜粋したものです。*
これらのタスクは投資家にとっては些細な作業ですが、意思決定のプロセスを言語化しようとする際に躓くことがあります。以下に、ニュース記事を投資専門家の関心に関連するかどうか分類する例を示します。

関連なし
ft.comTrump insists Greenland is his
© Jeremy Banx
関連あり
ft.com 米国株、トランプ氏の新たな中国関税発言で急落
S&P500指数が4月以来となる最大の一日安を記録し、数週間にわたる上昇相場に終止符を打った © AFP/Getty Images
*金融記事の米国市場への関連性を判断する例。出典:フィナンシャル・タイムズ*
記事の文脈を考慮すればグリーンランドの事例は真剣に受け取られることは unlikely ですが、中国関税の問題は極めて重要です。しかし両方の事例は地政学と金融の交差点に触れています。
一方、私たちがテストした最先端モデルは投資家たちとは対照的に、驚くほど低い性能しか示しませんでした。Gemini、Claude、GPT の各バリアントを、6 つのタスクを実行するよう指示するプロンプトを与えたところ、平均精度はわずか約 50% に留まりました。
まず私たちは、より強力なプロンプトによって大規模言語モデル(LLM)のパフォーマンスを向上させようと試みました。専門家が実際のタスク記述に基づいて指示を作成し、一部のタスクの枠組みを見直すよう提案しました。例えば、小規模な IPO に関する記事は明らかに金融関連ですが、ブリッジウォーター・アソシエイツのようなマクロ経済投資家にとって興味深い広範な意義には欠けています。LLM は、ニュース記事を「関連性があり興味深い」「関連性はあっても興味がない」「無関係」の 3 つのラベルに分類するよう指示された際、記事分類タスクのパフォーマンスが向上しました。
これらの変更により、精度は単なる偶然の的中(50%)から70%台後半に向上しました。しかし、自動プロンプト最適化手法からはさらに精度を高める効果は見られませんでした。最良のプロンプトを使用しても、テストした最先端モデルの精度はまだ80%未満にとどまり、これは投資家が日常業務で信頼できるシステムとして期待する閾値には達していません。
47.2
77.2
50.1
74.3
47.2
75.8
48.5
78.2
45.6
78.0
*手動および自動プロンプトエンジニアリングを適用した後の最先端モデルの金融タスクにおける精度と正例クラスF1スコア。F1スコアは3つの分類タスクの平均値、精度は全6タスクの平均値です。
また、結果から、このタスクにおいて新しいモデルが急速に改善しているわけではないことが示唆されます。特にコスト対効果で言えば、GPT 5.4 は 5.2 よりも 43% 高価ですが、精度はわずかな向上にとどまっています。
明示的なプロンプトでは、専門家が言葉にできる直観のみを伝えることができますが、最も重要な判断の多くは言語化するのが困難です。ファインチューニングはこの課題を回避します。専門家の直観を静的なプロンプトに変形するのではなく、トレーニングプロセスを通じてモデル自身が独自の判断力を獲得させるのです。オープンウェイトモデルで、これらのタスクにおいて最先端モデルを上回る性能を出すことは可能でしょうか?
トレーニングデータセットの構築#
カスタムモデルを訓練する際の最初の課題は、高品質な投資家の嗜好を反映したデータセットを取得することでした。特に、多くの情報は投資のプロの判断を通じて初めて有用なものになります。
当初、非専門家がラベル付けを行ったベンダーからデータセットを入手しました。しかし、このデータセットで訓練されたモデルは依然として性能が低く、結果も芳しくありませんでした。モデルの推論プロセスを追跡したところ、データセット内のラベルに誤りが多いことが判明しました。専門家によるラベル付けはコストがかかるため、私たちは争点のある事例のみを専門家に送る検証スキームを考案しました。
このスキームの手順は以下の通りです。まず非専門家のラベル付けデータを基にモデルを訓練し、同じデータで評価を行いました。モデルの回答とラベル付け者の意見が一致しない事例については、専門家による再評価のために送付します。もしモデルが自身の学習セットから得た事例にも対応できないのであれば、その事例は本質的に困難であるか、あるいは元のラベルに誤りがあったかのどちらかです。この手順によって訓練データの品質を向上させ、最終的な評価は別途用意したテストセットに対して行いました。
訓練レシピ#
私たちは Thinking Machines Lab の「Tinker」を用いてモデルを訓練しました。Tinker を利用することで、GPU インフラの心配をせずに迅速に試行錯誤を進めることができました。
ベースモデルには、学術文献で微調整性能が広く研究されている Qwen3-235B を採用しました。
まずは批判者不要のシンプルな出発点として、標準的な GRPO と重要性サンプリング損失を用いて開始しました。このベースライン手法によりモデルのパフォーマンスは劇的に向上しましたが、それでも目標とする 80% の閾値には届きませんでした。
| モデル / 学習手法 | 平均精度 | 平均正例 F1 スコア |
|---|---|---|
| Qwen Base | 44.8% | 55.24% |
| Qwen + GRPO | 73.48% | 88.95% |
パフォーマンスをさらに引き上げるために、トレーニングレシピに以下の改良を加えました。
1. インターリーブバッチ処理#
マルチタスク学習のレシピにおいて、3 つのバッチ戦略を比較しました。1 つは各タスクを順次実行する方式、2 つ目はバッチ内で全タスクを完全に混合する方式、そして 3 つ目はラウンドロビン形式でタスクごとに 1 バッチずつインターリーブ(交互に配置)する方式です。その結果、インターリーブ方式が最も効果的であることが判明し、完全混合バッチと比較して精度を 12.1% 向上させることができました。
2. 非対称クリッピングを備えた CISPO ロス#
標準的な重要性サンプリングロスの代わりに、非対称クリッピングを備えた CISPO ロスCISPO loss with asymmetric clipping (arXiv) を採用しました。試行したロス関数やクリッピングスキームの中でこれが最も優れた結果を示し、重要性サンプリングベースラインと比較して精度を 10.1% 向上させました。
3. 強力な教師によるオンポリシー蒸留#
オンポリシー蒸留On-Policy Distillation (OPD) を用いて学習を行いました。この手法では、Kevin Lu 氏らとの共同研究に基づき、以下のようにアドバンテージを構築します:
r=reward−β⋅avg(student_lp−teacher_lp)
r = \text{reward} - \beta \cdot \operatorname{avg}(\text{student\_lp} - \text{teacher\_lp})
advi=ri−avg(r)
\text{adv}_i = r_i - \operatorname{avg}(r)
学生モデルが教師モデルの分布から逸脱すると報酬がペナルティとして減点され、タスク学習中のポリシーを正則化します。
20 歩ごとに現在のチェックポイントを教師モデルに昇格させますが、検証精度が新たな最高値を更新した場合に限ります。これにより、より弱いモデルへ知識蒸留が行われるのを防ぎ、凍結されたベースモデルを教師とした場合と比較してさらに 3.1% の性能向上を実現しました。
Results#
最適なトレーニングレシピを見つけるには、異なるアプローチを複数回試す必要がありました。Tinker の使いやすさのおかげで迅速な実験を行い、手法を洗練させることができました。
*精度と価格の比較:訓練済みモデルと最先端モデル。当社のモデルは世代を超えて両方の指標において最先端モデルを上回っています。*
訓練済みのモデルでは平均精度が 78.2% から 84.7% に向上し、評価した最も優れた最先端モデルと比較して誤りが 29.8% 減少しました。この精度レベルは、私たちの日常業務に十分であると判断しています。
また、モデルサイズが小さいため推論コストも大幅に削減され、タスクあたりのコストは 13.8 倍の低減となりました。今後は特定のタスク支援のために訓練された複数のモデルを活用し、組織全体で AI をスケールさせる予定ですが、その際のコストは重要な検討事項です。
トレーニングレシピの各部分をアブレーション(除去)実験し、それぞれの要素が性能にどのように寄与するかを検証しました。
| トレーニング手法の比較検討 | 平均精度 | 平均正例 F1 スコア |
|---|---|---|
| Qwen + 最終レシピ | 84.66% | 92.99% |
| インターリーブバッチ処理 | 72.18% | 89.01% |
| CISPO + 非対称クリップ | 74.56% | 90.64% |
| OPD | 72.39% | 87.93% |
| OPD w/ Best Val Accuracy Teacher | 81.55% | 89.41% |
各行は、その特定のコンポーネントを除外した最終レシピ(Leave-one-out アブレーション)を示しています。
結論
今回テストしたフロンティアモデルは、比較的単純な金融タスクにおいても苦戦しており、モデルの進化が性能向上に直結していないことがわかりました。一方、専門投資家がラベル付けを行った高品質な独自データセットを用いてファインチューニングを施すことで、フロンティアモデルを上回るカスタムモデルを構築できることを実証しました。この知見は、本稿で取り上げた 6 つのタスクに限定されるものではなく、より広範な領域でも有効であることが確認されています。
精度の高さだけでなく、カスタムモデルはコスト面でも大幅な優位性を持っています。今後、Tinker のような迅速な実験を可能にするトレーニング基盤が整備されることで、カスタムモデルの学習による生産性向上効果がさらに高まると期待されます。
今回の結果は、組織ごとの特定のニーズに合わせて調整されたカスタムモデルがフロンティアモデルを上回る「差別化された知能」への未来の可能性を示唆するものです。
引用
本論文を引用する際は、以下の形式をご利用ください:
Su, Sarah; Zhu, Kevin; Xiao, Emily; Alur, Rohan; Kang, Daniel (Bridgewater AIA Labs), "Learning to replicate expert judgment in financial tasks",
Thinking Machines Lab: News, June 2026.
または、BibTeX 形式の引用は以下を使用してください:
@article{su2026expertjudgment,
author = {Sarah Su, Kevin Zhu, Emily Xiao, Rohan Alur, Daniel Kang (Bridgewater AIA Labs)},
title = {Learning to replicate expert judgment in financial tasks},
journal = {Thinking Machines Lab: News},
year = {2026},
note = {https://thinkingmachines.ai/news/learning-to-replicate-expert-judgment-in-financial-tasks/}
}
原文を表示
Judging information#
Outperforming the market is hard. When every investor has access to the same sources of public information, alpha must come from unique insight built on taste and judgment. A strong investor’s judgment is difficult to articulate and teach directly to others, whether human or AI. It comes from experience.
Even when we decompose an investor’s job into its simplest constituent tasks, those tasks turn out to be surprisingly difficult for LLMs. In this post, we consider a simple special case: filtering and processing financial documents to surface information relevant to investment decisions.
Investors are bombarded with information every day: news articles, research reports, company documents, emails, internal write-ups, and more. Reading is the easy part. The real work is the small, repeated judgments carried over it — filtering, interpreting, segmenting, and identifying where the useful signal lies. These judgments are embedded throughout an investor’s daily workflow and consume substantial time.
We wanted to see if we could automate the information triage task: identifying what is relevant and interesting to read. This alone could greatly augment investors’ productivity, letting them spend their freed up attention on higher-level synthesis and decision making.
Given that LLMs perform poorly on simple financial tasks, we asked: is it possible to teach LLMs financial judgement? We find that with high-quality human annotations, we can teach LLMs to interpret text with expert-level taste and judgement. Our proprietary model outperforms all frontier models we tested on information accuracy and recall, at a fraction of their cost.
We describe our training process and results on a subset of data cleared for public release. Based on our results, we further describe the seeds of a vision of *differentiated intelligence*, with models tuned for specific organizational needs.
Frontier model performance#
We evaluated models on six information filtering tasks drawn from investors’ daily workflows. Beyond these tasks, we have many others internally that show similar patterns to these six tasks: frontier models we tested on underperform compared to our internally trained models.
We measured accuracy — the percentage of documents that were correctly labeled according to our investors. For classification tasks, we also calculated the F1 score.F-score (Wikipedia).
01
Financial Article Relevancy
Given a financial article, classify whether it is relevant to a C-suite investment professional.
EVAL METRICS
F1 score, Accuracy
02
Central Bank Document Relevancy
Given a central bank document, classify whether it signals the direction of future interest rate changes.
EVAL METRICS
F1 score, Accuracy
03
Generic Document Relevancy
Given an investor's question and a research document, classify whether the document helps answer it.
EVAL METRICS
F1 score, Accuracy
04
Ad Hoc Content Labeling
Research documents are either recurring (repeated boilerplate) or mixed (boilerplate plus one-off, issue-specific analysis). Classify which, and find the last page of issue-specific content.
EVAL METRICS
Accuracy
05
Document Truncation
Identify where boilerplate content begins in a document.
EVAL METRICS
Exact Match Accuracy
06
Email Truncation
Identify where boilerplate content begins in an email.
EVAL METRICS
Exact Match Accuracy
*The six financial tasks we evaluate in this blog post, each drawn from the routine work of an investor.*
These tasks are trivial for investors, but they get stuck when articulating their decision process. Consider the following example of classifying a news article as relevant to an investment professional below:

Not relevant
ft.comTrump insists Greenland is his
© Jeremy Banx
Relevant
ft.comUS stocks close sharply lower after Trump threatens new China tariffs
Biggest one-day drop in S&P 500 since April brings weeks long rally to a halt © AFP/Getty Images
*Example of judging the relevance of a financial article to US markets. Source: Financial Times.*
The Greenland example is unlikely to be taken seriously given the context of the article, while the China tariffs are highly relevant. Yet both examples touch on geopolitics and finance.
In contrast to our investors, frontier models we tested on perform surprisingly poorly. Variants of Gemini, Claude, and GPT averaged a mere ~50% accuracy when given a prompt that simply states each of the six tasks to perform.
We first tried to improve LLM performance with stronger prompting. Our experts wrote instructions based on real task descriptions, and also suggested reframing certain tasks. For example, while an article about a small IPO is clearly financially relevant, it lacks the broad significance that would make it interesting to a macroeconomic investor at Bridgewater. LLM performance on the article classification task improved when they were asked to sort news stories into three labels: relevant and interesting, relevant but uninteresting, and irrelevant.
These changes boosted their accuracy from a coin flip to the mid-70s. We saw no further gains in accuracy from automatic prompt-optimization methods. With our best prompts the frontier models we tested on still achieved less than 80% accuracy — the threshold investors expect from a system they could trust in their daily workflow.
47.2
77.2
50.1
74.3
47.2
75.8
48.5
78.2
45.6
78.0
*Accuracy & Positive Class F1 score of frontier models on our financial tasks after manual and automatic prompt engineering. F1 score is averaged across our 3 classification tasks, and accuracy is averaged across all 6 tasks.*
Our results also suggest that newer models aren’t improving rapidly at this task, especially per dollar spent. GPT 5.4 costs 43% more than 5.2 but is only marginally more accurate.
An explicit prompt can only convey the intuition an expert is able to put into words, while the judgments that matter most are often the hardest to articulate. Fine-tuning sidesteps this: rather than contorting the expert’s intuition into a static prompt, the training process lets the model develop its own judgment. Could we train open-weight models to outperform frontier models we tested on these tasks?
Training dataset construction#
The first challenge of training a custom model was acquiring a dataset that reflects high-quality investor taste. In particular, much of the information is only useful when filtered through an investment professional’s judgment.
We initially sourced a dataset from vendors providing non-expert labeling. Models trained on this dataset still performed poorly. After examining the reasoning traces of the model we realized that the labels in the dataset were often wrong. Since expert labelers are costly, we devised a verification scheme that routes only the contested examples to experts.
The scheme worked as follows: we trained a model on the dataset from non-expert labelers, then evaluated it on the same data. Examples where the model’s answer differed from the labelers’ were sent to our experts for reevaluation — if a model couldn’t match an example from its own training set then either the example is genuinely difficult, or the original label was wrong. This procedure was used to clean the training set data; the final evaluation was done on a held out test set.
Training recipe#
We trained our models on Tinker from Thinking Machines Lab.Tinker. Tinker allowed us to iterate quickly without worrying about GPU infrastructure.
We chose Qwen3-235B as the base model as its fine-tuning performance is widely studied in the academic literature.
We began with standard GRPO and importance-sampling loss as a simple, critic-free starting point. This baseline approach resulted in a massive jump in the model performance, but it still fell short of our desired 80% threshold.
| Model / Training | Average Accuracy | Average Pos F1 |
|---|---|---|
| Qwen Base | 44.8% | 55.24% |
| Qwen + GRPO | 73.48% | 88.95% |
We make the following modifications to our training recipe to push performance farther:
1. Interleaved batching#
For our multi-task training recipe, we compared three batching strategies: training each task sequentially, fully mixing tasks within a batch, and interleaving one batch per task in round-robin order. We found interleaving worked best, improving accuracy by 12.1% over fully mixed batches.
2. CISPO loss with asymmetric clipping#
We used CISPO loss with asymmetric clippingCISPO loss with asymmetric clipping (arXiv). to replace the standard importance-sampling loss. Across the loss functions and clipping schemes we tried, this performed best, improving accuracy by 10.1% over the importance-sampling baseline.
3. On-policy distillation with strong teachers#
We train with on-policy distillationOn-Policy Distillation, Kevin Lu in collaboration with others (Thinking Machines). (OPD), constructing the advantage as follows:
r=reward−β⋅avg(student_lp−teacher_lp)
r = \text{reward} - \beta \cdot \operatorname{avg}(\text{student\_lp} - \text{teacher\_lp})
advi=ri−avg(r)
\text{adv}_i = r_i - \operatorname{avg}(r)
The reward is penalized when the student drifts from the teacher’s distribution, regularizing the policy while it learns the task.
Every 20 steps, we promote the current checkpoint to the teacher — but only if validation accuracy has reached a new high, so we never distill toward a weaker model. This gave a further 3.1% gain over a frozen base-model teacher.
Results#
Finding the optimal training recipe required several iterations of different approaches. Tinker’s accessibility allowed us to run fast experiments and refine our approach.
*Accuracy versus price for our trained model and frontier models. Our model outperforms frontier models on both dimensions across generations.*
Our trained model improves average accuracy from 78.2% to 84.7%, meaning the trained model makes 29.8% fewer mistakes than the best frontier model we evaluated. We find this level of accuracy is sufficient for our daily work.
Our trained model is also vastly cheaper due to its smaller size: a 13.8x reduction in inference costs per task. As we plan to rely on more models trained to help with specific tasks and to scale AI across the organization, cost is an important consideration.
We ablated each part of our training recipe to show how each portion contributes to performance.
| Training Method Ablations | Average Accuracy | Avg Pos F1 |
|---|---|---|
| Qwen + Final Recipe | 84.66% | 92.99% |
| Interleaved Batching | 72.18% | 89.01% |
| CISPO + Asymmetric Clips | 74.56% | 90.64% |
| OPD | 72.39% | 87.93% |
| OPD w/ Best Val Accuracy Teacher | 81.55% | 89.41% |
*Each row shows the final recipe with that single component removed (leave one out ablations)*
Conclusion#
Frontier models we tested on struggle with relatively simple financial tasks, and model advances don’t improve performance much. In contrast, we’ve shown that high-quality proprietary datasets labeled by expert investors and used for fine-tuning produce custom models that exceed frontier performance on our tasks. We have found that this outcome holds true well beyond the six tasks we’ve discussed in this post.
Aside from higher accuracy, custom models are also substantially cheaper. We expect to see more productivity gains from custom model training in the future, especially with the availability of training infrastructure like Tinker that enables rapid experimentation.
Our results show the possibility of a future of differentiated intelligence, where custom models tuned to specific organizational needs outperform frontier models.
Citation#
Please cite this work as:
Su, Sarah; Zhu, Kevin; Xiao, Emily; Alur, Rohan; Kang, Daniel (Bridgewater AIA Labs), "Learning to replicate expert judgment in financial tasks",
Thinking Machines Lab: News, June 2026.
Or use the BibTeX citation:
@article{su2026expertjudgment,
author = {Sarah Su, Kevin Zhu, Emily Xiao, Rohan Alur, Daniel Kang (Bridgewater AIA Labs)},
title = {Learning to replicate expert judgment in financial tasks},
journal = {Thinking Machines Lab: News},
year = {2026},
note = {https://thinkingmachines.ai/news/learning-to-replicate-expert-judgment-in-financial-tasks/}
}
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み