Harvey、Kimi K3 ベースの法務用モデル「Tenet」を公開
本文の状態
日本語全文を表示中
詳細モードで約30分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
Harvey は Kimi K3 ベースモデルを後訓練した「Harvey Tenet」を発表し、長期の法的タスク実行能力とコスト効率において既存のベースモデルを上回る性能を示した。
AI深層分析を開く2026年8月22日 22:04
AI深層分析
キーポイント
Harvey Tenet の発表と構成
Harvey は Fireworks Research と共同で Kimi K3 ベースモデルを後訓練し、長期の法的タスク実行に特化した「Harvey Tenet」を開発した。
ベンチマークにおける性能向上
合成データと専門家データを組み合わせた学習により、LAB の保持外タスクでほぼ2倍、契約書タスクでは20%の成功率向上を達成し、SOTA 水準に達した。
汎用性と転移学習の効果
訓練中に触れていない Mercor の APEX Agents や Crosby の Redline Bench でも K3 ベースモデルを大幅に上回る性能を発揮し、学習された行動が異なるベンチマークへ堅牢に転移することが確認された。
コスト効率とインフラの改善
トレーニングおよびタスク実行をより効果的に行うためのハネス(harness)の改良により、性能向上と同時にコスト効率も両立していることが示された。
コスト効率の向上
報酬整形を通じて効率的なツール使用と推論を促すことで、同等のパフォーマンスを保ちながら推論時のトークン消費量を削減している。これにより、パフォーマンスを大幅に向上させつつコストを安定させている。
重要な引用
Harvey Tenet is a Kimi K3 base that we post-trained together with Fireworks research for long-horizon legal work.
Our model successfully completes almost twice as many held out tasks on LAB and 20% more on LAB contracts than base Kimi K3.
These gains are particularly interesting, as our model had not seen these benchmarks during training, and thus demonstrate that the model's learned behaviors transfer robustly across benchmarks and harnesses.
In post-training, we incentivized efficient tool use and reasoning through reward shaping, preferring trajectories that would reduce tokens consumed at inference-time given equivalent performance.
編集コメントを表示
編集コメント
Harvey が Kimi K3 ベースモデルを法的ドメインに特化させるための後訓練アプローチを示した事例は、汎用モデルの活用における重要な知見となる。特に訓練データに含まれていないベンチマークでも高い性能を発揮することは、実務環境での信頼性向上に寄与する有望な結果である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
過去半年間、ハーベイの研究アジェンダは以下の2つの目標に注力してきました。
- オープンウェイトモデルを用いた最先端の法務知能の構築
- 法律事務所が独自の専門モデルを構築し、自社の知能を所有できるシステムの創出
本日、これらの研究取り組みに関する更新情報を共有いたします。その一環として、初めてポストトレーニングしたモデル「Harvey Tenet」の初期結果も発表します。Harvey Tenetは、Fireworks Researchと共同で長期的な法務作業に特化してポストトレーニングを行ったKimi K3ベースモデルです。さらに、トレーニングとタスク実行の効率を高めるためのハネス(訓練環境)の改善も取り入れています。
初期の結果では、パフォーマンスとコスト効率の両面で有望な成果が得られています。
品質向上
トレーニングの目的は、モデルによる長期的かつ自律的な法務タスクの実行能力を向上させることにありました。合成データ、公開されている法務データ、そして人間の専門家データからなる複合コーパスに対するポストトレーニングにより、LAB(Legal Agent Benchmark)のホールドアウトタスクにおけるモデルのパフォーマンスが大幅に改善されました。
その結果、Harvey TenetはベースラインであるKimi K3と比較して、LABのホールドアウトタスクをほぼ2倍、LAB契約関連タスクでは20%多く成功させることができました。これにより、両者の全通過率(all-pass rate)はそれぞれ9ポイントと2ポイント向上しています。
これらの改善は、回答の詳細さや引用の質といった広範な基準における向上に基づいています。これはNVIDIA's Nemotronなどの他のモデル群でも確認された傾向と同様です。その結果、LAB Contractsでは最先端のパフォーマンスを達成し、LAB全体では2位に位置しています。
これらの改善は、Mercor の APEX Agents(企業法務)や Crosby の Redline Bench など、他の主要なエージェントベンチマークにも一般化することが分かりました。特に注目すべきは、本モデルがこれらのベンチマークを学習段階で一度も見たことがないにもかかわらず、ベースとなる K3 モデルを大幅に上回る性能を示した点です。これは、モデルが学習した振る舞いが、異なるベンチマークや活用方法の間でも堅牢に転移することを裏付けています。
当モデルは、エージェント機能ではなく法的知識を評価するベンチマークにおいても、高いパフォーマンスを維持しています。これには、Mercor の APEX v1 や Scale の Professional Reasoning Benchmark といった法分野に特化したパラメトリック推論ベンチマーク、さらに LegalBench、CUAD、MAUD といった法的知識を測るベンチマークが含まれます。エージェント機能の向上は、教科書的な法概念に対するベースモデルの広範な理解を損なうことなく実現できることが分かりました。
ベンチマークの詳細については付録をご覧ください。
コスト効率性
2 つ目の主要な発見は、コスト効率性の高さです。オープンモデルへのポストトレーニングにより、モデルのコスト構造を改善する 2 つの機会が生まれます。1 つ目は明白で、オープンウェイトモデルはトークンあたりの価格が安価です。しかし、コストはトークン単価と使用量という 2 つの要素によって決まります。ポストトレーニングでは、報酬形状化を通じて効率的なツール利用と推論を促し、同等のパフォーマンスを達成しつつ推論時のトークン消費量を削減する経路を優先しました。これにより、コストとパフォーマンスの両方を同時に最適化することが可能になり、コストを安定させたまま大幅な性能向上を実現できました。
図 1: LAB ホールドアウトタスクにおける品質とコストのパレートフロンティア。ベースラインスコアのソース:Vals LAB リーダーボード。
データ、特に人間の専門家によるデータは、ポストトレーニングにおいて極めて重要な要素です。私たちは Mercor 社などとの緊密な協力のもと、人間専門家のデータセットを構築・拡張し、合成生成データのレビューと修正を行いました。これにより、大規模な学習が可能となりました。
本稿の後半では、当社のトレーニングおよび評価手法について概説し、研究を通じて開発した追加機能についても取り上げます。これらの機能は順次 Harvey の製品に組み込まれていきます。
トレーニング
Kimi K3 ベースモデルを出発点とし、Fireworks 社と連携して現実的な法務作業環境下で非同期強化学習(RL)を用いたポストトレーニングを行いました。学習中、エージェントには長期にわたる法務タスクが与えられ、効率的に実質的な法務成果物を生成できる能力に基づいて評価されます。
学習環境は Legal Agent Benchmark のタスクと同様の構造を持っています。各環境には以下の要素が含まれます:
タスク指示は、パートナーからの業務依頼として記述されます。
関連する書類や資料を含むクライアント事件の概要も用意します。
さらに、高品質な成果物に必須となる実質的なポイントを網羅した専門家による評価基準(ルブリック)を定義します。
各トレーニングロールアウトでは、エージェントはサンドボックス化されたワークスペースで初期化されます。そこにはタスクに関連するクライアント事件のファイルと、事件を検索し、文書を読み込み、成果物を起草するために必要なツールが用意されています。作業完了後、エージェントは最終的な納品物をディスクに書き出し、エピソードは終了します。
各ロールアウトの評価は、LLM-as-a-judge(LLMによる評価者)を用いてタスクのルブリックに基づいて行われます。報酬は、満たされたルブリック基準の割合に対する微細な項の加重和と、解決された基礎的な法的問題の数に注目する包括的な項の合計として定義されます。さらに、すべての基準で満点を得たロールアウトにはボーナスが加算されます。
評価用モデルの比較実験(アブレーションスタディ)を実施した結果、品質と効率性のバランスにおいて最適な評価者として Kimi 2.6 を採用しました。
エージェントのポリシーは、グループシーケンス・ポリシー最適化(GSPO)を用いて最適化します。各トレーニングタスクでは、独立したロールアウトのグループをサンプリングし、前述の評価基準でそれぞれにスコアリングして、グループ内のアドバンテージを計算します。この際、グループ分散で正規化处理を行います。スコアがほぼ同点となるグループについては再審査を行い、勾配へのノイズ混入を防ぎます。また、簡潔な成果物を促すため、グループ内での長さに関する項を小さく設定しています。
トークンにはシーケンス平均化された重要性重みを用いて加重付けを行います。これにより、安定した効率的なトレーニングが可能となります。さらに、RL(強化学習)の崩壊を防ぐために、両側からのクリッピングを適用し、重要性比率の高いトークンをマスクアウトしています。
追加的なトレーニングの詳細は付録をご覧ください。当社のポストトレーニングにおけるすべての取り組みで、顧客データは一切使用していません。
機能
コアとなる法的タスクの実行に関する研究に加え、標準的なエージェント・ハーンでは提供されていない能力を備えるようモデルをポストトレーニングする手法も探求しました。これらの能力はそれぞれ個別に訓練されており、特定のタスクに取り組むためにモデルがルーティング先として呼び出せるツールやサブエージェントとして展開可能です。
本節では、以下の分野における初期結果について概説します:
M&A 実務におけるデューデリジェンスでは、RLM(Reinforcement Learning from Models)を活用したポストトレーニングにより、大規模かつ長期にわたるタスクの調整をモデルが効果的に行えるようになりました。
文書の大量レビューや構造化データ抽出においては、モデルの効率性と有効性が向上しています。
ファーム固有の知識については、メモリ機能と突発的に現れる構造化分類体系を通じて理解するようトレーニングされたモデルにより、検索の質と効率が大幅に改善されています。
各機能は、タスク固有のデータセットを用いて訓練・評価が行われています。それぞれの機能について、使用したデータセット、現在の成果、そしてなぜこれらがより優れた法務エージェントへの有望な方向性を示すと考えているのかを解説します。
M&A デューデリジェンス
M&A デューデリジェンスのような複雑で文書依存度の高いタスクに対するポストトレーニングにより、著しい進歩が見られました。LAB Diligence における単一のタスクでは、リスクの特定とデューデリジェンスメモの作成のために、最大で 80M トークンに及ぶ文書コンテキストをたどる必要があります。最先端モデルや汎用コーディングエージェントですらこの環境では苦戦しており、LAB Diligence の評価基準において 43.8% を超える基準を満たすベースラインは存在しません。
Figure 2: モデル、ハネス、トレーニング構成ごとの LAB Diligence タスクにおける平均基準通過率。RLM の行は、自己蒸留 SFT を行う前後の同じ GLM-5.2 ルートモデルを比較しています。
このような専門タスクにおいてエージェントのパフォーマンスを最大化するには、そのドメインに適したプリミティブを持つエージェントハネス内でポストトレーニングを行うことが重要です。Baseten 研究との連携により、私たちはエージェントハネスに再帰型言語モデル(RLM)の機能を追加しました。RLM の構成では、ルートエージェントが REPL 環境で各タスクのデータルームを保持し、プログラムによって検索やスライスを行い、文脈ウィンドウを持つサブエージェントへ調査や分析を委任します。GLM-5.2 オーケストレーターのみで RLM ハネスに移行すると、基準通過率は 46.1% に向上します。さらに、高カバレッジのトレースを用いた自己蒸留によって GLM-5.2 を RLM ハネス内でポストトレーニングすることで、通過率は 60.1% に改善されます。これは主に、ベースモデルがデータルームのレビューを過小評価する傾向を修正した結果です。
これらは初期の研究結果であり、詳細は近々公開される技術レポートで共有いたします。
Review Table
ポストトレーニングによって大幅な向上が見られたもう一つの領域は、Harvey の Review Table に代表される高ボリューム文書分析タスクです。Review Table では、ユーザーは最大 10,000 ドキュメントを一度に処理しながら構造化されたデータ抽出と分析を行います。各モデルの回答は正確な引用を含み、明確に構造化されている必要があります。

図 3: Review Table プロダクトのスナップショット。完了したセル、推論プロセス、および引用元を示しています。
Applied Compute と協力し、最先端モデルの能力ギャップを埋めることを目的とした合成データと公開データのコーパスを構築しました。このコーパスを用いて、構造化されたデータ抽出と分析を行う GLM-5.2 モデルをポストトレーニングしました。報酬設計により、正確な引用を含むフォーマットが整った簡潔な回答を生成するようモデルに学習させます。
最強のベースラインと比較すると、ポストトレーニング後のモデルは回答品質で 3.6 ポイント、引用品質で 12.1 ポイント向上し、セルあたりのコストは約 1/10 に抑えられました。これにより、コストと品質のパレート最適 frontier を押し広げることができました。
強化学習(RL)によるポストトレーニング中、モデルは「質問が文書に適用されない場合は回答を控える」といった有用な振る舞いや、「根拠となる証拠を正確に引用し、不要な引用で埋め尽くさない」という行動を学習しました。
図4:各モデルの回答品質と引用品質(0〜1 スケール)。実線黒は訓練済み GLM-5.2(ours)、点線枠は GLM 5.2 ベース。軸は観測範囲に切り詰められています。
これらの改善は、レビューテーブルの生産環境ハーンネス内で直接トレーニングできる能力によるものです。ここではモデルが、検索コンテキストのナビゲーション方法やスキーマ制約、引用要件を学習できます。詳細については当社のブログ記事をご覧ください blog post。
法務知識
同様に有望な成果が、蓄積された法務知識を検索・推論するエージェントのポストトレーニングでも確認されています。LAB: Firm Knowledge におけるプレセデント検索タスクでは、約 100M トークンの合成クライアント事案に対する推論が必要です。法務知識を一切持たないベースモデルは、反復的で徹底的な検索に頼り、すべてのクエリで同じワークスペースを読み返す傾向があります。
これを解決するため、私たちは Engram と提携し、Qwen3.8-27B モデルを訓練しました。このモデルは法務知識をパラメトリックメモリと構造化テキストノートとして符号化します。トレーニング中、エージェントは法務知識ベースを検索し、特徴を 1M トークンの構造化知識に統合します。さらに、自己生成データに対する蒸留と強化学習を通じて、その語彙を重みの中に内在化させます。
図 5: 学習コストと推論コストの比較。各クエリあたりの平均推論トークン数に対する評価基準スコアの推移を示す。実線は Harvey の学習コスト増加時の結果、破線は Opus 4.8 の推論コスト増加時の結果である。学習コストを高めることで品質が向上し、クエリあたりのトークン数が削減されるのに対し、推論コストを増やすとトークン使用量が増大する一方、得られる効果は限定的となる。
この手法により、モデルは関連する法律事務所の知識を効率的に特定する能力を大幅に強化しました。その結果、評価基準の合格率が 15% 以上向上し、タスク完了率は約 10% 高まりました。また、メモリ機能を活用した検索がより的確になることで、完了した処理全体のトークン使用量を 58% 削減することに成功しています。
これらの改善を組み合わせることで、本モデルは最先端の他社モデルと肩を並べる精度で関連事案を特定しタスクを完遂しながら、クエリあたりのコストを 90% 削減しました。
さらに、トークンあたりの知能スコアも約 3 倍に向上しました。これは、推論トークン 10 万個あたりでエージェントが成功裏に完了できる評価基準ポイントの平均値として測定されています。当社のモデルは「トークンあたりの知能」指標で 190.8 を記録し、最良の最先端構成(129.3)や同等サイズの他モデル(37.2)を大きく上回っています。
トークン効率の高い推論経路は、単に高速かつ低コストであるだけでなく、エージェントのコンテキストウィンドウに無関係な結果が混入するのを防ぎます。その結果、大規模な文脈においても、より正確で精密な回答が可能になります。
詳細については、Engram の ブログ記事 をご覧ください。
次のステップ
Tenet とその拡張機能は、環境設計、ポストトレーニング、オープンウェイトモデルにおける当社の研究の大きなマイルストーンです。適切な環境で学習させることで、最も複雑な法的タスクに対してもモデルの能力を向上させられることが示されました。また、新たな研究と組み合わせることで、より効果的な法務エージェントを構築するための全く新しいアプローチが開拓されます。
今後は LAB をより多くの管轄区域、実務分野、ワークフローに拡張し、多様な法的業務全体におけるエージェントのパフォーマンスを理解・改善することを目指します。同時に、計算リソースの拡大にも注力しており、これにより Harvey において新たな一般化モデルと機能を備えた研究から本番環境への移行を加速させます。
LAB とポストトレーニングは、より優れた専門サービス型エージェントの実現に有望な道筋を示しています。私たちはこの約束を果たし、お客様にお届けできるよう取り組むことを楽しみにしています。
謝辞
本研究の達成には、多くのチームとパートナーの皆様の支援が不可欠でした。特に、これらのモデルや環境の構築における連携に尽力いただいた Fireworks、Engram、Baseten、Applied Compute、NVIDIA、Mercor、Snorkel AI に感謝いたします。また、プロジェクトへの貢献をしてくださった Spencer Poff 氏、Nick Gillies 氏、Karl de la Roche 氏にも深くお礼申し上げます。さらに、研究と製品開発の架け橋となるために尽力された Harvey の Assistant、AI Platform、Review Table、Security チームの皆様にも特別な感謝を表します。
付録
ここでは、環境設計、トレーニング、評価手法についてより詳細に説明します。
環境設計
本モデルは、LAB で見られるようなエージェント型法律環境でトレーニングされ、法的ベンチマークを通じて評価されました。各タスクは「クローズド・ユニバース」の顧客事案として構築されています。指示文は簡潔で平均約50語程度であり、期待される出力の詳細な仕様書というよりは、パートナーからの業務依頼として記述されています。
事案ファイルには主要書類と周辺書類が混在しており、法的な争点は複数のファイルにまたまって埋め込まれています。そのためエージェントは、曖昧な指示から始め、事案全体を通じて文脈を構築し、その文脈を活用してレビュー可能な成果物を生成する必要があります。
各タスクの専門家向け評価基準(ルブリック)では、パートナーレベルでのレビューを「事実」「結論」「引用」「重大度評価」「推奨事項」「書式」といった原子レベルの二値判定(合格/不合格)に分解しています。各基準は特定の納品ファイルと紐付けられています。平均して1つのタスクには50項目の評価基準が含まれており、最も大規模なタスクでは数百項目に達します。
事案ファイルの広範さと評価基準の細かさゆえに、単一のロールアウト(実行)で1,000ターン以上を要し、数十万トークンを使用することもあります。
学習の詳細
学習は、Kimi K3 ネットワーク全体に対してランク 64 の LoRA を適用する GSPO(Generalized State-Action Policy Optimization)を用いて行われました。これにより、アテンション層、MLP(多層パーセプトロン)、ルーティングされたエキスパートの重みすべてが適応されます(約 500k のエキスパートテンソル)。最適化には AdamW を採用し、各オプティマイザステップで 8 つのタスクグループ、それぞれに 8 ロールアウトを消費します。学習データセットは前述の通り、約 1,750 のエージェント型法律業務環境から構成されています。各エポックでは、10,000 以上の個別ロールアウトにわたって 150 ステップのオプティマイザ処理を実行しました。このモデルは、2 か月間にわたり約 150 台の NVIDIA B300 GPU で学習されました。
学習中は、エージェントのロールアウトとトレーニング自体が非同期ループで動作します。ローカルオーケストレーターは、ロールアウトデプロイメントに対して複数のエージェントロールアウトを並行して実行し続けながら、完了したグループをトレーナー側が消費します。ここで「硬い古さ(staleness)バウンド」を設定することで、ポリシーからの逸脱を抑制しています。オプティマイザステップのたびに新しい重みが即座にデプロイメントへホットロードされるため、デプロイメントは再起動されず、生成プロセスもほとんど停止しません。
強化学習(RL)の成功には、トレーニング時と推論時の整合性が不可欠です。Kimi K3 は有限精度で提供される大規模混合専門家モデルですが、トレーナーはロールアウト生成時とは異なる数値環境で対数尤度比を再計算します。この不一致が RL の不安定化を招く可能性があります。これを解決するため、Fireworks 学習プラットフォームではカーネルレベルでトレーナーとロールアウトデプロイメントを共同構築し、数値的な乖離を最小限に抑えています(バッチ不変カーネルに関するブログ記事も参照)。ループレベルではさらに、トークン入力・トークン出力の設計を採用し、ルーティングの再実行を通じて整合性を保証しています。
ベンチマーク
ここでは、各ベンチマークで一貫性と精度を確保するために使用した評価手法について説明します。また、他の公開レポートとの実装上の相違点や乖離についても記述します。
上記の表には、当社の評価で追加実行したベースラインモデルも含まれています。
Legal Agent Bench (LAB)
「LAB」は、ハーベイが公開したオープンソースのベンチマークで、24 の専門分野にわたる 1,200 件以上の長期法律エージェントタスクを対象としています。すべての評価スコアは、公式のホールドアウトセットに基づいて算出されました。当社のモデルは標準的なパブリックハーンネスで実行され、トレーニング最適化の一環として追加された「終了ツール(finish tool)」がテスト時のハーンネスにも引き継がれています。なお、Vals の hlab ベンチマーク に基づくベースモデルのスコアも併記しますが、これは同社の評価方法が当社の内部手法とより一致しているためです。
ハーンネスやジャッジの実装における差異により、当社の内部実行結果は他の外部報告パートナーの結果とは異なります。Artificial Analysis は、そのハーンネスとジャッジの手法についてこちらで公開しています。同社は「Stirrup」ハーンネスでモデルを実行し、評価者として Gemini 3.1 Pro を使用しています。Vals はその手法をこちらで文書化しており、LAB を独自の LLM 抽象化レイヤーである「Valkyrie」インフラ上で実行し、キャッシュ効率を高めるためにジャッジの指示を変更しています。Vals によると、これらの変更は純粋にインフラ上の調整であり、標準的なセットアップから本質的な変化をもたらすものではありません。
LAB: コントラクト
「LAB: Contracts」(https://www.harvey.ai/en-US/blog/legal-agent-benchmark-in-house-contracting) は、同ベンチマークを社内契約業務へ拡張したもので、契約書作成、レビュー、交渉にわたる 500 のエージェントタスクを含んでいます。上記の LAB で説明されたハネス、採点、報告手法と同じものを採用し、公式のホールドアウトセット(50 タスク)でのスコアを報告しています。なお、LAB: Contracts に公開リーダーボードはまだ存在しないため、報告されているすべてのスコアは Harvey 内部の実行結果です。
APEX Agents
APEX Agents は、Mercor がシミュレーションされた業務環境で設定した、長期にわたる専門サービスタスクのベンチマークです。私たちはこのモデルを、実務弁護士が作成したタスクからなる「企業法務弁護士」サブセットで評価しました。内部ハネス上でポストトレーニング済みの Kimi K3 チェックポイントを実行し、各タスクの世界に対するファイルシステムマウントへの直接 bash アクセスを提供しました。これは Mercor の標準的な実装とは異なります。Mercor の実装では、ReAct や Loop エージェントハネスによって駆動される構造化された MCP サーバーを通じて、モデルにファイルやツールを公開します(Archipelago)。
他のモデルへの影響は様々で、スコアが向上したモデルもあれば(Fable 5: +7.7%)、低下したモデルもありました(GPT-5.6 Sol: -4.2%)。そのため、すべてのモデルの主要な評価指標として、Mercor が公開している スコア を採用しました。Kimi K3 を当社の評価環境で実行すると、報告されていた 58.8% から 67.5% に改善されます。
Mercor の公式採点(Gemini 3 Flash, thinking=low)はそのまま使用し、8 回の試行における Pass@1 を報告します。
Redline Bench
RedlineBench は、Crosby が提供する多段階の契約交渉ベンチマークです。モデルは交渉ラウンドを交互に繰り返しながらドキュメント固有の修正(レッドライン)を生成し、弁護士が作成した評価基準に基づいて採点されます。当社のモデルには、Crosby の標準的な Harbor 設定(デフォルトの評価環境、スキルスクリプト、判定セットを含む)を使用しました。Hugging Face でスコアが報告されていない一部のモデルについては、Crosby チームから内部スコアを提供してもらい、許可を得てここに掲載しています。
APEX
APEX は、APEX-Agents の前身であり、Mercor が専門家の評価に基づいて構築したベンチマークです。経済的価値の高い知識労働を 4 つの職業分野で測定するものです。当社のモデルは、その法律関連サブセットである Big Law Associate で評価を行いました。APEX は Mercor によって、保有している 100 のタスクセットに対して実行・採点されました。
Mercor には当社のモデルへのアクセス権限を提供しましたが、同社は実行内容を開示せず、タスクやタスクレベルのスコアについても明かさずに独自に評価を行いました。その他のモデルについては、Mercor の リーダーボード に公開されているスコアを引用しています。
PRBench
PRBench は、Scale AI が法務・金融分野における高リスクな専門的推論能力を測定するためのベンチマークです。専門家によって作成されたタスクが、専門家による評価基準に基づいて採点されます。当社のモデルについては、標準的な設定と評価実装を用いてフルセットで評価を行いました。その他のモデルのスコアは、Scale の 公開スコア を参照してください。
PRBench には「ハード」カテゴリに分類される 250 問のサブセットが存在しますが、同ベンチマークの公開リーダーボードは発表時点では最新情報が反映されておらず、最新の報告モデルは GPT-5 となっています。当社のモデルは K3 ベースラインと比較して、このハードサブセットにおいて統計的に有意な改善は見られませんでしたが、基準通過率は 36.0% から 36.8% に向上しました。
LegalBench
「LegalBench」は、6 つの法的推論タイプにまたがる 162 の法務タスクと、合計 20 万件を超える評価例からなるベンチマークです。私たちは、Hugging Face で利用可能な 161 タスクにおいてモデルを評価しました。なお、「Rule QA」タスクは手動採点を前提としているため除外しています。他のモデルについては、長年運営されている LegalBench リーダーボードを管理する Vals のスコアを採用しました。Vals チームは各タスクタイプから 200 例をサンプリングしており、タスク全体で集計した場合、追加の例がスコアに実質的な影響を与えないことを確認しています。そのため、私たちの評価でもこの 200 例という慣習に従いました。
LegalBench は、従来の完結型モデル(completion models)を対象として設計されています。これは、数少ない出力例(few-shot output)を提示して、モデルに結果を完成させるようプロンプトする形式です。例えば以下の通りです。

現代のチャットモデルでは、こうしたプロンプトがノイズの多い応答を生み出す傾向があります。具体的には、「マークは空想的である」や「答え:空想的」といった追加のプロンプトがないにもかかわらず、応答の前や後に余計な文脈を付与してしまうケースです。これらの応答は、回答と期待される正解(例:"fanciful")との完全一致のみを判定する標準的な LegalBench グレーダーでは不合格となります。
完全一致による採点が必要なケースでは、LegalBench のプロンプト以外に実質的な内容を追加することなく、構造化された出力形式を用いてモデルの応答を統一しました。
契約理解アティカスデータセット (CUAD)
CUAD は、法律専門家によって 41 の条項カテゴリに注釈が付けられた 510 件の商用契約からなる契約レビュー用のデータセットです。当モデルは、このフルデータセットに対して評価を行いました。
もともと CUAD では、モデルの不確実性を推定するために対数尤度(log probabilities)が必要とされ、主要な指標として精度・再現率曲線下面積(AUPR)が報告されています。
しかし、多くの主要なクローズドソースモデルは対数尤度を公開していないため、この指標を独立して計算することは不可能です。そのため、当チームでは CUAD の標準的な 実装 を用いて、すべてのモデルに対して生成 F1 スコアという代替指標を報告します。
この指標は、契約から取得したパスages(記述部分)が、人間が注釈をつけた正解セットと Jaccard 類似度で少なくとも 0.5 の値を示す場合の精度と再現率を測定するものです。Harvey を実行した場合の結果スコアをすべてのモデルについて報告します。
M&A 契約理解用データセット「MAUD」
MAUD は、M&A 契約の理解を目的としたデータセットです。152 の公開企業を対象とした M&A 契約書に対し、92 の取引ポイントに関する質問を用意しています。
本研究では、すべてのモデルについて 92 の質問に対して評価を行いました。各質問につき最大 20 例をサンプリングし、MAUD の全テストセットは約 42,000 行に相当します。このサンプリング手法により、各質問が必ず含まれるようにしつつ、実行規模を抑え、すべてのモデルの標準誤差を 1% 未満に抑えることができました。
MAUD の標準指標は AUPR ですが、これはモデルの対数尤度(log probabilities)へのアクセスがないと計算できません。MAUD では代替となる標準的な評価手法を提供していませんが、同データセットの本質は M&A 契約書の文章を特定のカテゴリに分類するタスクにあります。そのため、本研究では生成された回答の一致率(answer-matching accuracy)を報告します。
この評価方法では、モデルに対して抜粋文、質問、および回答選択肢を表示させ、モデルが選択した回答が正解と一致した場合に正解としてカウントします。各モデルについて、1,840 件のサンプリングされた質問に対する平均精度を報告しています。
原文を表示
Over the past six months, Harvey’s research agenda has focused on two goals:
- Building frontier legal intelligence using open-weight models; and
- Creating systems to allow law firms to build their own specialized models and own their intelligence.
Today, we’re sharing an update on that research effort, including initial results from our first post-trained model, which we’re calling Harvey Tenet. Harvey Tenet is a Kimi K3 base that we post-trained together with Fireworks research for long-horizon legal work. In addition, it incorporates harness improvements to make training and task execution more effective. Our initial work shows promising results for both performance and cost-efficiency.
Quality
Our goal in training was to improve the model's ability to perform long-horizon, agentic legal tasks. Post-training on a combined corpus of synthetic data, publicly-available legal data, and human expert data substantially improves model performance on LAB hold-out tasks. Our model successfully completes almost twice as many held out tasks on LAB and 20% more on LAB contracts than base Kimi K3, increasing all-pass rate by 9 and 2 percentage points, respectively. These improvements are grounded in broad criteria gains around answer detail and citations similar to those we’ve seen across models like NVIDIA’s Nemotron. It achieves state-of-the-art performance on LAB Contracts and places second on LAB.
We find these improvements generalize to other leading agent benchmarks including Mercor’s APEX Agents (Corporate Law) and Crosby’s Redline Bench where our model substantially outperforms the K3 base. These gains are particularly interesting, as our model had not seen these benchmarks during training, and thus demonstrate that the model’s learned behaviors transfer robustly across benchmarks and harnesses.
Our model also maintains strong performance on benchmarks that test legal knowledge, rather than agentic capabilities. This includes legal-specific parametric reasoning benchmarks, such as Mercor’s APEX v1 and, Scale’s Professional Reasoning Benchmark, and legal knowledge benchmarks like LegalBench, CUAD, and MAUD. We find that agentic capabilities can be improved without affecting the base model’s broader understanding of textbook legal concepts.
Full details on our benchmarking can be found in the appendix.
Cost-efficiency
Our second major finding is cost-efficiency. Post-training open models provides two opportunities for improving a model’s cost profile. The first is obvious, open-weight models have cheaper per token prices. But cost is a function of both token prices and tokens used. In post-training, we incentivized efficient tool use and reasoning through reward shaping, preferring trajectories that would reduce tokens consumed at inference-time given equivalent performance. This allowed us to co-optimize for both cost and performance, gaining significant performance while keeping cost stable.
Data, and specifically human expert data, is a crucial component of our post-training work. We collaborated closely with Mercor and others to build and scale human expert datasets and for review and remediation of synthetically-generated data, which made training at this scale possible.
In the rest of this post, we outline our training and evaluation methodology, and highlight additional capabilities that we have developed through our research and will be incorporating into the Harvey product over time.
Training
Starting from a Kimi K3 base, we worked with Fireworks to post-train our model via asynchronous reinforcement learning (RL) in realistic legal work settings. During training, an agent is given long-horizon legal tasks to solve and is graded on its ability to produce substantive legal work product efficiently.
Training environments share the structure of Legal Agent Benchmark tasks. Each environment consists of:
- A task instruction written as a request for work from a partner;
- A client matter containing the relevant documents and materials; and
- An expert rubric that outlines the substantive points a high-quality work product must contain.
For each training rollout, the agent is initialized in a sandboxed workspace containing the task's client matter files and the tools it needs to search the matter, read documents, and draft work product. After completing its work, the agent writes its final deliverables to disk, which ends the episode.
Each rollout is graded via LLM-as-a-judge over the task's rubric. Reward is defined as the weighted sum of a granular term for the fraction of rubric criteria satisfied and a holistic term counting the number of underlying legal issues solved, plus a bonus for rollouts that score perfectly on all criteria. We ran ablations comparing candidate judge models against heavier frontier models for grading and converged on Kimi 2.6 as the optimal judge for quality and efficiency.
We optimize the agent's policy with group-sequence policy optimization (GSPO). For each training task, we sample a group of independent rollouts, score each with the rubric reward described above, and compute advantages within the group, normalized by group variance. Near-tied groups are rejudged to reduce noise entering the gradient, and a small intra-group length term favors concise deliverables. Tokens are weighted via sequence-averaged importance weights, which we find results in stable and efficient training. We additionally enforce double-sided clipping and mask out high-importance-ratio tokens to prevent RL collapse.
Additional training details can be found in the appendix. We did not use any customer data in any of our post-training efforts.
Capabilities
In addition to our work on core legal task execution, we have explored ways to post-train models around capabilities beyond those available in a standard agent harness. These capabilities are each trained separately and can be deployed as tools and sub-agents that our model can route to in order to tackle specific tasks.
In this section, we outline initial results across:
- M&A Diligence: Post-training in RLM harnesses to enable models to effectively coordinate high-scale, long-horizon tasks like M&A diligence.
- Review Tables: Making models more effective and efficient at high-volume document review and structured data extraction.
- Firm Knowledge: Models that are trained to understand firm knowledge through memory and emergent, structured taxonomies, enabling higher-quality and more efficient search.
Each capability is trained and evaluated on a task-specific dataset. For each capability, we describe the dataset, our current results, and why we think they provide a promising direction to better legal agents.
M&A Diligence
We have seen significant gains from post-training on complex, document-intensive tasks like M&A Diligence. On LAB Diligence, a single task requires traversing up to 80M tokens of document context to identify risks and produce a diligence memo. Both frontier models and off-the-shelf coding agents struggle in this setting, with no baseline passing more than 43.8% of rubric criteria on LAB Diligence.
Maximizing agent performance for specialized tasks like this comes from post-training within an agent harness that has the right primitives for the domain. Together with Baseten research, we extended our agent harness to include Recursive Language Models (RLMs). In the RLM construction, a root agent holds each task’s dataroom in a REPL environment, can search and slice it programmatically, and delegate research and analysis to sub-agents with their own context windows. Moving to the RLM harness with a GLM-5.2 orchestrator alone lifts criteria pass rate to 46.1%. Post-training GLM-5.2 within the RLM harness via self-distillation over high-coverage traces further improves the pass rate to 60.1%, largely by correcting the base model's tendency to under-delegate review of the dataroom.
These are early research results and we will be sharing more detail in a forthcoming technical report.
Review Table
Another area where we have seen substantial gains from post-training is in high-volume document analysis tasks like Harvey's Review Table. With Review Table, users perform structured data extraction and analysis at scale, over up to 10,000 documents at a time, and each model response must be well-structured with accurate citations.

Together with Applied Compute, we built a corpus of synthetic and public data targeting capability gaps of frontier models in representative Review Table tasks. We used this corpus to post-train a GLM-5.2 model for structured data extraction and analysis, with reward shaping to encourage well-formatted, concise answers with precise citations. Relative to the strongest baselines, the post-trained model improves answer quality by 3.6 points and citation quality by 12.1 points at roughly one-tenth the cost per cell, pushing the cost-quality Pareto frontier. During RL post-training, the model learns useful behaviors like abstaining when a question does not apply to a document and citing precise supporting evidence rather than padding citations.
We attribute these gains to our ability to train directly inside the Review Table production harness, where the model can learn how to navigate retrieval context, schema constraints, and citation requirements during training. Full details are in our blog post.
Firm Knowledge
We have seen similarly promising gains from post-training agents to search and reason over a firm's accumulated knowledge. On LAB: Firm Knowledge, completing precedent search tasks requires reasoning over ~100M tokens of synthetic client matters. Lacking any knowledge of the firm, base models default to repetitive, exhaustive search, re-reading the same workspace on every query.
To address this, we partnered with Engram to train a Qwen3.8-27B model that encodes the firm’s knowledge into parametric memory and structured text notes. During training, the agent explores the firm knowledge base, consolidates features into 1M tokens of structured knowledge, and internalizes the corpus in its weights through distillation and RL over self-generated data.
This approach substantially improves the model’s ability to efficiently identify relevant law firm knowledge, improving criteria pass rate more than 15% and task completion rate by nearly 10%. Memory also leads to more targeted searches, cutting total tokens in completed trajectories by 58%. Combined, the post-trained model identifies relevant matters and completes tasks competitively with leading frontier models while reducing cost per query by 90%.
It also triples intelligence-per-token, which is measured as the mean rubric criteria points that an agent completes successfully per 100k inference tokens. Our model scores 190.8 in our intelligence-per-token metric against 129.3 for the best frontier configuration and 37.2 for equivalently sized models. Token-efficient trajectories are not only faster and cheaper, they avoid polluting an agent’s context window with irrelevant results allowing for more accurate and precise answers over large-scale context.
Full details are in Engram’s blog post.
What’s next
Tenet and its extended capabilities are a major milestone in our research across environment design, post-training, and open weight models. It shows that training on the right environments can improve model capabilities even for the most complex legal tasks and, when coupled with novel research, unlock entirely new approaches to building more effective legal agents.
Next, we are scaling LAB to more jurisdictions, practice areas, and workflows to help understand and improve agent performance across the full diversity of legal work. We are also focused on scaling our compute so that we can bring our work from research to production with new generalist models and capabilities in Harvey.
LAB and post-training show promise for better professional service agents, and we’re excited to work to deliver on that promise for our customers.
Acknowledgements
This work would not have been possible without the support of many teams and partners. We’re especially grateful to Fireworks, Engram, Baseten, Applied Compute, NVIDIA, Mercor, and Snorkel AI for their partnership in building these models and environments, and to Spencer Poff, Nick Gillies and Karl de la Roche for their contributions to these projects. And a special thanks to the Assistant, AI Platform, Review Table, Security within Harvey for helping to bridge research and product.
Appendix
Here we describe environment design, training, and evaluation methodologies in more detail.
Environment design
Our model was trained on agentic legal environments, similar to those found in LAB, and evaluated across legal benchmarks. Each task is built as a closed-universe client matter. Instructions are short, averaging around 50 words, and are written as a request for work from a partner rather than a detailed specification of the expected output. The matter files mix key and peripheral documents, with the underlying legal issues embedded across multiple files, so the agent must start from a loose instruction, build context across the matter, and use that context to produce reviewable work product.
Each task's expert rubric decomposes partner-level review into atomic, binary pass/fail criteria – facts, conclusions, citations, severity ratings, recommendations, and formatting – with each criterion tied to a specific deliverable file. On average, a task contains 50 rubric criteria, with the largest tasks containing hundreds. Given the breadth of the matter files and the granularity of the rubrics, a single rollout can span >1,000 turns and use hundreds of thousands of tokens.
Training details
Training consisted of GSPO with a rank-64 LoRA over the full Kimi K3 network, adapting all attention, MLP, and routed-expert weights (~500k expert tensors). For optimization, we leveraged AdamW, with each optimizer step consuming 8 task groups of 8 rollouts. The training dataset comprised ~1,750 agentic legal task environments as described above. For each training epoch, we ran 150 optimizer steps across more than 10,000 individual rollouts. The model was trained on approximately 150 NVIDIA B300 GPUs over the course of 2 months.
During training, agent rollouts and the trainer itself run in an asynchronous loop. A local orchestrator keeps a fleet of agent rollouts in flight against the rollout deployments while the trainer consumes finished groups, with a hard staleness bound keeping off-policyness in check. After every optimizer step the new weights are hot-loaded into the deployments in place, so the deployments never restart and generation rarely stops.
A key to successful RL is alignment between training and inference. Kimi K3 is a large mixture-of-experts model served in finite precision, and the trainer recomputes logprobs in a different numerical world than the one that generated the rollout. This misalignment can destabilize RL. To solve this, the Fireworks training platform co-builds the trainer and rollout deployments at the kernel level to minimize numerical gaps (see the blog on batch-invariant kernels). At the loop level, we further adopt a token-in-token-out design with router replay to guarantee alignment.
Benchmarks
Here we discuss the evaluation methodology we used to ensure consistency and accuracy on each reported benchmark. We also describe any implementation divergences or divergences from other public reports of those benchmarks.
The table above includes additional baseline models that we ran in our evaluation.
Legal Agent Bench (LAB)
LAB is Harvey’s open-source benchmark of long-horizon legal agent tasks, consisting of more than 1,200 tasks across 24 practice areas. All evaluations for LAB were scored on the official hold-out set. Our model was run in the standard public harness with the addition of a finish tool, which was added as a training optimization and carried over to the test-time harness. We report scores for base models from Vals as we find they track more closely to our internal methodology.
Our internal runs differ from those of other external reporting partners due to differences in harness and judge implementation. Artificial Analysis documents its harness and judge methodology here. They run models in their stirrup harness and use Gemini 3.1 Pro as a grader. Vals documents its methodology here. They run LAB on their internal LLM abstractions, Valkyrie infrastructure and change judge instructions to allow for more efficient caching. Vals reports that these changes are purely infrastructural and do not represent a substantive change from the canonical setup.
LAB: Contracts
LAB: Contracts extends LAB to in-house contracting work, containing 500 agentic tasks across contract drafting, review, and negotiation. We follow the same harness, grading, and reporting methodology described for LAB above and report scores on a 50-task official hold-out set. All reported scores are internal Harvey runs, as there is not yet a public leaderboard for LAB: Contracts.
APEX Agents
APEX Agents is Mercor's benchmark of long-horizon, professional-services tasks set in simulated work environments. We evaluated our model on the corporate-lawyer subset, whose tasks were authored by practicing attorneys. We ran our post-trained Kimi K3 checkpoint in our internal harness with direct bash access to a filesystem mount of each task's world. This differs from Mercor's canonical implementation, which exposes files/tools to the model through structured MCP servers driven by a ReAct or Loop agent harness (Archipelago).
We found that our approach had a mixed effect on other models, with some scoring better (Fable 5: +7.7%) and some regressing (GPT-5.6 Sol: -4.2%). Accordingly, we chose to report Mercor’s published scores for all models as our primary leaderboard metric. Running Kimi K3 in our harness improves it from its reported score of 58.8% to 67.5%.
We otherwise keep Mercor's official grading (Gemini 3 Flash, thinking=low) and report Pass@1 over 8 trajectories.
Redline Bench
RedlineBench is Crosby's multi-turn, contract negotiation benchmark, in which models produce document-native redlines across alternating negotiation turns and are graded against attorney-authored rubrics. We ran our model using Crosby’s standard Harbor config including their default harness, skill scripts, and judging set up. For some models not reported on Hugging Face, we received internal scores from the Crosby team and reproduce them here with their permission.
APEX
APEX, the predecessor to APEX-Agents, is Mercor's expert-graded benchmark of economically valuable knowledge work across four professions. We evaluated our model on its legal subset, Big Law Associate. APEX was run and graded by Mercor against their held-out set of 100 tasks.
We provided Mercor with access to our model, which they ran independently without disclosing the runs, tasks, or task-level scores. For other models, we use publicly reported scores from Mercor's leaderboard.
PRBench
PRBench is Scale AI’s benchmark of high-stakes professional reasoning in law and finance, with expert-authored tasks graded against expert-curated rubrics. We evaluated our model on the full benchmark using the standard configs and evaluation implementation. For other models, we report Scale's public scores.
PRBench has a 250-question “hard” subset but its public leaderboard is stale as of the time of publication with the most recent model reported being GPT-5. Our model showed non-statistically significant improvement over K3 base on the hard subset, increasing from 36.0% to 36.8% criteria pass rate.
LegalBench
LegalBench is a benchmark of 162 legal tasks spanning six types of legal reasoning and totaling more than 200,000 evaluation examples. We evaluated our model across 161 tasks available on Hugging Face, omitting the Rule QA task that states it is meant to be hand-graded. For other models, we report scores from Vals who run a longstanding LegalBench leaderboard. The Vals team samples 200 examples per task type, finding that additional examples do not materially change scores when aggregated across tasks. We adopted this 200-example convention in our own runs.
LegalBench is designed to run on legacy completion models, with models being prompted to complete the result from a few-shot output. For example:

These types of prompts create noisy responses in modern chat models, which often provide additional context before or after a response without additional prompting such as “the mark is fanciful” or “answer: fanciful”. These responses fail the standard LegalBench grader which does exact text matching against the response and an expected golden answer (“fanciful”). For cases involving exact match grading, we used structured outputs to standardize our model’s responses without providing any additional substantive content beyond the LegalBench prompt.
Contract Understanding Atticus Dataset (CUAD)
CUAD is a contract-review dataset of 510 commercial contracts annotated by legal experts for 41 clause categories. We evaluated our model on the full dataset. As originally described, CUAD requires log probabilities to estimate model uncertainty and reports Area Under the Precision-Recall curve (AUPR) as its headline metric.
Most leading closed-source models do not expose log probabilities, making this metric impossible to compute independently. We instead report CUAD's alternative metric of generative F1 for all models, using CUAD's standard implementation. This metric measures precision and recall of passages retrieved from a contract with at least 0.5 Jaccard similarity to a human-annotated gold set. We report Harvey-run scores for every model.
Merger Agreement Understanding Dataset (MAUD)
MAUD is a dataset for merger agreement understanding, posing 92 deal-point questions over 152 public-target merger agreements. We evaluated our model across all 92 questions, sampling up to 20 examples per question. MAUD's full test set is roughly 42,000 rows; the 20-example cap keeps every question represented while bounding the run size and landing full run standard errors for all models below 1%. We report scores for all models under this sampling methodology.
Like CUAD, MAUD's standard metric is AUPR, which cannot be computed without access to a model's log probabilities. MAUD does not provide a standard alternative grading approach, but the dataset largely consists of classifying passages of a merger agreement into specific categories. We therefore report generative answer-matching accuracy. To compute this, the model is shown the excerpt, question, and answer options and generates its choice, which is scored correct when it matches the gold answer. We report mean per-question accuracy across the 1,840 sampled questions for every model.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み