AI エージェント推論評価ベンチマーク「Legora Benchmark」発表
Legora は顧客の実際の利用環境を再現する「Legora BAR」ベンチマークを発表し、合成データによるスコア操作を防ぎ、システム全体の実用品質を継続的に評価する方針を示した。
AI深層分析を開く2026年7月28日 00:11
AI深層分析
キーポイント
実環境ベースの評価基準
Legora は既存の合成環境やオープンソースケースに依存せず、顧客が日常的に使用する Legora aOS™とハッチング内で実際の法務ケースを評価するベンチマークを採用した。
システム全体の品質測定
単なるモデルの性能だけでなく、ツール、スキル、法的ソース、ワークフローを含むシステム全体が連携して生み出す成果物の質に焦点を当てて評価を行う。
継続的な改善ループ
ベンチマークは一度きりの評価ではなく、毎回の評価結果に基づいて改善点を特定し、即座にプラットフォームへ反映させる継続的なサイクルを構築している。
重要な引用
The Legora harness equips the model with the tools, skills, legal sources, and workflows needed to perform legal work.
We're taking a different approach. The quality of a legal AI product is about more than the underlying model.
編集コメントを表示
編集コメント
法務分野における AI の信頼性を高めるため、ベンチマークの設計思想を「実環境での再現性」にシフトさせた点は評価できる。このアプローチは、他の垂直領域(Vertical AI)においても同様の課題解決のモデルとなり得るだろう。
なぜこのベンチマークを構築したのか
毎日、何千もの法務チームが Legora 上で最も重要な業務を遂行しています。この信頼に応えるためには、最良のモデルを実際に活用し、Legora プロダクトの改善に注力していく責任があります。そのためには、再現性があり信頼できる方法で自社の成果を測定する必要があります。それが「Legora BAR」です。
プロダクトの品質を高めるには、顧客が実際に利用している姿に合わせて評価しなければなりません。多くの法務向けベンチマークは、技術的な評価として設計されており、ベンチマーク用に特別に作られた簡易なハーンネス(環境)の中で動作する合成データ上で実行されます。こうした環境も機能はしますが、法務チームが現場で直面する現実を完全に反映しているわけではありません。また、オープンソースの事例を用いると、研究機関がモデルの学習データとして利用できてしまい、公的なベンチマークでのスコアが人為的に高くなる恐れがあります。
私たちは異なるアプローチを採用しました。法務 AI プロダクトの品質は、基盤となるモデルだけで決まるわけではありません。モデルそのものと、それを動かすハーンネス、アクセス可能なツール群、そしてシステム全体が組み合わさって初めて成り立ちます。Legora BAR は、実在する法務事例に基づき、Legora ハーンネスおよび Legora aOS™ 内で動作するモデルを評価します。これは顧客が毎日使っているのと同じ環境です。現実の環境で得られた結果こそが、顧客が実際に体験している姿を表しています。これが私たちが測定基準として掲げる標準です。
Legora の BAR(ベンチマーク)の目的は、単に進捗を測定することだけではありません。その加速こそが狙いです。私たちはこのベンチマークを継続的に実行しており、各評価を通じて改善すべき箇所を特定し、その成果を即座に Legora プラットフォームへ反映させています。これにより、「評価→特定→改善」のサイクルが常時回され、製品は着実に高品質化し、結果として顧客にとっての価値も時間とともに増大していきます。
ベンチマークの基本原則
法務チームが AI を評価する際、最も重視するのは「成果物の質」です。この質を決定するのはシステム全体が連携して働くことなので、Legora ではシステム全体を対象としたベンチマークを実施しています。これこそが、法務チームが実際に体験する品質を正確に測る唯一の方法だからです。
- Legora ハーネスは、法律業務を実行するために必要なツール、スキル、法的ソース、ワークフローをモデルに付与します。
- 入力プロンプトは、Legora に指示するタスクです。法務担当者の多様性を反映し、目的明確かつ範囲が適切に設定されたものです。
- ケース環境とは、エージェントが作業を行う対象となる事案のことです。ここには、弁護士が依存するドキュメント、プレイブック、判例、テンプレートなどの資料が格納されます。
- モデルは推論エンジンです。AI の中核機能を担い、私たちは主要な研究機関から提供される最良のモデルを採用しています。
Legora BAR
Legora では現在、28 の実務分野にわたる 5,000 件以上の事例からなる評価コーパスを継続的に拡張しています。各事例は法律専門家によってレビュー済みです。
Legora BAR は、当社の法務パートナーや社内法務エンジニアが作成した数百件の事例から慎重に選定されたサブセットです。これは、実務分野の構成、タスクの種類、難易度のバランスにおいてフルセットと同一であり、実際の法務業務における Legora のパフォーマンスを代表する視点を提供します。
28
高ボリュームレビューからエンドツーエンドの文書作成・助言まで、主要な実務分野すべてを網羅しています。
5,161
事例数
11,075
ソースドキュメント数
事例を構成するファイル群には、元資料、判例、過去のドラフトなどが含まれます。
M&A: 850
IP & tech: 700
不動産: 600
ガバナンス: 345
訴訟: 286
資本市場: 275
PE & VC: 200
信託・相続: 155
ファンド: 145
税務: 135
商事法務: 125
データ・サイバーセキュリティ: 120
ライフサイエンス: 105
特許: 100
独占禁止法: 95
ESG: 90
再編: 85
エネルギー: 80
融資: 80
保険: 75
白领犯罪: 75
雇用法務: 75
仲裁: 70
構造化金融: 65
移民法務: 60
貿易法務: 60
移転価格税制: 55
制裁関連: 55
<120
120–274
275–599
≥600 cases
各評価事例(以下、eval と略記)は、法務業務が作成・実行されるプロセスを反映した 4 つの要素で構成されています。
ケースの概要は、弁護士が Legora で実行する一連のワークフローを示す入力プロンプトです。各ケースは「短時間」「中程度」「長時間」のいずれかに分類されます(後述)。
事案ファイルルームには、関連するすべての文書が保管されるフォルダ構造と対応するドキュメントが含まれます。これには、ファーム固有のテンプレート、プレイブック、先例、およびケース固有の文書などが含まれますが、これらに限定されません。
Legora の回答は、Legora から出力された結果です。単一の文書である場合もあれば、PDF、Word ドキュメント、スプレッドシート、赤線(修正履歴)、またはチャット内での回答といった一連の調整済みのファイル群である場合もあります。
評価基準(ルブリック)は、Legora の回答と比較するための指標です。これらは専門家によって作成されたもので、事実、分析、引用、推奨事項を網羅する個別に検証可能なバイナリ形式の項目で構成されています。これは「ゴールドスタンダード」とされる出力を定義したものです。
各ケースは難易度を含む複数のカテゴリに基づいて分類されます。Legora は単純なプロンプトから複雑なものまで幅広く対応できる必要があります。そのため、ベンチマークには難易度の次元を追加しました。この難易度は法律専門家にとって非常に馴染み深い基準、つまり「法律の専門家が業務を完了するために要する時間」を定義の根拠としています。これは法務ファームが業務範囲を設定し、請求書を作成する際の単位であり、実務分野を超えて通用する指標です。
対応するプロンプスの例は以下の通りです。
難易度レベルと内容:
短時間:数時間、最大で半日程度。限定された範囲で要件が明確な作業(特定の質問への回答や、数項目をチェックする単一の文書作成など)
中程度:1〜3 日
弁護士が担当する典型的な成果物には、メモ、プレイブックとの比較修正版(レッドライン)、訴状の特定セクション、開示スケジュールなどが含まれます。
長期的な案件では、1 週間以上、あるいは数百時間に及ぶ作業が必要となります。
一方、エンドツーエンドの納品物は、事件全体を前進させるものです。例えば、完全なポジションペーパー、完結した訴状、取引文書セットなどが該当します。これには、大規模で複雑な記録全体にわたる、持続的な上級弁護士の判断力が求められます。
なぜケースは非公開なのか
本ベンチマークは、著名なパートナー弁護士から提供された実際の法律事務所のユースケースに基づいて実行されます。データは合成または匿名化されています。参加企業の実務ノウハウや知的財産(IP)を活用しているため、私たちはオープンソース化していません。この制約こそが、むしろ強みとなっています。
ベンチマーク用ケースを公開してしまうと、次世代のモデルの学習データに含まれてしまう恐れがあります。その結果、高いスコアは能力の高さというよりも、テスト内容への熟悉度を反映している可能性が出てきます。そのため、当社のケースは非公開にしています。すべてのスコアは、弁護士が初めて作業に取り組むのと同じように、エージェントがその仕事を初めて見る状態での評価を反映したものです。
透明性の確保のため、AfterQuery と共同で作成した一つのケースをオープンソース化しました。これは BAR 内の非公開ケースと構造を同じくする完全な合成ケースであり、私たちがどのようにケースを構築し、スコアリングしているかを具体的に示しています。AfterQuery がこれをレビューし、当社のデータキュレーション手法が業界基準を満たしていることを確認済みです。
品質の定義 – スコアリング方法
各タスクには、法律専門家が作成した評価基準(ルブリック)が定義されています。これは「優れた回答」に必須の項目を網羅したチェックリストであり、提出された成果物がこの基準をどの程度満たしているかを点数化します。これを「品質ルブリックスコア」と呼びます。
Legora プラットフォーム内では、評価プロセスを完全に隔離されたサンドボックス環境で実行しています。これにより、クライアントが実際に利用する環境と機能的に同一の条件下で、アクセス可能なデータのみを利用した孤立した動作を検証できます。各ケースは統計的な有意性を確保するため連続 3 回実行され、生成された出力は LLM-as-a-judge(LLM を審査員として用いた評価)によって採点されます。その後、その結果が対応するルブリックと照合されます。この手順は対象となるすべてのモデルに対して繰り返し行われます。
なお、当社の評価プラットフォームの技術的セットアップに関する詳細なブログ記事は、近日公開予定です。
01
弁護士によるチェックリスト作成
優れた回答に必須の要素:事実関係、分析内容、結論、および出典情報です。
02
システムによるタスク実行
弁護士が依頼されたのと同様のプロセスで案件を処理し、成果物を生成します。
03
正答箇所の採点
チェックリストを満たした項目の割合がスコアとなります。重要度の高い項目はより重く評価されます。
合意内容レビュー
90%
品質ルブリックスコア
✓
支配権変更条項における重要な変更点をすべて特定
高
✓
解除権を正しく指摘
高
✓
主要なリスクの商業的影響を説明
中
✓
適切な訂正案(レッドライン)を提案
低
✕
二次提出期限の注意点
重要度:低
この回答は、高重要性の項目を正しく満たしており、わずかな軽微な項目を見落としたに過ぎないため、高い評価を得ています。ただし、高重要性の項目を見落とすと、得点は軽微な項目を見落とした場合よりも大幅に低下します。
一部のエージェントベンチマークでは、すべての基準を満たすか、あるいは完全に失敗するかの二択で採点されます。このような「すべてまたはなし」の評価方式は、結果を単なる合格・不合格の二値に圧縮してしまい、システム間の実際の違いや、その差が最も重要となる困難な作業の部分を見えなくしてしまう欠点があります。
当方の評価では、重み付けスコアリングを採用しています。これは、パートナーが成果物レビューを行う際の姿勢を模倣したものです。定義用語の省略と、重要な支配権変更条項の省略はどちらも起草上のミスですが、どのパートナーもこれらを同等に深刻な問題とは見なしません。この重み付けの手順は、ケースレビューを行う法律エンジニアにとって不可欠な業務の一部です。彼らは専門知識に基づいて判断し、成果物がステークホルダーに与えるリスクの程度に応じて、各評価項目を「高」「中」「低」の重要度としてマークします。高重要性の項目を見落とした場合のペナルティが重いため、実質的に不完全な成果物であっても、そのスコアは低いものとなります。
主な知見
ベンチマークの比較対象は、モデルそのものです。各モデルを同じケースとタスクに対して Legora ハーネスで実行し、品質判定器を用いてスコアリングを行います。評価には、出力全体の質、コスト、レイテンシ、引用カバレッジなど複数の次元が含まれます。
出力の質は測定する上で重要な要素ですが、高得点を得たからといって、パフォーマンスの全体像が完全に把握できるわけではありません。反復的な短タスクや、表形式レビューにおける大量処理においては、スピードとコストこそが法務チームが最適化すべき主要指標となる可能性があります。
Legora はモデルに依存しないアプローチを採用しています。私たちは意図的に多様なモデルを評価対象とし、Legora プラットフォーム上に存在するものもあれば、そうでないものもあります。最先端の研究機関とは緊密に協力し、新モデルの評価や、Legora ハーネス内および法務分野でのパフォーマンス向上に向けた最適化方法についてのフィードバックを提供しています。これにより、私たちは AI 技術の最前線に常に立ち、最高のパフォーマンスを発揮するモデルをリアルタイムで顧客へ展開することが可能になります。
今回のベンチマークでは、Anthropic、OpenAI、SpaceXAI から選出された多様なモデルを対象に評価を実施しました。
以下のグラフは、各モデルと難易度レベルにおける出力品質スコアの相対比較を示しています。黒い横線は平均値を表します。
短い法的タスクでは各モデルの性能は概ね同等ですが、タスクが長くなり複雑になるにつれて差が見えてきます。この段階では、事案全体を通じて判断を維持することが極めて重要です。
最も実行時間の長いタスクにおいては、Fable 5、Grok 4.5、Opus 4.8、Sonnet 5 のいずれも、全モデル群の平均を大きく上回る結果を示しました。
難易度別の品質
各難易度パネルは、その難易度における7つのモデルの平均値に対して正規化されています。
「The Legora Benchmark for Agentic Reasoning (4 minute read)」
Claude(Anthropic)、GPT(OpenAI)、Grok(SpaceX AI)の 7 モデルを難易度別に平均したスコアは以下の通りです。
Short(短問)
- 5.6 Luna:0.97×
- 5.6 Sol:0.98×
- 5.6 Terra:0.98×
- Grok 4.5:1.00×
- Opus 5:1.01×
- Fable 5:1.01×
- Opus 4.8:1.02×
- Sonnet 5:1.03×
Medium(中問)
- 5.6 Luna:0.92×
- 5.6 Terra:0.92×
- 5.6 Sol:0.99×
- Sonnet 5:0.99×
- Opus 4.8:1.00×
- Grok 4.5:1.04×
- Fable 5:1.05×
- Opus 5:1.09×
Long(長問)
- 5.6 Terra:0.80×
- 5.6 Luna:0.82×
- 5.6 Sol:0.96×
- Sonnet 5:1.00×
- Opus 4.8:1.04×
- Grok 4.5:1.08×
- Fable 5:1.10×
- Opus 5:1.19×
品質とレイテンシ、コストのトレードオフ
品質と速度のバランスを見ると、モデル群全体にばらつきが見られます。Grok 4.5 は、高い出力品質を維持しながら、ケースあたりの処理時間が最も短いことが確認できました。
品質とケースあたりの中央値時間
7 モデル平均に対する相対的な全体的な品質を、ケースあたりの中央値時間(分)に対してプロットしたグラフです。横軸の右側ほど処理に時間がかかります。
- Grok 4.5:最速かつ最高品質
- Opus 4.8
- Fable 5
- GPT-5.6 Sol
- Sonnet 5
- GPT-5.6 Luna
- GPT-5.6 Terra
- Opus 5
品質とケースあたりのコスト
7 モデル平均に対する相対的な全体的な品質を、公開リスト価格に基づくケースあたりの平均コストに対してプロットしたグラフです。横軸の右側ほど高価になります。
- Grok 4.5:最もコストパフォーマンスに優れる
- Opus 4.8
- Fable 5
- GPT-5.6 Sol
- Sonnet 5
- GPT-5.6 Luna
- GPT-5.6 Terra
- Opus 5
Grok 4.5、Sonnet 5、Opus 4.8 は、平均コストに対して高い品質を提供しており、特に優れた価値を示しています。
モデル固有の評価に加え、「Legora BAR」は、Legora ハーネスの性能向上を導くための重要なツールです。Legora aOS のリリース以降、BAR の結果を活用してハーネスを改善し、本番環境で稼働しているすべてのモデルにおいて、相対的な出力品質を平均 5% 向上させることに成功しました。
同じモデル、1 ヶ月後の比較 - ハーネスの改善効果
- 6 月:2026 年 6 月
- 7 月:2026 年 7 月
グラフは、同一モデルセットを 6 月と 7 月の Legora ハーネスで評価した際の平均ベンチマーク品質を示しています。モデル自体に変更はないため、品質の向上は個々のモデルではなく、Legora ハーネスおよびプラットフォーム側の進化によるものです。このチャートは絶対値ではなく、改善の方向性と概算規模を示すものです。
引用パフォーマンス
法的分析に AI を活用するには高い信頼性が不可欠であり、その信頼は情報源の透明性なしには得られません。弁護士は Legora が生成するすべての回答を事実確認できる必要があります。そのため、Legora はすべての回答に引用(シテーション)を付与しています。
Legora BAR は実際の事件ファイルルーム内で実行されるため、出力品質を表す最も重要な指標の一つである「引用スコア」も測定できます。Legora は生成するすべての回答に引用を必須としているため、以下の 2 つの項目を評価します。
- Cited Answers スコア:各主張に対して、少なくとも 1 つの引用が含まれているかを確認します。
- Grounding スコア:AI が回答を生成するために参照したソース文書が、実際に引用に含まれているかを確認します。つまり、「依存した情報をすべて引用しているか」を検証します。
下のグラフでは、OpenAI モデルが「Cited Answers スコア」で最高位となっています。一方、「Grounding スコア」においては、Fable 5 と Opus 4.8 が特に高いスコアを示しています。
エージェント推論のための Legora ベンチマーク(読み込み時間:約 4 分)
Claude (Anthropic)、GPT (OpenAI)、Grok (SpaceX AI) の各モデルを比較した結果、すべてのモデルの平均スコアは 1.00×と設定されています。これに対し、Claude は 0.80×、GPT は 0.90×、Grok は 1.10×の性能を示しました。
引用された回答の数では、Grok 4.5 が 0.86×、Sonnet 5 が 0.98×の結果を記録しています。
原文を表示
Why we built this benchmark
Thousands of legal teams do their most important work on Legora every day. That trust comes with a responsibility to put the best models to work, and to focus our efforts on making the Legora product measurably better over time. To do this successfully, we need to measure ourselves in a repeatable, reliable way. The Legora BAR is how we do this.
To improve the quality of our product, we have to measure it the way our customers actually use it. Many legal benchmarks are designed as technical evaluations that run in synthetic environments with simplified harnesses built specifically for benchmarking. Those environments are functional but don't fully reflect what a legal team experiences in practice. Additionally, open-sourced cases enable labs to train their models on them, which can create artificially high scores on public benchmarks.
We're taking a different approach. The quality of a legal AI product is about more than the underlying model. It’s the combination of the model, the harness, the tools it has access to, and the system it operates within. Our benchmark evaluates models on real legal cases across practice areas, inside the Legora harness and the Legora aOS™. This is the same harness and system our customers use every day. The environment is real, so the results represent what our customers actually experience. This is the standard we are committed to measuring.
Our intention with the Legora BAR is not just to measure progress, but to accelerate it. We run the benchmark continuously: every evaluation helps our team identify where to improve next, and each improvement ships directly into the Legora platform. This creates a constant loop of evaluation, identification, and improvement that results in a measurably better product and, ultimately, more value for our customers over time.
The Benchmark Fundamentals
When legal teams evaluate AI, they care about one thing: the quality of the work it produces. That quality comes from the whole system working together, which is why we benchmark the full system. It’s the only way to measure what a legal team actually experiences.
- The Legora harness equips the model with the tools, skills, legal sources, and workflows needed to perform legal work.
- The input prompt is the task you give Legora. Goal-oriented, well-scoped, and written with the variability of a real lawyer.
- The case environment is the matter the agent works in. It holds the documents, playbooks, precedents, templates and other materials lawyers rely on.
- The model is the reasoning engine. It supplies the core AI capability, and we run the best models from the leading labs.
The Legora BAR
We maintain a growing evaluation corpus of over 5,000 cases across 28 practice areas (see below), each reviewed by a legal professional. The Legora BAR is a carefully selected subset made up of hundreds of cases drafted by our law firm partners and in-house legal engineers. It reflects the same mix of practice areas, task types, and difficulty levels as the full suite, so the benchmark gives a representative view of Legora's performance across real legal work.
Practice Areas
28
Every major area, from high-volume review to end-to-end drafting and advisory work.
Cases
5,161
Source Documents
11,075
The files that make up the cases, from source material to precedents and prior drafts.
M&A
850
IP & tech
700
Real estate
600
Governance
345
Litigation
286
Capital markets
275
PE & VC
200
Trusts & estates
155
Funds
145
Tax
135
Commercial
125
Data & cyber
120
Life sciences
105
Patents
100
Competition
95
ESG
90
Restructuring
85
Energy
80
Lending
80
Insurance
75
White-collar
75
Employment
75
Arbitration
70
Structured fin.
65
Immigration
60
Trade
60
Transfer pricing
55
Sanctions
55
<120
120–274
275–599
≥600 cases
Each evaluation case, or eval for short, consists of four elements that mirror how legal work is created and executed:
- Case description — The input prompt: an end-to-end workflow a lawyer might run in Legora. Each is classified as short, medium, long (see below).
- The matter file room — The folder structure, and corresponding documents, that holds all documents relevant to the case. This includes, but is not limited to, firm-specific templates, playbooks, precedents, and case-specific documents.
- The Legora answer — The output from Legora. This can be a single document or a coordinated set of PDFs, Word documents, spreadsheets, redlines, or an in-chat answer.
- Rubrics — The criteria against which we compare the Legora answer. They are expert-written, binary, individually verifiable criteria covering facts, analysis, citations, and recommendations. They represent the gold standard output.
Each case is classified based on several categories, including its difficulty level. Legora needs to perform well on both simple and more complex prompts. Therefore, we added the difficulty dimension to our benchmark, where its definition is rooted in something very familiar to legal professionals: how much time the work would take a legal expert. It’s the unit firms scope and bill by, and it translates across practice areas. Examples of corresponding prompts can be found below.
Level
What it looks like
Short
A few hours, up to ~half a day
A bounded, well-specified piece of work: a discrete question, a single document checked against a few points.
Medium
One to three days
A complete work product an associate would be assigned: a memo, a redline against a playbook, a pleaded section, a disclosure schedule.
Long
A week or more, up to hundreds of hours
An end-to-end delivery that moves the whole matter forward: a full position paper, a complete pleading, a deal's document set. Sustained senior judgment across a large, messy record.
Why the cases stay private
The benchmark runs on real law firm use cases supplied by recognized partners with synthesized and/or anonymized data. As we use real know-how and IP from participants, we do not open-source them. That constraint is also a strength. When benchmark cases are made public, they end up in the training data of the next model generation. This means a high score can reflect familiarity with the test as much as capability. Our cases stay private. Every score reflects the agent seeing the work for the first time, the way a lawyer does.
For transparency, we have open-sourced one of the cases, co-created with AfterQuery. It is a fully-synthetic case that mirrors the private cases inside BAR, and shows exactly how we build and score a case. AfterQuery reviewed it to confirm our data-curation methodology meets industry standards.
Defining quality – how we score
Every task comes with a defined list of criteria: the rubrics, covering the items a great answer must get right, written by a legal professional. We score the work based on how much of that checklist it actually gets right. That percentage is the Quality Rubric Score. We run our evaluations in isolated sandbox environments within the Legora platform. This enables an environment to be functionally identical to our client's experience, one where it operates in isolation and restricted to only the accessible data. Each case is run three consecutive times to ensure statistical relevance, and the output gets scored by an LLM-as-a-judge, which then gets compared to the corresponding rubrics. This procedure is repeated for all models. A more in-depth blog post on the technical set-up of our evaluation platform will follow shortly.
01
A lawyer writes the checklist
The must-haves for a great answer: the facts, the analysis, the conclusions, and the sources.
02
The system does the task
It works through the matter and produces the deliverable, the same way a lawyer would be asked to.
03
We tick what it got right
The share of the checklist it satisfies is the score. Important items count for more than minor ones.
Agreement review
90%
Quality rubric score
✓
Identifies every material change-of-control provision
High
✓
Flags the termination rights correctly
High
✓
Explains the commercial impact of the key risks
Medium
✓
Proposes appropriate redline language
Low
✕
Notes the secondary filing deadline
Low
This answer got the high-importance points right and missed only a minor one, so it scores well. Miss something marked high and the score falls much further than missing a low item.
Some agent benchmarks grade all-or-nothing: a deliverable either satisfies every criterion or fails outright. All-or-nothing grading collapses every result into a single pass/fail and hides where systems actually differ, on the hard work where the differences matter most.
We weight our grading instead. Weighted scoring mirrors how a partner reviews work: omitting a defined term and omitting a material change-of-control provision are both drafting errors, but no partner treats them as equally serious. The weighting procedure is an integral part of the legal engineer's work when reviewing cases. They draw knowledge from their expertise, and mark rubrics as high/medium/low importance dependent on how much risk the deliverable exposes its stakeholders to if the rubric was not met. Since high-importance misses are penalized heavily, a materially incomplete deliverable still scores like one.
Key findings
The headline comparison for the benchmark are the models themselves; each model is run through the Legora harness on the same cases and tasks, scored using our quality judge. Our evaluation consists of a number of dimensions including overall output quality, cost, latency and citation coverage. While output quality is a critical component of what we measure, a model scoring high on quality does not necessarily represent the complete picture of performance. For repetitive short tasks, or bulk operations in Tabular Review, speed and cost might be the key metrics legal teams might optimize for.
Legora is model-agnostic. We deliberately evaluate across a range of models, some live in the Legora platform and some not. We collaborate closely with the frontier labs to evaluate new models and provide feedback on how they can be optimized to perform better in the Legora harness and in the legal vertical. This allows us to constantly stay at the frontier of AI technology and deploy the best performing models to our customers in real-time. In this benchmark, we ran our evaluations across a range of models from Anthropic, OpenAI and SpaceXAI.
The graph below showcases the relative output quality scores across various models and difficulty levels, with the black horizontal line representing the average. The models perform similarly on shorter legal work and begin to show variance as the work gets longer and more complex, where sustaining judgment across a full matter is critical. On the longest running tasks, Fable 5, Grok 4.5, Opus 4.8 and Sonnet 5 all significantly outperform the average across the full suite.
Quality by difficulty
Each difficulty panel is normalized to the average of the seven shown models within that difficulty.
Claude (Anthropic)GPT (OpenAI)Grok (SpaceXAI)- Average of the 7 shown models per difficulty (1.00×)0.75×1.00×1.10×1.25×Short5.6 Luna · 0.97×5.6 Luna5.6 Sol · 0.98×5.6 Sol5.6 Terra · 0.98×5.6 TerraGrok 4.5 · 1.00×Grok 4.5Opus 5 · 1.01×Opus 5Fable 5 · 1.01×Fable 5Opus 4.8 · 1.02×Opus 4.8Sonnet 5 · 1.03×Sonnet 5Medium5.6 Luna · 0.92×5.6 Luna5.6 Terra · 0.92×5.6 Terra5.6 Sol · 0.99×5.6 SolSonnet 5 · 0.99×Sonnet 5Opus 4.8 · 1.00×Opus 4.8Grok 4.5 · 1.04×Grok 4.5Fable 5 · 1.05×Fable 5Opus 5 · 1.09×Opus 5Long5.6 Terra · 0.80×5.6 Terra5.6 Luna · 0.82×5.6 Luna5.6 Sol · 0.96×5.6 SolSonnet 5 · 1.00×Sonnet 5Opus 4.8 · 1.04×Opus 4.8Grok 4.5 · 1.08×Grok 4.5Fable 5 · 1.10×Fable 5Opus 5 · 1.19×Opus 5
Quality vs latency and cost
When we look at the trade-off between quality and speed, we see a spread across the model set with Grok 4.5 showing up with lowest time per case, at a high quality output.Quality vs median time per caseOverall quality relative to the seven-model average, against median minutes per case.Claude (Anthropic)GPT (OpenAI)Grok (SpaceXAI)Best: Fast and highest quality0.90×1.00×1.10×Median time per case (further right = slower)Grok 4.5Grok 4.5Opus 4.8Opus 4.8Fable 5Fable 5GPT-5.6 SolGPT-5.6 SolSonnet 5Sonnet 5GPT-5.6 LunaGPT-5.6 LunaGPT-5.6 TerraGPT-5.6 TerraOpus 5Opus 5On Quality vs. Cost, models like Grok 4.5, Sonnet 5, Opus 4.8 deliver high quality and lower cost compared to the average of the set.Quality vs cost per caseOverall quality relative to the seven-model average, against average cost per case at public list prices.Claude (Anthropic)GPT (OpenAI)Grok (SpaceXAI)Best value: Strong and cheap0.90×1.00×1.10×Average cost per case (public list prices) (further right = pricier)Grok 4.5Grok 4.5Opus 4.8Opus 4.8Fable 5Fable 5GPT-5.6 SolGPT-5.6 SolSonnet 5Sonnet 5GPT-5.6 LunaGPT-5.6 LunaGPT-5.6 TerraGPT-5.6 TerraOpus 5Opus 5In addition to model-specific evaluations, The Legora BAR is a critical tool for informing the performance improvements we make to the Legora harness. Since launching the Legora aOS, we have used the BAR results to improve our harness, delivering a 5% increase in relative output quality across all models that are live in production.Same models, one month apart - the harness improvesQualityMay average≈ +5% average upliftJune to July, same modelsJUNE 2026JULY 2026The line shows average benchmark quality across the same set of models, evaluated on our harness in June and again in July. Because the models themselves did not change, the uplift reflects advances in the Legora harness and platform rather than in any individual model. The chart shows direction and approximate scale, not absolute scores.
Citation performance
Leveraging AI for legal analysis requires a high degree of trust, and that trust can not be earned without transparency around information sources. Lawyers must have the ability to fact check every answer Legora produces, which is why we back all of our answers with citations.Because the Legora BAR runs inside a real matter file room, it can also measure one of the most important indicators of output quality: the citation score. Legora requires every answer it produces to be backed by a citation, so we measure two things:The Cited Answers score checks whether each claim carries a citation at all.
- The Grounding score checks whether the source documents the AI consulted to generate an answer, are reflected in the citations. In other words, did it cite everything it relied on?
In the chart below we see that OpenAI models rank highest in Cited Answers scores. On the Grounding Scores, Fable 5 and Opus 4.8 show up particularly strong.
Claude (Anthropic)GPT (OpenAI)Grok (SpacexAI)Average across all models (1.00×)0.80×0.90×1.00×1.10×1.25×Cited answersGrok 4.5 · 0.86×Grok 4.5Sonnet 5 · 0.98×
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み