オープンウェイト LLM が規制・臨床タスクでクローズドモデルと精度同等
本文の状態
日本語全文を表示中
詳細モードで約32分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
TLDR AI の記事は、オープンウェイト大規模言語モデルの精度がクローズドソースモデルに追いついたという分析結果を提示し、開発者や企業にとってオープンソースモデルの実用性が確立された重要な転換点であると論じている。
AI深層分析を開く2026年7月31日 23:16
AI深層分析
キーポイント
オープンウェイトモデルの精度向上
最新のオープンウェイト大規模言語モデルは、従来のクローズドソースモデルと比較しても精度において遜色ないレベルに達していることを示すデータが提示されている。
ベンチマークにおける競争力
主要な評価基準においてオープンソースモデルがトップランナーと同等のスコアを記録しており、技術的な壁が崩れつつある状況が報告されている。
実用化への道筋
精度の向上により、オープンウェイトモデルは研究段階から実際のビジネスアプリケーションや開発現場での採用へと移行する可能性が高まっていると分析している。
規制されたライフサイエンス分野における精度の重要性
文献スクリーニングや臨床データからの構造化データ抽出など、誤答が規制提出書に悪影響を及ぼす文脈では、正確性が最も重要となる。
クローズドな最先端モデルへの依存仮説の崩壊
これらのタスクには閉鎖型の最先端モデルのみが信頼できるとする従来の前提はもはや成立しないことが示された。
重要な引用
Open-Weight LLMs Have Caught Up on Accuracy
21 minute read
In the life-sciences context, accuracy matters most given the high stakes.
The prevailing assumption has been that only closed, frontier models can be trusted with tasks like these. We put that assumption to the test — and found it no longer holds.
編集コメントを表示
編集コメント
この分析は、オープンソース AI の成熟度が劇的に高まったことを示唆しており、技術選定におけるパラダイムシフトを予感させる内容である。開発者はクローズドモデルへの依存度を下げ、自社のニーズに合わせた柔軟なモデル構築の可能性を探るべき時期に来ていると言えるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
規制の厳しいライフサイエンス分野では、市場監視のための文献スクリーニング、臨床文書からの構造化データ抽出、規制提出資料の作成、および臨床試験データに基づく表・リスト・図(TLF)の生成などにおいて、誤った回答は単なる不便ではありません。体系的レビューで論文を見逃したり、データ抽出テーブルに虚偽の数値が含まれたりすれば、それがそのまま規制提出資料へと波及する恐れがあります。ライフサイエンスという文脈では、そのリスクの高さから正確性が何よりも重要となります。
これまで一般的には、こうしたタスクを信頼できるのはクローズドな最先端モデルだけだと考えられてきました。しかし私たちはこの前提を検証し、もはや成り立たないことを確認しました。
ライフサイエンス分野におけるベンチマークの注目は、主に創薬発見に集中してきました。構造予測については、数十年続いている盲検評価「CASP」が業界の標準的な指標として定着しており、これはアルファフォールド(AlphaFold)によってさらに確固たるものとなりました。
また、「LAB-Bench」や「BixBench」は、AI エージェントが生物学文献を推論したり、実際のバイオインフォマティクス解析を実行できるかを検証するものです。特に BixBench は難易度が高く、最新モデルでもリリース時には単一桁から十数パーセント程度のスコアしか出せませんでした。
さらに「Coscientist」は、LLM が実験室のハードウェア上で自ら化学実験を計画・実行する様子を示しました。この成果は重要ですが、あくまでパイプラインの前半部分のみを測るものです。実際には、資金と年月が投入されるのは後半の部分です。
米国議会予算局(CBO)の試算によると、創薬および前臨床研究にかかる費用は製薬 R&D 支出全体の約 31%(承認された薬剤あたり約 4.74 億ドル)であり、一方、臨床試験には約 10.7 億ドルが投じられています。期間についても、臨床試験には平均 95 ヶ月がかかるのに対し、前臨床研究は約 31 ヶ月に過ぎません。
他の規制業界でも、この後半部分に相当する領域の評価が始まっています。
Harvey の Legal Agent Benchmark は、専門家によって作成された評価基準に基づき、1,250 件以上の実際の法律事務所の業務課題に対してエージェントを評価します。OpenAI の HealthBench や Stanford の MedHELM は臨床ケアのタスクを対象としており、FinanceBench は公開された提出書類に対してアナリストの質問を検証します。特定の業界に限定されない汎用的なベンチマークとしては、GDPval が「実際の専門的な成果物」に最も近いものですが、規制当局への提出書類に関する項目は含まれていません。規制対応や臨床試験のワークフローは、ライフサイエンス企業が時間と費用を多く費やす領域ですが、これに対応する公開ベンチマークはありませんでした。ClinReg はその課題に対する私たちの取り組みです。
「ClinReg」では、3 つの実務的な規制対応および臨床試験タスクを対象としています。主要なオープンウェイトモデルとクローズドソースモデルを実行し比較しました。注目すべき点は、これら 3 つのすべてのタスクにおいて、オープンウェイトモデルの精度がクローズドソースモデルの変動範囲内に収まっていることです。さらに、コスト面でも圧倒的な差があります。オープンウェイトモデルは、テストした中で最高スコアを示したモデルの約 3 分の 1、最も高価な最先端モデルに至っては約 10 分の 1のコストで動作します。
ここでは、測定した内容や驚きの事実、そして臨床・規制環境で AI を導入する人々にとってこれが何を意味するのかを順を追って解説します。
顧客にとって、高リスクなワークロードをオープンモデルへ移行させる意義は、主に以下の 3 つの価値提案に集約されます。
- コストを抑えつつ高い精度。オープンウェイトモデルが重要なタスクにおいて最前線のモデルと同等の性能を発揮できれば、経済性は根本から変わります。
- インフラの堅牢性。最先端モデルでは容量やインフラの問題が発生しがちですが、オープンウェイトモデルへ負荷を分散させることで、ワークロードの堅牢性と稼働率を向上できます。
- データ主権。精度と信頼性のほかにも、オープンウェイトモデルの価値は外部プロバイダーへのデータ流出を防ぐ「データ主権」の実現にあります。
本稿では、この 3 つのうち最初の点(経済性)を厳密に証明することに焦点を当てます。なぜなら、これが裏付けられなければ、2 番目と 3 番目の価値提案も意味を持たないからです。
ClinReg Benchmark のタスク、データセット、評価方法、スキル、ハルネス、推論プロバイダー、スコア計算の詳細は付録に記載されています。
19 のモデル(12 社製クローズドと 7 つのオープンウェイト)について「総合スコア」(詳細は付録参照)を測定しました。その結果、ClinReg ベンチマークでは GPT-5.6 Sol が最高スコア 88.4(標準偏差 4.6)を記録し、次点でオープンウェイトモデルの GLM-5.2(87.4; 標準偏差 5.6)と Kimi K3(86.9; 標準偏差 4.6)がこれに続きました。GLM-5.2 と Kimi K3 は、1 タスクあたりの平均コストで GPT-5.6 Sol のそれぞれ 33.8%、59.6% で済みます。つまり、オープンウェイトモデルはすでに最高性能のクローズドモデルと誤差範囲(標準偏差)内に収まっていることがわかります。数ヶ月前までは最高性能だった Opus 4.8/4.7 は、このタスクにおいて GLM-5.2 の約 9 倍のコストがかかっており、コストと精度の改善スピードがいかに速いかを示しています。
これらのモデルに次ぐのは、同程度の精度を持ちながらコストには大きなばらつきがあるクローズドモデル群です。GPT-5.6 Terra、GPT-5.5、Grok 4.5、Opus 5 がこれに含まれ、総合スコアは 84.3 から 86.3 の範囲ですが、コスト差は最大で 11 倍に達します。
一方、米国製のオープンウェイトモデルでは、Gemma 4(73.5)と Nemotron 3 Ultra(70.3)がリスト内で最も低いスコアとなりました。各モデルファミリー内では、後継バージョンが精度向上につながっています。クローズドモデルでも(例:GPT-5.5 → GPT-5.6 Sol、Opus 4-6 → 4.7 → 4.8 → 5)、オープンウェイトモデルでも(例:Kimi 2.6 → 3)同様です。GPT-5.6 ファミリーでは、総合スコアが Sol > Terra > Luna の順で、OpenAI が示したパフォーマンスとコストのランキング順序と一致しています。同様に、Opus シリーズは Sonnet よりも高い性能を示しました。
Fable 5 と Muse Spark 1.1 は、生物学的な安全性の観点からベンチマークタスクを拒否したため、今回の分析からは除外しました。また、Opus と Sonnet のコストは他のモデルファミリーよりも高いことが判明しています。この差が、実行環境(Claude Code)によるものなのか、それともモデル自体の特性によるものなのかについては、今後の調査課題として残っています。
文献スクリーニングでは、すべての評価対象モデルのパフォーマンスが 85〜94% の狭い範囲に集中しており、正解数が 40〜44 件とばらつきも小さいことから、どのモデルもこの分類タスクを適切に処理できていることがわかります。しかし、精度という数値だけでは、モデル間の行動上の有意な違いは見えてきません。各モデルが誤答した論文は異なり、採用・除外基準の適用方法にも差がありました。例えば Sonnet 4.6 や Gemma 4 は再現性(リコール)を重視する傾向があり、一方、MiniMax や DeepSeek といった慎重なモデルは、より多くの論文を除外していました。実務においては、単一のスコアでモデルを選ぶのではなく、その誤りのパターンがレビューチームの「見落とし許容度」と「追加の手作業負荷」のバランスに合致しているかが重要になります。
IND および TLF のタスクについては、エージェント実行時のトレースデータを収集し、エラーパターンの分析を行いました。
IND: 継続力よりも自制心が重要だった。 モデルは、ソースに明示的に記載された情報の抽出においては概ね信頼性があったが、必須項目が存在しない場合における違いは顕著であった。最も性能が高かったモデル(GPT 5.6 ファミリーと Kimi K3)は、網羅性と忠実性のバランスを最適化し、根拠のない推測で穴を埋めることなく、完全なドキュメントを作成した。一方、捏造を避けるために過度に慎重になり、実際のコンテンツすら省略してしまったモデル(Gemini 3.1 Pro など)もあれば、完全性を追求するあまり、明記されていない値をでっち上げてしまったモデル(MiniMax M3 など)もあった。それ以外の最先端モデル(Opus ファミリーや Sonnet 5)も、ソースに裏付けのない仕様や手法を採用してしまったことで、結果的に他社に遅れをとった。
「TLF」の評価では、モデルごとの振る舞いに明確な差が現れました。すべてのモデルは最初の試行で構文エラーのない Python コードを生成しましたが、決定的な違いは実行時のエラーへの対応方法と、検証ループの活用度合いにあります。
GPT モデルファミリーは素早くクリーンなコードを書き、発生したエラーをその場で修正する傾向がありました。ただし、まだ修正可能な問題が残っているにもかかわらず「タスク完了」と宣言してしまうケースも散見されました。一方、Opus 4.8 や Grok 4.5 はスクリプト全体を何度も書き直すことを繰り返しました。
GLM-5.2 や Kimi K3 といった高性能なオープンウェイトモデルは、検証チェックが通過するまで継続的に反復処理を行いました。その反面、Gemma や DeepSeek V4 は多くの試行を重ねても結果が安定せず、正しい答えに収束しませんでした。これは、単なる粘り強さだけでは能力の限界を補いきれないことを示唆しています。
今回の評価設定では、モデルの実力だけでなく「運用スタイル」もパフォーマンスに大きく影響しました。特にミスの修正方法や、「いつ作業を終了と判断するか」という判断基準が鍵となります。また、これらの結果は周辺環境(ハルネス)の影響の大きさも浮き彫りにしています。より強力なバリデーション機能や明確な停止条件を備えた評価枠組みであれば、早期に終了してしまう傾向のあるモデルでも、さらに良い成果を出せる可能性があります。
「TLF」と「IND」の評価結果を合わせると、モデルごとに異なる特性が浮かび上がります。「TLF」は持続的な反復処理を評価する一方、「IND」は抑制されたアプローチを重視します。モデルの能力が競争に参加できる土台を作るものの、最終的にどこで最も高いパフォーマンスを発揮するかは、タスク固有の振る舞いや、ハルネスに組み込まれた採用基準によって決定されます。**
このベンチマークの作成と評価手法(付録参照)を通じて、いくつかの教訓を得ました。特に人間による正解データが存在しないタスクにおいて、これらの知見は重要です。例えば当社の ClinReg ベンチマークでは、最終的なモジュール 3 の出力や、コード生成を含む TLFs の中間段階などが該当します。以下にそのいくつかを共有します。
LLM を単独の判定者として信頼してはいけません。他でも報告されている通り、0〜10 点の評価尺度を用いた LLM による採点には信頼性が欠けており、モデルによって同一の出力に対して異なる点数が付けられることが判明しました。 私たちが IND タスクで実施した予備分析では、批判者として機能させた際にも各モデルの評価基準に大きなばらつきがあることが分かりました。Gemini 3.1 Pro と GLM 5.2 は比較的寛容で平均点が 8/10 でしたが、GPT 5.5 は厳格で平均点は 6/10 でした。自己評価の傾向にも偏りがあり、Gemini 3.1 Pro と GLM 5.2 は自社の出力を好む一方で、GPT 5.5 は自身の出力に対して厳しい態度を示しました。
こうした差異の影響を最小限に抑えるため、私たちはいくつかの対策を採用しました。
IND スコアリングでは、4 人の審査員が得点ではなく発見事実を提示し、「議長」モデルがその事実を確認して重症度を評価した上で最終スコアを付与します。
TLF タスクの批評家パネルでは、3 名の批評家の中間値を採用しました。
決定論的指標の使用。IND タスクでは、108 のフィールドを対象とした中間データ抽出タスクに関連する 2 つの指標(省略率と捏造率)を導入しました。これらは審査モデルによって客観的に測定・採点が可能です。省略率は、ペナルティを避けるために存在するフィールドを欠如としてマークした場合に発生し、捏造率は存在するフィールドに対して誤った値を出力した場合に発生します。
興味深いことに、これらの指標は 2 つの異なる失敗モードを明らかにしました。Opus 4.8 は GLM 5.2 と比較して省略率が高い一方で、捏造率は低い傾向を示しています。
ループ内検証とタスク分解: タスク終了時の評価に加え、実行プロセス内で行う「ループ内検証」も重要な要素です。自律型ワークフローが再帰的な自己改善(出力を反復的に向上させる仕組み)を実現するためには、スキルに明確な停止ルールや受容基準を組み込んだ強力な検証機能が不可欠です。
検証ロジックが決定論的または客観的に定義されている場合は、その実装は比較的容易です。一方、LLM を審査員として用いる場合のように確率的・主観的な要素が含まれる場合でも、上記で述べたような配慮を行うことで、再帰的自己改善の恩恵を享受できます。
実際には、各タスクの詳細に最も精通している専門家が受容基準の策定において大きな役割を果たします。また、スキル内に中間サブタスク(例:IND タスクにおける中間データ抽出)を組み込むタスク分解は、抽象的な高レベルフィードバックに依存するのではなく、エージェントが自己改善の指標として活用できる具体的なマイルストーンを提供することで、精度向上に寄与します。
自律型実行のトレース記録も有用な設計パターンです。これらはエラー分析だけでなく、精度とコストをさらに高めるための後続的微調整や強化学習にも活用できます。
規制業務における精度において、オープンウェイトモデルはついにクローズドモデルと互角の水準に達しました。今回テストしたベストなクローズドモデル(GPT-5.6 Sol、88.4 点)と比較すると、セット内の最上位オープンウェイトモデルである GLM-5.2(87.4 点)や Kimi K3(86.9 点)は、統計的な誤差範囲内(1 標準偏差以内)に収まっています。文献スクリーニング、IND モジュール 3 の作成、TLF 生成といったタスクにおいて、「オープンかクローズドか」という区分けが、誰がトップを争うかを予測する指標ではもはやなくなりました。
コスト格差は片方向に開いており、その差は甚大です。GLM-5.2 は GPT-5.6 Sol の実行あたりのコストの約 3 分の 1、Opus 5 の約 10 分の 1で同様のスコアを達成しています。実際の規制プログラムが想定する処理量において、これは単なる経費項目の問題ではなく、予算規模そのものを変えるほどの差です。
互角であることは、すべてが均一であることを意味しません。見かけ上のスコアがほぼ同じモデルでも、失敗の仕方は明確に異なりました。あるモデルは穴を埋めるために数値を捏造する一方、別のモデルは数値自体を省略します。また、検証に合格するまで反復を試みるモデルもあれば、修正可能なエラーが残ったまま「完了」と宣言してしまうモデルもあります。リーダーボードで 1 ポイント違うからといって、現場では互換性があるわけではありません。ランク順ではなく、エラーの発生パターンに基づいて選定すべきです。
ハルネス(評価枠組み)自体がモデルの一部となります。ここで得られるエージェントの結果は、モデルそのものだけでなく、バリデーターの厳格さや停止基準、リトライ動作といった「足場」の影響も受けます。今回の設定では、どのモデルを使ったかよりも、どのように実行したかによって結果が大きく変動するケースが複数見られました。
米国のオープンウェイトモデルは、非米国製モデルに比べてまだ一歩遅れていました。今回の評価で最も高いスコアを記録したのは GLM-5.2 と Kimi K3 です。一方、米国製のオープンウェイトモデルでは最高性能を示した Gemma 4 31B(73.5 点)と Nemotron 3 Ultra(70.3 点)でも、今回のセット全体で見ると最下位グループに属していました。
しかし、重要な結論は「オープンウェイトモデルが精度面で追いついた」という点にあります。これにより、モデル選び自体はすでに容易な段階になりました。今後はエンジニアリングの課題が残ります。具体的には、「どのモデルとハネス(実行環境)をタスクに合わせて組み合わせるか」「スキルやループ内の検証でどのように受入基準を定義するか」、そして「コストとレイテンシのバランスを考慮して、どの推論プロバイダにルーティングするか」といった判断です。この複雑さが消えるわけではありませんが、その所在が変わるのです。
Everest と Log10 のチームはこの複雑さを吸収し、ライフサイエンス組織が成果物の提出に集中できる環境を提供します。ClinReg といった標準化されたベンチマークを通じて、規制や臨床試験ワークフローにおける AI の可能性への信頼が高まり、最終的には新薬や医療機器を市場に出すまでの時間とコスト削減に貢献することを願っています。
*当社の ClinReg ベンチマークタスクセットに対するフィードバックや追加提案を歓迎します。今後オープンソース化を予定していますので、ai@log10.io までご連絡ください*
私たちは、公開データを活用した実際の規制対応や臨床試験業務を模倣する 3 つのタスクを中心に、「ClinReg」というベンチマークスイートを設計しました。
最初のベンチマークタスクは文献選別です。このタスクでは、成人における肺結核およびリファムピン耐性の検出に関する Xpert MTB/RIF のコクラン診断精度レビュー(CD009593)に基づき、CLEF eHealth TAR のトピック CD009593 を使用します。CLEF TAR は、再現可能な選別タスクとして、実際のシステマティックレビューのワークフローを流用しています。
このタスクは、システマティックレビューの 2 つの段階を模倣しています。タイトル・アブストラクトの選別では、どの記録が全文レビューに進むかを決定し、全文選別では最終的な採用・除外とその理由を決定します。選別基準は、元のレビューで公開された適用基準を手作業で策定したものを基に作成され、評価前に固定化された順序の指示として翻訳されています。モデルは同じ指示に従って動作しますが、レビュープロトコル自体を生成したり書き換えたりするわけではありません。
この研究は独立したコクランレビューであり、正式な市販後監視プログラムではありませんが、そのエビデンス選別ワークフローは、市販後監視や診断装置の性能評価における文献監視コンポーネントに対する現実的で公開された代替指標として機能します。
現在、フルテキストスクリーニングの結果のみを報告しています。評価セットには PubMed Central やその他の Web ソースから入手可能なフルテキスト PDF が含まれる 47 件の論文が含まれており、最終的な人間のレビューラベルとして「20 件が採用」「27 件が除外」とされています。各モデルは、1 記事あたり 1 回ずつ採用・除外の判断を行います。
正規化された文献スクリーニングスコア(Normalized Literature Screening Score)は、全体の判断精度を測定する指標です。計算式は以下の通りです。
正規化された文献スクリーニングスコア = 100 * (正しい採用数と除外数の合計) / 47
また、47 件の論文からなる評価セットをスクリーンするための総コスト(プロバイダー費用)も報告します。これは記録されたトークン使用量に基づいて算出されます。Cochrane レビューでは、除外された研究に対する人間の記録による除外理由が提供されており、後でモデルの判断根拠を検証・監査することが可能です。
私たちは、規制提出書類のモデルを作成します。具体的には、FDAへのIND(新薬臨床試験申請)における「化学・製造・管理(CMC)」セクション(すなわちモジュール3)です。これは実際のIND作成作業を模倣したプロキシ(代理指標)として機能します。
本来の提出書類は、企業が保有する独自の製造および分析データから作成されますが、ベンチマーク化のために、同様の情報を提供する公開ソースに置き換えています。モデルには、承認された抗体薬複合体に対する欧州規制当局のレビュー報告書(EPAR)という単一の密度の高い評価レポートを提示し、そこからIND申請書のCMCセクションを再構築させます。
これは答え付きの一問一答形式ではありません。モデルが最初から最後まで実行する、持続的で自律的なタスクです。この自律的タスクの重要な側面は、EPAR報告書から108項目のフィールド情報を抽出し、モジュール3作成に適切に組み込むことです。
このタスクでは、「LLM-as-a-judge(LLMを審査員として用いる)」手法を用いて、各モデルがEPARからCMCコンテンツをどの程度忠実かつ完全に再構築したかを評価しました。詳細は後述しますが、要約すると、4人のモデルで構成される審査パネルがエラーの証拠を発見・相互検証し、その証拠をもとに1人の仲裁者が包括的な0〜10点の評価スコアを算出します。
4名の審査員(Gemini 3.1 Pro、GPT-5.5、Opus 4.8、GLM-5.2)は、それぞれ以下の2種類の構造化された証拠リストを作成しました。
「fabrications(出典に明記されていない値を埋め込んだ虚偽の項目、例えば架空の制限やメーカー名の記載)」と
- 「omissions(出典で明確に言及されているにもかかわらず、モデルが「MISSING」や曖昧な表現として扱った項目)」の 2 つを評価対象とします。
ICH Q1A や USP <71> といった標準的な引用は虚偽には含めません。これらは定型文だからです。出典で事実上言及されていない項目が欠落している場合でも、それは「誠実な空白」とみなし、問題視しません。
各候補項目は出典と照合して再検証され、4 人の評価者のうち少なくとも 3 人が正しいと認めたもののみを採用します。その後、「Chairman(議長)モデル」が 2 つの採用リストを比較・採点します。この際、議長モデルは出典と各モデルの出力結果を併せて読み込みます。
「Chairman モデル」に関する注記:当初は Fable 5 を選定しましたが、生物学的安全性に関するトリガーによりプロンプトを拒否し、代わりに Opus 4.8 に振り向けられました。これにより、Opus シリーズモデルに対する自己優遇バイアスの可能性が生じることは否定できません。
しかし、このバイアスが結果を歪める可能性は低いと考えられます。その理由は 2 つあります。第一に設計上の理由です。評価者は点数をつけるのではなく、出典に基づいた証拠を提示します。また、採用には 4 人中 3 人の賛成が必要であるため、1 人の評価者が自らの判断を否定する結果を抑えることはできません。
第二に、実際に測定した結果です。前述の予備的な批判分析において、Opus 4.8 はテストされた 4 モデルの中で自己優遇バイアスが最も小さかったことが確認されました(0〜10 のスケールで Δ = −0.14)。一方、Gemini 3.1 Pro は +1.23、GLM-5.2 は +0.67 でした。つまり、実証データに基づけば、Opus 4.8 は利用可能なモデルの中で最も自己に偏らない評価者であったと言えます。
IND スコア(0〜100 点)は、0〜10 点を評価した結果に 10 を乗じて算出されます。各モデルは 3 回実行され、その平均値が採用されました。
FDA が医薬品を承認する際、必ず基盤となるのは「試験の表・リスト・図(TLF)」を中心とした技術資料です。これは、誰にどの治療が行われ、どのような副作用がどの程度の頻度で発生し、主要な有効性エンドポイントがスポンサーが主張した通りに変化したかを示す事前指定された分析データです。
FDA は生データの症例報告書(CRF)を直接参照するわけではありません。提出された数値の背後にある標準化されたデータセットを確認し、それらを生成した分析プログラムを実行して再検証します。スポンサーは、すべての結果を追跡可能な形でパッケージ化した電子提出資料を提出します。これには、分析に使用できるデータセット、その元となった標準化された集計表、そして全体を生成したプログラムが含まれており、審査担当者はこれを一貫して実行することで、報告された主要な結果の再現が可能になります。
臨床試験では、生物統計プログラマーのチームが数週間から数ヶ月をかけてコードを記述します。彼らは診療所で収集した生データ(CRF:症例報告書)を取り込み、CDISC が定める標準データ集計モデル(SDTM)と分析用データモデル(ADaM)という 2 つの標準化されたデータモデルを経由させ、最終的に審査担当者が目にする特定の表、リスト、図を作成します。このプロセスのどの段階でも小さな導出エラーが発生すると、それがパイプライン全体に静かに伝播し、下流の結果が誤っているにもかかわらず表面的には整合性が保たれてしまうことがあります。具体的な例としては、許可された訪問期間を超えて値を保持する「最終観測値推移(LOCF)」ルールや、行番号が 1 つずれた分析対象集団のフラグ、あるいは古くなったバージョンの制御語彙に対してデコードが行われるケースなどが挙げられます。
このベンチマークは、こうしたエンドツーエンドのワークフローを単一の自律的なタスクに圧縮したものです。生データ(CRF)とその仕様書が与えられた場合、各モデルは自律的に CDISC 提出用パッケージ全体を作成しなければなりません。つまり、生 CRF から SDTM、ADaM を経て TLF(表・リスト・図)に至るまでの一連の処理を、単一の自律的な実行ターンで完結させる必要があります。
具体的には、1 つの自律的なターンでモデルが 3 つの Python スクリプトを書き、かつそれらを実行します:
- raw_to_sdtm.py は、生 CRF データを変換し、DM(被験者)、AE(有害事象)、LB(検査値)、VS(バイタルサイン)などを含む 10 の標準化された SDTM ドメインを作成します。
- sdtm_to_adam.py は、SDTM ドメインから分析準備が整った ADaM データセット 7 つを導出します。
- adam_to_tlf.py は、ADaM データセットから TLF を 13 件生成します。
このパイプラインには4つのデータレイヤーがあり、モデルはそれらの間で3つのPythonスクリプトを記述します。これらはすべて1つの自律的な改善ループ内で完結しています。
10のSDTMドメインと7つのADaMデータセットは、CDISCパイロットデータセットによって定義されています。このうち7つは、出力の基盤となる分析準備完了版のサブセットです。一方、13のTLF(Tables, Listings, Figures)は、完全な研究報告書に含まれるはずの表やリスト、図の中からベンチマークが選定した代表的な集合です。
10のSDTMドメインと7つのADaMデータセットは、CDISCパイロットそのものから直接引用されています。SDTMでは収集されたデータの種類ごとに標準化されたドメインにデータを整理します。今回の研究では、人口統計情報、有害事象、検査値、バイタルサイン、併用薬、曝露量など10種類のデータタイプを収集しました。一方、7つのADaMは、これらの出力が参照する分析準備完了版のサブセットです。
13のTLFについては、唯一、私たちが選定基準を決めた箇所です。実際の研究報告書には数百もの表が含まれますが、ここではICH E3セクション14(「表、図、グラフ」に関する基準)に従い、人口統計情報、有効性、安全性、検査値、リスト、そして1つの図という主要な出力タイプを網羅する代表的な13件を選定しました。これらはすべてCDISCパイロットのADaMデータセットから決定論的に計算されています。
つまり、このパイプラインはCDISCの公開データをエンドツーエンドで活用しており、選定された13件の出力は、実際の提出資料に含まれるべき内容の一部をベンチマークが抽出したものです。
CDISC PILOT01 は、完全な CDISC パイプラインの公開参照実装です。これはアルツハイマー病における Xanomeline とプラセボを比較した 24 週間の第 III 相試験で、被験者数は 254 名、群は 3 つ(プラセボ、低用量 Xanomeline、高用量 Xanomeline)です。主要評価項目には ADAS-Cog (ACTOT) が採用されています。このデータセットを選定した理由は、SDTM、ADaM、TLF の各レイヤーにおいて公開された人間および SAS による参照結果が存在し、各段階で正解となる基準(グランドトゥルース)が明確であるためです。また、モデルに依存しない公開データでもあります。なお、TLF 生成の過程では回答鍵はモデルに対して非開示としました。
各ランニングの評価方法について
すべての実行結果は 3 つの観点から採点されますが、リーダーボードの順位決定に直接影響するのはそのうち 2 つのみです(図を参照)。
インループバリデーター(決定論的、0〜10 点)は、モデルが動作している間、参照データではなく決定論的な構造化チェックに対して出力を繰り返し評価します。このバリデーターはパイプラインが正しく構築されているかを確認し、必要なファイルが存在するか、必須の列が含まれているか、行数が範囲内にあるか(被験者レベル分析用データセットには正確に 254 人の被験者が含まれる)、主要変数が一意であるか、値が適切な統制語彙を使用しているか、導出の不変条件(例:AVAL = BASE + CHG)が満たされているか、各データセットがその下のレイヤーのみから構築されたかなどを検証します。発見された問題の深刻度に応じて、初期スコア 10 点から減点されます(エラーは 1.0 点減、警告は 0.3 点減、情報ノートは 0.1 点減)。各項目には上限があり、最終的なスコアの下限は 1.0 点です。参照解答はモデルに開示されず、モデルは入力とバリデーターからのフィードバックのみを元に動作します。このスコアは、モデルが反復処理を継続するかどうかを判断するための自己フィードバック信号であり、合格・不合格のゲートとしても機能しますが、複数のモデルを順位付けするものではありません。
報告されている 2 つの指標:
データ精度は客観的な指標(0〜1)です。モデルの表、リスト、図に含まれるすべての数値セルを、公式の CDISC リファレンスと比較します。許容誤差は±0.01 です。「平均 (標準偏差)」や最小値〜最大値範囲といった複合セルについては、それぞれの数字が独立してチェックされるようトークン化処理を行います。
ランニングスコアは各ファイルごとの一致率の平均値で算出されます。すべての出力表に同様の重み付けを行うため、モデルは巨大なリスト(約 8.8 万セルと約 1.2 万セル)を正しく出力してスコアを水増しすることはできなくなります。重要なのは、実際に意味を持つ出力である 28 セルの主要有効性表です。
提供された入力から再現できない表、つまり生の CRF に含まれていない中央ラボやベンダーからのフィードが必要となる表は評価対象から除外されます。その結果、13 の出力のうち 10 個が数値に基づいてスコアリングされます。
レビューパネルは主観的な指標(LLM による採点、0〜10)です。Gemini 3.1 Pro、GPT-5.5、Opus 4.8 という 3 つの異なるモデルファミリーから選出された 3 名の批評家が、FDA レビューアーの評価基準に基づき、コード品質、実行可能性、仕様適合性、出力精度という 4 つの軸で各提出物を独立して評価します。また、すべての発見事項には深刻度のタグが付されます。
最終スコアは 3 名による評価の中央値を採用しています。なお、このタスクにおいて GLM 5.2 が精度でトップとなった際も、これら 3 名のモデル批評家はすべてクローズドモデルでした。
「オープンウェイトのLLM(大規模言語モデル)が、精度において追いついた」
画像:Open-Weight LLMs Have Caught Up on Accuracy (21 minute read)
原文を表示
In regulated life-sciences work — screening literature for post-market surveillance, extracting structured data from clinical documents and generating regulatory submissions, and generating tables, listings, and figures (TLFs) based on clinical trial data — a wrong answer isn’t just an inconvenience. A silently missed paper in a systematic review or a fabricated value in a data-extraction table can propagate into a regulatory submission. In the life-sciences context, accuracy matters most given the high stakes.
The prevailing assumption has been that only closed, frontier models can be trusted with tasks like these. We put that assumption to the test — and found it no longer holds.
Within life sciences, most benchmarking attention has gone to discovery. Structure prediction has CASP, the decades-old blind assessment that AlphaFold turned into the field’s canonical measuring stick. LAB-Bench and BixBench probe whether agents can reason over biology literature and run real bioinformatics analyses — BixBench, notably, is hard enough that frontier models scored in the single-digit-to-teens percent range at release. Coscientist showed an LLM planning and executing its own chemistry experiments on lab hardware. That work matters, but it measures the front of the pipeline. The money and the years are at the back: the Congressional Budget Office puts discovery and preclinical work at about 31% of pharmaceutical R&D spend — roughly $474M per approved drug — against about $1.07B for clinical trials, and about 95 months in the clinic versus 31 months preclinical. Other regulated professions have started measuring their own equivalent of that back half. Harvey’sLegal Agent Benchmark scores agents on more than 1,250 real law-firm assignments against expert-written rubrics; OpenAI’s HealthBench and Stanford’s MedHELM cover clinical care tasks; FinanceBench tests analyst questions against public filings. Outside any one vertical, GDPval is the closest to a general “real professional work product” benchmark, but does not cover regulatory submissions. Regulatory and clinical-trial workflows — where a life-sciences program spends most of its time and money — has had no public equivalent. ClinReg is our attempt at one.
“ClinReg” spans three real regulatory and clinical-trial tasks. We ran the leading open-weight and closed-weight models through it. The headline: across all three tasks, open-weight models now fall within the variance band of closed-source models on accuracy. And they do it at a fraction of the cost — the cost gap is significant, with open-weight models running 3x cheaper than the best scoring model we tested, and ~10x cheaper than the most expensive frontier models.
We’ll now walk through what we measured, what surprised us, and what it means for anyone deploying AI in clinical and regulatory settings.
For our customers, moving high-stakes workloads onto open models comes down to three value propositions:
- Accuracy at a fraction of the cost. If open-weight models match frontier on the task that matters, the economics change completely.
- Infrastructure robustness. Frontier models often run into capacity and infrastructure issues. By shifting the load to open-weight models, teams can achieve enhanced robustness and higher uptimes on their workloads.
- Data sovereignty. Beyond accuracy and reliability, the open-weight model value proposition is also about data sovereignty — not leaking data to external providers.
The rest of this post is about proving the first point rigorously, because without it the second and third ones don’t matter.
Details on the ClinReg Benchmark tasks, datasets, evaluations, skills, harness, inference provider and score computation are provided in the Appendix.
We measured the “Overall Score” (see Appendix) across 19 models (12 proprietary and 7 open-weight). We found that GPT 5.6 Sol had the highest overall score (88.4; std dev=4.6) on the ClinReg benchmark, followed closely by open-weight models GLM 5.2 (87.4; std dev=5.6) and Kimi K3 (86.9; std dev=4.6). GLM 5.2 and Kimi K3 were 33.8% and 59.6% of the cost of running GPT 5.6 Sol on average per task. Thus, we see that open-weight models are now well within a standard deviation of the best performing proprietary model. Just a few months ago, the best performing models were Opus 4.8/4.7 which were ~9x the cost of GLM 5.2 on this task, indicative of the fast-pace of cost and accuracy improvements.
Just behind these models, we saw a cluster of proprietary models with similar accuracy but a wide-spread in cost: GPT 5.6 Terra, GPT 5.5, Grok 4.5 and Opus 5 ranging from 84.3-86.3 overall score, but with a 11x cost spread.
The best US open-weight models however were the lowest performers in our list: Gemma 4 (73.5) and Nemotron 3 Ultra (70.3). Within each model family we see that subsequent versions of models led to better accuracy for both proprietary (for e.g. see GPT 5.5 -> GPT 5.6 Sol, or Opus 4-6 -> 4.7 -> 4.8 -> 5) and open-weight models (for e.g. Kimi 2.6 -> 3). Within the GPT 5.6 model family, the overall score (Sol > Terra > Luna) tracked the performance and cost rankings from OpenAI. Similarly the Opus family performed better than Sonnet.
Fable 5 and Muse Spark 1.1 refused our benchmark tasks on biological-safety grounds and were omitted from this analysis. We found that the Opus and Sonnet costs were higher than those of other model families. How much this was impacted by the harness (Claude Code) vs. the model itself is something we plan to study next.**
Literature screening.** Across the literature screening task, performance was tightly clustered at 85–94% (a spread of 40-44 correct articles), suggesting that all evaluated models handled this classification setting well. Accuracy alone, however, obscures meaningful behavioral differences: models made errors on different papers and applied inclusion or exclusion criteria differently. Some models favored recall (e.g. Sonnet 4.6 and Gemma 4), while more conservative models such as MiniMax and DeepSeek excluded more articles. In practice, this makes model selection less about a single headline score and more about whether its error profile matches the review team’s tolerance for missed studies versus additional manual review.
For the IND and TLF tasks, we collected traces of the agentic runs, and analyzed these to understand the error patterns.
IND: restraint mattered more than persistence. Models were broadly reliable at extracting information explicitly present in the source; the larger differences appeared when required fields were absent. The best-performing models (GPT 5.6 family and Kimi K3) balanced coverage with faithfulness, drafting complete documents without filling gaps with plausible but unsupported details. Some models were highly cautious and avoided fabrication but omitted real content (e.g. Gemini 3.1 Pro), while others pursued completeness by inventing unstated values (MiniMax M3). Otherwise capable frontier models (Opus family and Sonnet 5) also fell behind when they committed to specifications or methods that the source did not support.
TLF: iteration separated the models. Every model produced syntactically valid Python on its first attempt; the differentiator was how it responded to runtime errors and used the validation loop. The GPT model family wrote clean code quickly and patched errors in place, but occasionally declared the task complete while fixable issues remained. Others such as Opus 4.8 and Grok 4.5 repeatedly refactored entire scripts, while top open-weight performers such as GLM-5.2 and Kimi K3 continued iterating until their checks passed. At the other end, Gemma and DeepSeek V4 invested many iterations without consistently converging on the correct result, suggesting that persistence cannot fully compensate for capability limits. Under this evaluation setup, performance reflected both model capability and operating style—particularly how models repaired mistakes and decided when the work was complete. These results also highlight the influence of the surrounding harness: stronger validators and more explicit stopping criteria could improve outcomes, especially for models prone to exiting early.
Together, TLF and IND reveal complementary model behaviors: TLF rewards sustained iteration, while IND rewards restraint. Capability puts a model in contention, but task-specific behavior—and the acceptance criteria encoded in the harness—can determine where it ultimately performs best.**
In the process of creating this benchmark and constructing the evaluations (see Appendix), we had several learnings. These are especially relevant for tasks when there isn’t human ground truth – for our ClinReg benchmark this was applicable to the final Module 3 output, and intermediate stages of TLFs such as code gen. We share a few of these observations below.
Don’t trust a single LLM judge. **As has been reported elsewhere, we also found LLM-as-a-judge scoring on a 0–10 scale to be unreliable with different models assigning different scores to the same output. In our own preliminary analysis on the IND task, we found the models when used as critics to be operating on different scales – Gemini 3.1 Pro and GLM 5.2 were lenient (mean score of 8/10), and GPT 5.5 was strict (mean score of 6/10). Even self-preference patterns were uneven: Gemini 3.1 Pro and GLM 5.2 preferred their own output, but GPT 5.5 was harsher on itself.
In order to minimize the impact of these differences, we used a couple of strategies:
- For IND scoring the four judges surface findings rather than scores, and a “chairman” model reviews those findings, assesses severity, and assigns the score.
- For the critic panel in the TLF task, we used the median score from the 3 critics
Use deterministic metrics. For the IND task, we introduced 2 metrics (Omission rate and fabrication rate) related to the intermediate data extraction task (108 fields) which could be objectively measured and scored by the judge models. Omission rate and fabrication rate are fully deterministic — omission rate captures when a model marks a present field as missing to avoid penalty, and fabrication rate is when a model outputs an incorrect value for a present field. Interestingly, these metrics revealed two different failure modes: Opus 4.8 had a higher omission rate but a lower fabrication rate compared to GLM 5.2.
In-loop validation and task decomposition: Along with the evaluation performed at the end of the task, another important factor is in-loop validation – i.e., validation done within the execution of the task. Agentic workflows need strong validators, stopping rules and acceptance criteria defined into their Skills for recursive self-improvement (whereby the output improves iteratively) to work. Defining validators is straightforward when they are deterministic or objectively defined. When the validator is probabilistic or subjective (as in the case of LLM-as-a-judge) similar considerations to what we outlined above apply in order to get the benefit of recursive self-improvement.
In practice, subject matter experts can play a big role in defining the acceptance criteria as they are closest to the details of each task. Decomposing tasks with intermediate sub-tasks within the Skill (for e.g. the intermediate data extraction task in the IND task) has the added benefit of improving accuracy by providing concrete milestones for the agent to self-improve against, rather than having to rely on abstract, high-level feedback.
Tracing the agentic runs is another useful design pattern. Traces can be useful for error analysis, as well as for downstream fine-tuning or reinforcement learning to further improve accuracy and costs.**
- Open-weight models have reached accuracy parity on regulatory work. The best open-weight models in our set — GLM-5.2 (87.4) and Kimi K3 (86.9) — land within one standard deviation of the best closed model we tested (GPT-5.6 Sol, 88.4). Across literature screening, IND Module 3 authoring, and TLF generation, the open-vs-closed line no longer predicts who is at the top.
- The cost gap runs one direction, and it is large. GLM-5.2 reaches that score at roughly a third of GPT-5.6 Sol’s cost per run and about a tenth of Opus 5's. At the volume a real regulatory program runs, that is the difference between a line item and a budget conversation.
- Parity is not uniformity. Models with nearly identical headline scores failed in visibly different ways: some fabricated values to fill gaps while others omitted values; some iterated until validation passed while others declared victory with fixable errors still in place. Two models a point apart on the leaderboard are not interchangeable in production. Choose on error profile, not on rank.
- The harness is part of the model. Every agentic result here is a model and a scaffold — validator strictness, stopping criteria, retry behavior. Under our setup, several results moved more with how a model was run than with which model it was.
- US open-weight models lagged non-US counterparts. The strongest open-weight results came from GLM-5.2 and Kimi K3. The best US open-weight models were the lowest performers in our set — Gemma 4 31B (73.5) and Nemotron 3 Ultra (70.3).
The headline is that open-weight models have caught up. The consequence is that model choice has become the easy part. What’s left is engineering: which model and harness fit the task, how skills and in-loop validation encode acceptance criteria, and which inference provider you route to at what cost and latency. That complexity doesn’t disappear — it moves. Everest and the Log10 team absorb this complexity so life-sciences organizations can focus on the deliverable. We hope standardized benchmarks like ClinReg raise confidence in what AI can do for regulatory and clinical-trial workflows, and ultimately help cut the time and cost of bringing new drugs and devices to market.
*We welcome feedback and additions to our set of ClinReg benchmark tasks, which we plan to open source in the future. Please contact ai@log10.io*
We designed a benchmark suite (“ClinReg”) around three tasks that mirror real regulatory and clinical-trial work using publicly available data:
Literature screening is the first benchmark task. It uses CLEF eHealth TAR topic CD009593, based on a Cochrane diagnostic-accuracy review of Xpert MTB/RIF for pulmonary tuberculosis and rifampicin-resistance detection in adults (link). CLEF TAR repurposes real systematic-review workflows as reproducible screening tasks.
The task mirrors two stages of a systematic review: title/abstract screening determines which records proceed to full-text review; full-text screening determines final inclusion or exclusion and the applicable reason. Screening criteria were human-curated from the source review’s published eligibility criteria and translated into a fixed, ordered instruction before evaluation. Models apply the same instruction; they do not generate or rewrite the review protocol.
Although this is an independent Cochrane review rather than a formal post-market-surveillance programme, its evidence-screening workflow is a realistic, public proxy for the literature-surveillance component of post-market surveillance and diagnostic-device performance evaluation.
We currently report full-text screening results only. The evaluation set contains 47 articles with publicly available full-text PDFs from PubMed Central and other web sources, and final human review labels: 20 included studies and 27 excluded studies. Each model makes one include/exclude decision per article.
Normalized Literature Screening Score measures overall decision accuracy:
Normalized Literature Screening Score = 100 * (#correct inclusions and exclusions) / 47
We also report the total provider cost to screen the 47-article evaluation set, calculated from recorded token usage. The Cochrane review provides human-recorded exclusion reasons for excluded studies, enabling later audit of a model’s rationale.
We model a draft regulatory filing: the Chemistry-Manufacturing-Controls section (i.e., Module 3) of an IND application (to the FDA). This is a proxy for the actual work of authoring an IND. A real submission is drafted from a company’s own proprietary manufacturing and analytical data; to turn it into a benchmark, we swap that private source for a public one that carries similar information. We hand the model a single dense assessment report, the European regulator’s published review (EPAR) of an approved antibody-drug conjugate, and ask it to reconstruct the Chemistry-Manufacturing-Controls section of an IND application, from that source alone. It is not a single-turn question with an answer; it is a sustained, agentic** job the model has to run end to end. A key aspect of the agentic job is extracting 108 fields from the EPAR report and incorporating those data points in the right way for the Module 3 creation.
This task used an LLM-as-a-judge setup to assess how faithfully and completely each model reconstructed the CMC content from the EPAR. Details are below, but in brief: a four-model judge panel finds and cross-verifies evidence of errors, and a single arbitrator turns that evidence into one holistic 0–10 quality score.
Four judges (Gemini 3.1 Pro, GPT-5.5, Opus 4.8, and GLM-5.2) each list two kinds of structured evidence:
- fabrications (populated fields whose specific value isn’t stated in the source, like invented limits or manufacturer identities) and
- omissions (fields the source clearly states but the model marked MISSING or vague).
Standard citations (ICH Q1A, USP <71>) are excluded from fabrications; they’re boilerplate. Fields the source genuinely leaves silent are not omissions; honest gaps are correct. Every candidate is re-verified against the source, and only those confirmed by at least three of the four judges count as correct. A single “Chairman” model then scores the two confirmed lists, reading the source and the model’s output alongside them.
Note on “Chairman” model: We initially chose Fable 5 (for some degree of model family independence), but it declined our prompts due to its biological safety trigger, and routed to Opus 4.8 instead. We acknowledge that this creates the possibility of a self-preference bias for Opus family models. Two things argue against it distorting our results. First, the design: judges emit source-verified evidence rather than scores, and confirmation needs three of four votes, so no single judge can suppress a finding against itself. Second, we measured it. In the preliminary critic analysis above, Opus 4.8 showed the smallest self-preference of the four models tested (Δ = −0.14 on a 0–10 scale, against +1.23 for Gemini 3.1 Pro and +0.67 for GLM-5.2) — it was, empirically, the least self-favoring judge available to us.
Normalized IND score (out of 100) is calculated by multiplying the 0-10 score by 10. Each model was run 3 times and the mean score across the 3 runs was used.
Every drug approval at the FDA rests on a specific technical dossier built around the *trial’s tables, listings, and figures (TLF)*: the pre-specified analyses showing who was treated with what, which adverse events happened at what rate, and whether the primary efficacy endpoint moved the way the sponsor claimed. The FDA doesn’t work from the raw case-report forms. It reviews the standardized datasets behind every reported number and re-runs the analysis programs that produced them. The sponsor submits that whole traceable stack, the analysis-ready datasets, the standardized tabulations they were derived from, and the programs that generated all of it, packaged into an electronic submission a reviewer can execute end to end to reproduce any headline result.
On every clinical trial, a team of biostatistical programmers spends weeks to months writing that code. They take the raw case report form (CRF) data collected at the clinic and push it through two standardized data models, CDISC’s Standard Data Tabulation Model (SDTM) and the Analysis Data Model (ADaM), finally into the specific tables, listings, and figures the reviewer reads. A small derivation error at any stage can silently propagate through the entire pipeline, leaving downstream results wrong but superficially consistent. Examples include a last observation carried forward (LOCF) rule that carries a value beyond its permitted visit window, an analysis population flag derived off by one row, or a controlled vocabulary decoded against a stale version.
This benchmark compresses that end to end workflow into a single autonomous task. Given raw case report form (CRF) and its specifications, each model must autonomously produce a complete CDISC submission – from raw CRF through SDTM and ADaM to TLFs.
Specifically, in a single autonomous turn, the model writes and executes three Python scripts:
- raw_to_sdtm.py transforms the raw CRF data into 10 standardized SDTM domains, including DM, AE, LB, and VS.
- sdtm_to_adam.py derives 7 analysis-ready ADaM datasets from the SDTM domains.
- adam_to_tlf.py generates 13 TLFs from the ADaM datasets.
*The four data layers of the pipeline and the three Python scripts the model writes between them, inside one autonomous refinement loop.*
The 10 SDTM domains and 7 ADaM datasets are set by the CDISC Pilot dataset (the 7 being the analysis-ready subset our outputs draw on). The 13 TLFs are the benchmark’s chosen subset of the tables, listings, and figures that would appear in a full study report.The 10 SDTM domains and 7 ADaM datasets come straight from the CDISC Pilot itself: SDTM organizes a trial’s data into one standardized domain per type of data collected, and this study collected ten — demographics, adverse events, labs, vitals, con-meds, exposure, and so on — while the 7 ADaM are the analysis-ready subset those outputs draw on. The 13 TLFs are the one place we made a choice. A real study report contains hundreds of tables, so we selected a representative set of thirteen spanning the core output types — demographics, efficacy, safety, labs, listings, and a figure — following ICH E3 Section 14 standards on Tables, Figures and Graphs and computing each deterministically from the Pilot's ADaM. So the pipeline is grounded in the public CDISC data end to end; the 13 outputs are the benchmark’s chosen slice of what a full submission would contain.
CDISC PILOT01, the public reference implementation of a full CDISC pipeline: a 24-week Phase III trial of Xanomeline vs placebo in Alzheimer’s, 254 subjects, 3 arms (Placebo, Xanomeline Low/High Dose), ADAS-Cog (ACTOT) primary endpoint. Chosen because a published human/SAS reference exists at every layer (SDTM, ADaM, TLF), giving ground truth at each stage, and it is public and model-neutral. Answer key was withheld from the models during TLF generation.
How we score each run
Every run is graded three ways, but only two of them decide the leaderboard (see figure).
An in-loop validator (deterministic, 0–10). While the model works, it repeatedly scores its own output against a deterministic structural check — never against the reference. The validator confirms the pipeline is well-formed: expected files exist, required columns are present, row counts are in range (the subject-level analysis dataset has exactly 254 subjects), key variables are unique, values use the correct controlled vocabulary, derivation invariants hold (e.g. AVAL = BASE + CHG), and each dataset was built only from the layer beneath it. Findings lower the score from a starting 10 in proportion to severity — −1.0 per error, −0.3 per warning, −0.1 per information note, each capped — with a floor of 1.0. The reference answer key is withheld from the model throughout; it works from its inputs and the validator’s feedback. This score is the model’s own feedback signal (how it decides whether to keep iterating) and a pass/fail gate; it does not rank models.
The two reported metrics:
Data accuracy — the objective one (cell-match, 0–1). Every numeric cell in the model’s tables, listings, and figures is compared to the official CDISC reference at a ±0.01 tolerance; composite cells like “mean (SD)” or a min–max range are tokenized so each number is checked independently. The run’s score is the mean of the per-file match rates, weighting every output table equally — so a model can’t inflate its score by nailing the two huge listings (≈88k and ≈12k cells) while missing the 28-cell primary efficacy table, which is the output that actually matters. Tables that can’t be reproduced from the provided inputs —those needing central-lab or vendor feeds absent from the raw CRFs — are excluded, leaving 10 of the 13 outputs scored on their values.
Reviewer panel — the subjective one (LLM-graded, 0–10). Three critics from three model families (Gemini 3.1 Pro, GPT-5.5, Opus 4.8) independently grade each submission against an FDA-reviewer rubric across four axes — code quality, executability, spec conformance, output correctness — and tag every finding by severity. We take the median of the three. Note that all 3 model critics were closed models even as GLM 5.2 topped accuracy on this task.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み