2026 年における AI エージェント最適化プラットフォームの選定ガイド
本文の状態
日本語全文を表示中
詳細モードで約24分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Pydantic Blog
Pydantic Blog は2026年のAIエージェント最適化プラットフォーム市場における買収と統合の動向を分析し、観測から改善へのループを閉じる能力を持つツールが重要であると指摘する。
AI深層分析を開く2026年8月4日 00:19
AI深層分析
キーポイント
市場の統合と買収動向
2026年はツールリングの統合年であり、LangfuseのClickHouseによる買収やOpenAIによるPromptfoo取得など、主要プレイヤーが相次いでいる。
最適化プラットフォームの定義基準
単なるダッシュボード表示ではなく、生産環境のトレースに基づいて変更を提案し、安全なロールアウトと即時のデプロイを実現する機能が不可欠である。
評価ルールの見直し
スコアごとの課金制限なくオンライン・オフライン両方の評価を行い、データポータビリティを確保するためのオープン標準(OpenTelemetry)の採用が推奨される。
Pydantic Logfire の包括的な最適化機能
分散トレーシング全体を一つのプラットフォームで統合し、問題がエージェント自体かインフラの欠陥かを特定して修正を提案する。バージョン管理された設定とラベルによるロールバックで、診断から修正の実行までを完結させる。
Braintrust の評価駆動型アプローチ
自然言語での目標記述によりスコアラーの草案作成や評価データセットの構築を行い、本番環境の失敗事例を恒久的な評価ケースに変換する。
重要な引用
The gap between those two verbs is the whole category this post is about.
A platform that optimizes agents does the next thing: it tells you whether the fault is the agent or something else in your stack, proposes a specific change, and gives you a safe way to ship that change and watch the result.
You optimize the thing that is actually broken, instead of tuning a prompt against an infrastructure bug for a week.
Customers cite real outcomes, including Notion's team going from three to thirty issues fixed a day.
編集コメントを表示
編集コメント
記事は2026年という未来の視点から市場統合を予測・分析しており、単なるツールの列挙ではなく、開発者がツール選定時に重視すべき本質的な価値基準(ループの閉鎖性)を提示している。特にOpenAIが自社製品を停止する動きや、買収によるサービス終了リスクは、データ所有権を重視する組織にとって重要な示唆となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
「AI エージェント」プラットフォームの多くは、エージェントを監視するに留まっています。それを変化させることができるものは、さらに少ないのです。この二つの動詞の間にこそ、本稿が扱うカテゴリ全体があります。
エージェントはリリースされ、稼働しますが、その回答の一部には、評価スイートでは予測もできなかった形で誤りが生じます。ダッシュボードはスコアが低下したことを告げるだけですが、優れたプラットフォームは、実際の生産環境でのトレースに基づいて「なぜそうなったのか」を説明します。そして、エージェントを最適化するプラットフォームはさらに一歩進み、問題の原因がエージェント自体にあるのか、それともスタック内の他の要素にあるのかを特定し、具体的な改善策を提案します。さらに、その変更を安全にリリースし、結果を確認できる方法も提供します。つまり、「観察」し、「評価」した上で「改善」するのです。ここでいう改善とは、単なる Jira チケットの作成ではなく、実際にデプロイして実行に移すことを意味します。
2026 年は、この分野のツールリングにおいて統合が進んだ年となりました。Langfuse は ClickHouse に買収され、OpenAI は Promptfoo を取得し、11 月 30 日に自社で提供していたホスト型 Evals プロダクトを廃止しました。Cisco は Galileo の買収に動き、Helicone は Mintlify に統合されてメンテナンスモードに移行しています。一方、最も資金調達に成功した独立系ベンダーである Braintrust は、独自に 8000 万ドルの資金調達を果たしました。
現在、私たちが問うべきは単一の質問ではありません。二つの重要な問いがあります。「そのプラットフォームは改善のループを完結させるのか、それとも監視だけなのか」、そして「誰がまだ独立しており、このループとあなたのデータがあなた自身のものとして保たれているのか」です。
エージェント最適化プラットフォームと観測ツールの違い
これらの基準は、ループを閉じる機能を重視して設定しました。なぜなら、それが本質的なカテゴリだからです。単にダッシュボードが必要であれば、以下のツールのほとんどが対応可能です。その場合は、当社の AI 観測性比較記事の方がより有益な情報となるでしょう。
実際にプロダクション環境でエージェントを改善するためには、以下の点が重要です。
ループを閉じることです。プラットフォームは、実運用のトレースに基づいて変更案を提示し、バージョン管理、ターゲティング、ロールバック機能を備えた管理された設定を通じて変更を実行します。多くのツールはここで止まり、ダッシュボードや評価レポートで終わってしまいます。
エージェント全体を見渡せることです。エージェントの失敗は、 retriever(検索)、データベース、ツール、インフラストラクチャの障害が原因で、結果として不適切な回答として現れることがよくあります。LLM のスパンだけを記録するプラットフォームでは、証拠の半分しか見ておらず、「プロンプト自体に問題がなかった」という判断もできません。
評価を制限しないことです。スコアごとの計測機能を持たないオンラインおよびオフライン評価により、カバレッジは品質上の判断基準となり、請求額による制約にはなりません。
オープンスタンダードと独立性です。OpenTelemetry ネイティブなデータ取り込みとポータブルなデータを採用することで、自社のインストゥルメンテーションが特定のベンダーのロードマップ、買収、またはサービス終了に左右されることを防ぎます。
安全なロールアウトです。変更不可バージョン、コホートターゲティング、加重キャノリー展開、ラベルの変更によるロールバック機能により、ライブのエージェントへの更新はいつでも元に戻せます。
マーケティングの喧騒を切り裂く一つの指針があります。どの名詞を対象に最適化しているかを見極めることです。一部のプラットフォームは「評価指標」そのものを最適化し、スコアラーや審査員を調整します。これは測定対象ではなく、定規を研ぐ行為に過ぎません。例えば LangSmith のアライメントツールも、アプリ自体ではなく評価器のチューニングに焦点を当てています。エージェントの最適化とは、本質的にプロダクション環境で動作するものを変更することであり、単に得られる点数を厳格化するものではありません。
各プラットフォームの評価順位
#1. Pydantic Logfire:エージェント最適化において総合最高
Pydantic Validation や Pydantic AI を開発したチームが提供する「Pydantic Logfire」は、最も完結したループを実現するプラットフォームです。ブラウザ、サービス、データベース、モデル呼び出し、ツール呼び出し、インフラストラクチャに至るまで、分散トレーシング全体を 1 つの OpenTelemetry ネイティブ製品に統合します。評価スコアも同じタイムライン上に紐付けられるため、一貫した可視性が得られます。
そのオプティマイザーは、評価が低かったランと、その周辺にある数千のランを読み込み、パターンを見つけ、根拠に基づいた変更案を提示します。最も重要なのは、修正よりも前の段階です。モデルのスパンだけでなく全体のトレースを見渡せるため、問題がエージェント自体にあるのか、検索が遅いのか、レート制限がかかる上位システムがあるのか、あるいはデータベースのタイムアウトが悪回答として現れているのかを特定できます。これにより、インフラストラクチャの不具合に対してプロンプトを数週間微調整するのではなく、実際に壊れている部分を最適化できます。
その後、管理された設定を通じて変更が展開されます。エージェントの指示、モデル、設定、ツール定義は、ターゲティングと加重ロールアウト機能を備えたバージョン管理された設定として保存され、ロールバックもラベルの変更だけで行えます。これが自己改善型エージェントに求められる本質的なプロセスです。診断し、提案し、展開し、ロールバックする。この言葉がスローガンで終わらないのは、一つのプラットフォームが全体のトレースにおいてこれら4 つの機能をすべて実行できるからです。
既存のトレースをそのまま評価に活用できます。オンライン・オフライン問わず、スコアリング専用の計測器は不要です。スコアは単なる観測値であり、スパーンと同じフラットレートで課金されます。また、無料枠も手厚く用意されています。
OpenTelemetry ネイティブであるため、OTel をサポートするあらゆるフレームワークが同じ機能を即座に利用可能になります。Pydantic AI も最初から組み込まれています。コーディングエージェントは PostgreSQL 互換の SQL で MCP サーバーを通じてトレース全体を照会し、得られた情報に基づいてアクションを実行できます。
2026 年、Langfuse が ClickHouse を採用し、Promptfoo が OpenAI と提携する中、Logfire は独立したオープンな選択肢として残ります。いずれの機能も吸収されることも、終了することもありません。
正直に申し上げますと、評価機能のみを求め、予算が問題でない場合、Braintrust の方がエッジケースでのエージェント評価を行うチームにはより拡張性の高いプラットフォームです。
向いているのは:何を修正すべきかを知り、修正をリリースし、フルスタックのコンテキストを一つの場所で管理したいエンジニアリングチームです。オープンスタンダードに基づいています。
#2. Braintrust: 評価駆動型チームに最適
Braintrust は最も確立された評価ファーストのプラットフォームであり、Loop は実用的なオプティマイザーです。自然言語で目標を記述するだけで、スコアリングルールをドラフト作成し、ログから評価データセットを構築します。また、本番環境での障害を恒久的な評価ケースに変換します。
顧客からは具体的な成果が報告されています。例えば Notion のチームでは、1 日に修正される問題の数が 3 から 30 に増加しました。さらに、修正の実装も可能です。プロンプトレジストリはランタイム時に取得され、コードデプロイを伴わずに新しいバージョンへプロモートされます。ただし、評価に合格していることが条件となります。
制限事項は、スコープの狭さとオープン性の欠如です。Braintrust はプロンプトと LLM の出力を最適化しますが、エージェント全体の構成管理は行いません。ドキュメント自体が「LLM スパンに焦点を当てている」と明記している通り、実際の障害がプロンプトではなく検索ステップやデータベースにある場合でも、それを特定することはできません。
proprietary データストア上で動作し、セルフホスティングは Enterprise プランのみ対応です。さらに評価そのものにも従量課金(メーター)を導入しており、プランに応じて 1,000 スコアあたり 1.50 ドルから 2.50 ドルが請求されます。これは LLM-as-judge がスコアを付与する際に既に発生しているモデルトークンのコストに加え、処理されたデータ量(ギガバイト単位)にも課金されるものです。
本番環境のスケールでは、この「スコアごとの課金」が主要なコスト要因となります。月間 300 万回の評価実行を Pro プランで行う場合、スコア関連費用だけで約 4,400 ドルになります。一方、Logfire では同じ 300 万回の実行は単なる観測データとして扱われ、数ドルで済み、スコア課金自体が存在しません。
Braintrust は AI の出力に関するループを閉じる点では優れていますが、その範囲は限定的です。システムの残りの部分を支える可観測性スタックとは別に、AI 出力に特化した狭い領域での運用となります。
こんなチームにおすすめ: 評価カバレッジや CI/CD の回帰テストがボトルネックとなっているチーム
料金プラン: フリーティアムあり;Pro プランは月額 249 ドル(Braintrust のドキュメントによると、2026 年 9 月には 100 ドルに引き下げられる見込み);Enterprise はカスタム見積もり。
詳細比較:Logfire vs Braintrust
#3. Arize AX: OpenInference に完全移行しているチームに最適
Arize は AX を「自己改善型エージェント」のためのプラットフォームとして位置付けていますが、真の自己改善型エージェントとは、トレース全体にわたってフィードバックループを完結させるものです。その点で Logfire の方が上位に位置します。Arize が実際に強みを持つのは、Phoenix コンポーネントによる OpenInference ネイティブなトレーシングと、ドリフト検出やモデルパフォーマンス指標といった深掘りされた ML モニタリングの系譜です。
同社のプロンプト最適化機能も本物です。評価フィードバックから改善されたプロンプトを生成するメタプロンプティングを行い、それを Prompt Hub でバージョン管理し、数回のクリックで展開できます。
一方で課題となるのは、スコープの狭さ、ベンダーロックイン、そしてコストです。Arize は ML モニタリングの世界出身であり、アプリケーショントレースよりもモデル指標を重視する考え方を持っています。Phoenix は OSI のオープンソースライセンスではなく、Elastic License に基づくソースコード公開版です。また、本番環境での最適化機能は商用製品 AX に含まれており、その課金体系はスパン数とペイロードのギガバイト数という二軸方式で構成されています。RAG ワークロードではコストが急増し、Logfire の 100 万トークンあたり一律$2 という定額モデルよりも遥かに高額になります。さらに、Phoenix は自己ホスト環境において、バージョンごとにアップグレードが破綻する歴史があり、その詳細は Arize 自身の移行ノートに記録されています。
向いているのは:OpenInference と Phoenix に標準化されたチーム、あるいはドリフトやモデルパフォーマンス指標を重視する ML チームです。料金体系:無料枠あり;Pro は月額$50 から;Enterprise はカスタム見積もり。詳細比較は「Logfire vs Arize」をご覧ください。
#4. LangSmith: LangChain および LangGraph スタックに最適
LangSmith は、LangChain や LangGraph との統合が最も密接で、強力なトレーシングと評価機能を提供します。また、プロンプトハブでは、完全なデプロイを行わずにプロンプトのバージョン管理や更新が可能です。ただし、その最適化機能は主にプレイグラウンドのアシスタントや外部のレシピ集といった位置づけであり、製品内に組み込まれた最適化ツールというわけではありません。最近追加されたアライメントツールも、アプリケーション自体を調整するものではなく、評価器のチューニングに焦点を当てています。
表面上はフレームワークに依存しませんが、その実力は LangChain や LangGraph のエコシステム内において最も発揮されます。また、クローズドソースであり、セルフホスティングは Enterprise プラン限定です。料金体系はユーザー数とトレーシング数に応じた従量課金で、スケールするにつれてコストが膨らみ、チームは請求額を抑えるためにトレーシングデータを大幅にサンプリングせざるを得なくなります。これは、観測可能性(オバザビリティ)本来の目的とは逆の事態です。
データの保持期間はデフォルトで 14 日ですが、400 日間保持しようとすると、1 トレーシングあたりのコストが約 10 倍になります。この上限は、自社サーバーで運用している場合でも適用されます。
最も適しているのは:LangChain と LangGraph にコミットするチームです。
料金プラン:無料の Developer タイア、Plus は月額 39 ドル/ユーザーから、Enterprise はカスタム見積もり。
詳細比較:Logfire vs LangSmith。
#5. Langfuse: オープンソース版として最適
Langfuse は、トレーシング、プロンプト管理、データセット、評価機能を備えた広範なオープンソースプラットフォームです。MIT ライセンスの下で提供され、AI 観測ツールの予算がなく、自社ホスト・自社管理を望むチームにとってデファクトスタンダードとなっています(2026 年 1 月には ClickHouse が Langfuse を買収しましたが、引き続きオープンソースかつ自己ホスティング可能です)。
プロンプトの管理や実験実行、教科書的なバージョン管理、ラベルベースのロールアウトに対応していますが、トレーシングに基づく修正を提案するのではなく、観測とスコアリングに留まります。つまり、最適化ループの半分は自分自身で組み立てる必要があります。自己ホスティングには独自の再帰的なコストが伴います。Postgres、ClickHouse、Redis、オブジェクトストレージを運用することになり、これら自体も監視が必要となるため、結局「観測ツールを監視するための観測ソリューション」が必要になるという皮肉な状況に陥ります。また、重要なガバナンス機能(プロジェクトレベルの RBAC、SCIM、監査ログ)は、自己ホスティングの場合でも商用ライセンスキーが必要です。
こんなチームにおすすめ:ホームラボでインフラを動かしたい開発者や、無料で自己ホスティング可能なプラットフォームが必要で、その上に独自の最適化機能を構築するチーム向け。料金プラン:無料の自己ホスティング版、クラウド版は月額 29 ドルから。詳細比較:Logfire vs Langfuse。
#6. 評価特化型ツール:DeepEval、Promptfoo、Patronus
DeepEval、Promptfoo、Patronus は、リリース前の評価フェーズにおいて非常に優れたツールです。DeepEval は pytest スタイルの LLM-as-judge テストと、本番環境のトレースではなくゴールドデータに対してチューニングを行うプロンプト最適化機能を備えています。Promptfoo は設定ファイル駆動型のローカル評価とレッドチーム演習を提供します。Patronus は幻覚検出やエージェントのストレステストに特化した専用Judge モデルを扱います。
これらはあくまでテスト・評価フレームワークであり、本番環境での最適化プラットフォームではありません。実行前後の問題点を指摘するものであり、稼働中のエージェント設定を管理したり、修正を本番へ直接反映させたりする機能は持ち合わせていません。2026 年の市場再編について注目に値するのは、Promptfoo が OpenAI に買収されたことと、OpenAI が提供するホスト型 Evals サービスが 11 月 30 日に終了することです(ただし、オープンソースの openai/evals リポジトリには影響ありません)。
推奨される活用シーン:エージェントループを処理するプラットフォームよりも上位の CI パイプラインに、厳格な評価とレッドチーム演習のステップを組み込む場合。
#7. Elastic:LLM ダッシュボードを備えた検索企業
Elastic は検索およびログ管理企業ですが、その APM 製品に LLM の可観測性を付加しています。このツールはエージェントのプロンプトやレスポンス、トークン数、レイテンシを表示し、「AI エージェントの最適化」を謳っています。しかし、その主張の根拠を詳しく見ると、実はあっさりとしてしまいます。Elastic が語る「最適化」という話は、エンジニアが手動でプロンプトを書き換えて社内ワークロードのコストを削減したという、短いエンジニアブログシリーズに過ぎません。製品機能としての最適化ではありません。2026 年初頭に登場した Agent Builder は Elasticsearch データ上で検索エージェントを構築するものであり、既存のものを改善するものではありません。また、LLM の計測機能は、エージェントの実行フローに合わせて設計されたトレーシングではなく、検索エンジン上に最近追加された統合レイヤーに過ぎません。
さらに Elastic は運用自体が重く、クエリ DSL、インデックスライフサイクル管理、クラスターチューニングなど、適切に運用するには「Elastic に関する博士号が必要」なほど難易度が高いという評判があります。この評判は、設定ミスを防ぐために AutoOps をリリースしたことで、Elastic 側も半分認めたようなものです。
つまり Elastic はプロンプトを表示するだけで、管理機能はありません。バージョン管理もターゲティングもロールアウトやロールバックもなく、修正を提案する最適化機能も、失敗した実行を本番変更へと導く仕組みも存在しません。「エージェント最適化」という言葉の隣に置かれているのは、単なる可観測性のビューに過ぎず、記事が製品ではないため、ここでは最下位としました。
向いているチーム:Elastic 環境に深く浸透しており、LLM の呼び出しを他のログと同じダッシュボードで確認したいチーム。ただし、実際の最適化作業は別の場所で行う必要があります。
#Comparison at a glance
プラットフォーム
ループを完結させるか(提案+実装)?
トレース範囲
オープン標準
評価/スコアリングと価格設定
最も適している用途
Pydantic Logfire
はい:障害診断、オプティマイザーによる提案、管理された構成の実装
分散全体でのトレース
OTel ネイティブで独立した仕組み
スコア計測器なし、フラットなスパン
本番環境全体のコンテキストにおけるエージェント最適化
Braintrust
はい(プロンプトの場合のみ:ループ+レジストリ)
LLM スパン(OTel 取り込み対応)
独自データストア(Brainstore)
スコア単位課金(1,000 トークンあたり$1.50〜2.50)+GB 単位のストレージ課金
評価駆動型の CI/CD チーム向け
Arize AX
はい:メタプロンプトオプティマイザーと Prompt Hub を提供
LLM/エージェントスパン
Phoenix ELv2、AX は独自仕様
二軸評価:スパン数+GB 単位のペイロード課金
エンタープライズ ML チームおよび LLM チーム向け
LangSmith
部分的:プロンプトハブとアシスタント機能を提供
エージェント/LLM トレース
クローズドシステム
ユーザー数(シート)+トレース量ベースの課金
LangChain および LangGraph を活用するチーム向け
Langfuse
組み込みオプティマイザーなし
LLM トレース(OTel 互換)
オープンソース(MIT ライセンス)
スコアを含むユニット課金
セルフホスト志向、OSS ファーストのチーム向け
評価専門ツール
いいえ:実装前のテストとスコアリングに特化
本番環境ではなくテスト時実行
ほぼオープンソース
料金体系は多様
厳格な CI 評価およびレッドチーム演習向け
Elastic
いいえ:プロンプト表示機能はあるが管理機能なし
検索エンジンに追加された LLM ダッシュボード
OTel 対応だが、LLM レイヤーは後付け
該当せず
既存の Elastic ユーザー向け
選び方
ボトルネックから始めましょう。エージェントの失敗理由が見えない場合、あるいはその原因がエージェント自体にあるのかインフラ側にあるのか判断できない場合は、まず全体トレーサビリティ(トレース)の可視化が必要です。この段階では、LLM 専用ツールよりもフルスタック型の選択肢の方が優れています。
一方、失敗は把握できるものの、デプロイサイクルを経ずに修正を本番環境に反映させる手段がない場合は、フィードバックループを完結できるプラットフォームが求められます。Logfire、Braintrust の Loop、Arize AX がまさにこの領域で競合しています。
自社のインフラを完全に管理したい場合は、Langfuse やセルフホスト型の Phoenix といったオープンソースから始めるのが良いでしょう。本番リリース前の回帰テストやレッドチーム(攻撃シミュレーション)が課題であれば、どのプラットフォームを生産ループに採用するかに関わらず、評価専門ツールを CI に追加してください。
オープンで多言語対応のスタック上で、実稼働中のエージェントを改善する大多数のチームにとって、Pydantic Logfire が最も強力な出発点となります。これは、スコアごとの計測や特定のベンダーの独自形式への依存を一切必要とせず、全体トレーサビリティを通じて「何を修正すべきか」を示し、「どのように修正するか」を提案し、実際に本番に反映させるまでを一貫して支援するツールです。
よくある質問
AI エージェント最適化とは何ですか?
失敗した本番環境の運用を改善されたリリースへと転換させるのは、このループです。トレーシングデータを収集し、評価を行い、問題がエージェント自体にあるのか周囲のシステムにあるのかを特定します。そして、エージェントのプロンプトや設定に対する具体的な変更案を提示し、安全なロールアウトでその変更を実装、結果を測定します。
最適化(Optimization)は、単に何が起こったかを示すだけにとどまる可観測性(Observability)とは異なります。
エージェント最適化における Braintrust の代替として最適なものはどれでしょうか?
Pydantic Logfire が有力です。Braintrust の Loop は強力であり、プロンプトの変更を配信する機能も備えていますが、LLM の出力を Braintrust 内で最適化する点に留まります。一方、Logfire は分散トレーシング全体を対象とするため、真の問題がプロンプトではなく、データ取得(retrieval)やデータベースにある場合でも特定できます。また、エージェントの構成全体を管理されたロールアウトで配信し、OpenTelemetry を基盤としながら、スコアごとの計測器を必要としない点も特徴です。
Elastic は AI エージェント最適化に対応しているのでしょうか?
実のところ、あまり対応していません。Elastic はエージェントのプロンプトやトレーシングデータを表示し、「AI エージェント最適化」を謳っていますが、その最適化機能は製品機能というより、エンジニアがコスト削減のために手動でプロンプトを書き換えるための数編の技術ブログ記事に過ぎません。Elastic はプロンプトを表示するだけで、バージョン管理やターゲティング、ロールアウト、ロールバックを行うことはなく、修正を提案する最適化エンジンも備えていません。これはエージェントの可観測性を提供するビューであり、エージェント自体を改善するためのプラットフォームではありません。
オープンソースで利用可能なエージェント最適化の選択肢はあるのでしょうか?
Langfuse は MIT ライセンスでセルフホスト可能、Arize Phoenix も Elastic License のオープンソースです。どちらも管理型の最適化ツールではなく、観測性と評価機能を提供し、それを基盤にループを構築できる環境を与えます。Pydantic Logfire は OpenTelemetry ネイティブなので、どこで実行しても計装コードの移植性が保たれます。
2026 年の AI 評価ツールの状況はどうなったのでしょうか?
業界は急速に再編されました。1 月に ClickHouse が Langfuse を買収し、3 月には OpenAI が Promptfoo を取得すると同時に、自社が提供するホスト型 Evals プロダクトを 11 月 30 日に終了することを発表しました。また Cisco は Galileo の買収に動いています。GitHub 上のオープンソースプロジェクトである openai/evals フレームワーク自体は影響を受けませんが、重要な点はこの傾向です。次回の買収が誰であっても、ループとデータがあなたのものであるために、OpenTelemetry ネイティブで独立したプラットフォームを基盤として構築すべきなのです。
#Pydantic Logfire を無料で試す
数分で、エージェントのトレース全体、オンライン・オフラインの評価、そして最適化ループを実行できるようになります。無料プランでは月間 1000 万スパンを提供します。
Pydantic Logfire で無料開始
AI は依然としてエンジニアリングです。
原文を表示
Most "AI agent" platforms watch your agent. Fewer of them change it. The gap between those two verbs is the whole category this post is about.
An agent ships, it runs, and some fraction of its answers are wrong in ways your eval suite never predicted. A dashboard tells you the score dropped. A good platform tells you why, grounded in the actual production traces. A platform that optimizes agents does the next thing: it tells you whether the fault is the agent or something else in your stack, proposes a specific change, and gives you a safe way to ship that change and watch the result. Observe, evaluate, and then improve, where improve means a deploy, not a Jira ticket.
2026 has been a consolidation year for this tooling. Langfuse was acquired by ClickHouse, OpenAI acquired Promptfoo and is shutting down its own hosted Evals product on November 30, Cisco moved to acquire Galileo, and Helicone was folded into Mintlify and put in maintenance mode. Braintrust, the best-funded holdout, raised $80 million of its own. So there are two honest questions now, not one: does the platform close the loop or just watch, and who is still independent enough that the loop, and your data, stay yours?
#What separates an agent-optimization platform from an observability tool
We wrote these criteria to favor closing the loop, because that is the category. If all you need is a dashboard, most of the tools below will do, and our own AI observability comparison is the better read. For actually improving an agent in production, these are what matter:
It closes the loop. The platform proposes a change grounded in production traces, and ships it through managed configuration with versioning, targeting, and rollback. A dashboard or an eval report is where most tools stop.
It sees the whole agent. An agent failure is often a retrieval, database, tool, or infrastructure failure that shows up as a bad answer. A platform that only stores LLM spans is optimizing against half the evidence, and cannot tell you when the prompt was never the problem.
Evaluation you do not ration. Online and offline evals without a per-score meter, so coverage is a quality decision rather than a billing one.
Open standards and independence. OpenTelemetry-native ingestion and portable data, so your instrumentation is not hostage to one vendor's roadmap, acquisition, or shutdown.
Safe rollout. Immutable versions, cohort targeting, weighted canaries, and rollback by moving a label, so shipping a change to a live agent is reversible.
One tell cuts through the marketing: watch which noun a platform optimizes. Some optimize your evals, the scorers and the judges, which is sharpening the ruler rather than the thing it measures. LangSmith's alignment tooling, for one, tunes the evaluator, not the app. Optimizing an agent means changing what runs in production, not tightening the grade it gets.
#The platforms, ranked
#1. Pydantic Logfire: best overall for agent optimization
Pydantic Logfire, from the team behind Pydantic Validation and Pydantic AI, is the platform that most completely closes the loop. It keeps the whole distributed trace in one OpenTelemetry-native product: browser, services, databases, model calls, tool calls, and infrastructure, with evaluation scores attached to the same timeline.
Its optimizer reads the runs that scored badly and the thousands around them, finds the pattern, and proposes one evidence-cited change. The part that matters most is upstream of the fix: because it sees the whole trace and not just the model spans, it tells you whether the problem is the agent at all, or a slow retrieval, a rate-limited upstream, or a database timeout showing up as a bad answer. You optimize the thing that is actually broken, instead of tuning a prompt against an infrastructure bug for a week. Then managed configuration ships the change: an agent's instructions, model, settings, and tool definitions live as versioned config with targeting and weighted rollout, and rollback is moving a label. That is what a self-improving agent actually takes: diagnose, propose, ship, roll back. The phrase is a slogan until one platform does all four over the whole trace.
Evaluation runs on the same traces you already emit, online and offline, with no separate per-score meter: a score is an observation, billed at the same flat rate as any span, with a generous free tier. Because it is OpenTelemetry-native, any framework that speaks OTel lights up the same surfaces, and Pydantic AI is wired in out of the box. Your coding agent can query the whole trace over the MCP server in PostgreSQL-compatible SQL and act on what it finds. And in a year when Langfuse went to ClickHouse and Promptfoo went to OpenAI, Logfire is the independent, open option: nothing here is getting absorbed or sunset.
Honest limitation: if all you want is evals and budget is no concern, Braintrust is the more extensible platform for teams on the bleeding edge of agent evals.
Best for: engineering teams optimizing production agents who want to know what to fix, ship the fix, and keep the full-stack context in one place, on open standards.
#2. Braintrust: best for eval-driven teams
Braintrust is the most established eval-first platform, and Loop is a real optimizer: describe a goal in natural language and it drafts scorers, builds eval datasets from your logs, and turns production failures into permanent eval cases. Customers cite real outcomes, including Notion's team going from three to thirty issues fixed a day. And it does ship the fix: a prompt registry fetched at runtime promotes a new version without a code deploy, gated on passing evals.
The limits are scope and openness. Braintrust optimizes the prompt and the LLM output; it does not manage the agent's whole configuration, and its own documentation says it focuses on LLM spans, so it cannot tell you when the real fault was the retrieval step or the database rather than the prompt. It runs on a proprietary datastore, self-hosting is Enterprise-only, and it puts a meter on evaluation itself: $1.50 to $2.50 per thousand scores depending on tier, charged on top of the model tokens each LLM-as-judge score already burns, plus processed data by the gigabyte. At production scale that per-score meter dominates. Three million scored runs a month is roughly $4,400 in score charges on the Pro plan, where the same three million on Logfire are just observations, a few dollars, with no score line at all. It closes a loop, a narrower one, on the AI output, beside the observability stack that holds the rest of your system.
Best for: teams whose bottleneck is eval coverage and CI/CD regression gates. Pricing: free tier; Pro at $249/month (Braintrust's docs say this drops to $100 in September 2026); Enterprise custom. Full breakdown: Logfire vs Braintrust.
#3. Arize AX: best if you are all-in on OpenInference
Arize markets AX as the platform for "self-improving agents," but a self-improving agent is one that closes the loop over the whole trace, which is exactly why Logfire sits above it here. Where Arize actually leads is elsewhere: OpenInference-native tracing through its Phoenix component, and a deep ML-monitoring lineage of drift detection and model-performance metrics. Its prompt-optimization is real: meta-prompting that generates an improved prompt from eval feedback, versions it in a Prompt Hub, and promotes it in a few clicks.
The tradeoffs are scope, lock-in, and cost. Arize comes from the ML-monitoring world and thinks in model metrics more than application traces; Phoenix is source-available under the Elastic License, not OSI open source; and the production optimization lives in the commercial AX product, whose dual-axis billing, charged per span and per gigabyte of payload, climbs fast on RAG workloads and runs well above Logfire's flat $2 per million. Phoenix also carries a recurring history of breaking upgrades that self-hosters absorb version to version, catalogued in Arize's own migration notes.
Best for: teams standardized on OpenInference and Phoenix, or ML teams that think in drift and model-performance metrics. Pricing: free tier; Pro from $50/month; Enterprise custom. Full breakdown: Logfire vs Arize.
#4. LangSmith: best for LangChain and LangGraph stacks
LangSmith offers strong tracing and evaluation with the tightest integration into LangChain and LangGraph, and a prompt hub that lets you version and update prompts without a full deploy. Its optimization is mostly a playground assistant and external cookbook recipes rather than an in-product optimizer, and its recent alignment tooling tunes the evaluator, not the app. It is framework-agnostic on paper but strongest inside that ecosystem, closed source, with self-hosting reserved for Enterprise and a per-seat plus per-trace price that, at scale, pushes teams to sample their traces down to a fraction just to control the bill, which is the opposite of what observability is for. Base retention is only 14 days; keeping traces the full 400 days costs about ten times as much per trace, and that ceiling applies even on your own hardware.
Best for: teams committed to LangChain and LangGraph. Pricing: free Developer tier; Plus from $39/seat/month; Enterprise custom. Full breakdown: Logfire vs LangSmith.
#5. Langfuse: best open-source option
Langfuse is a broad open-source platform, MIT-licensed, with tracing, prompt management, datasets, and evaluation, and it is the default choice for teams that have no budget for an AI observability tool and want to self-host and self-manage (ClickHouse acquired Langfuse in January 2026; it stays open source and self-hostable). It manages prompts and runs experiments, with textbook versioning and label-based rollout, but it observes and scores rather than proposing a trace-backed fix, so the optimizer half of the loop is one you assemble yourself. Self-hosting has its own recursive cost: you now run Postgres, ClickHouse, Redis, and object storage, infrastructure that itself needs watching, so you end up needing an observability solution to monitor your observability solution. And important governance features (project-level RBAC, SCIM, audit logs) require a commercial license key even when self-hosted.
Best for: devs who want to run infrastructure on their homelabs and teams that need a free and self-hostable platform and will build their own optimizer on top. Pricing: free self-host; cloud from $29/month. Full breakdown: Logfire vs Langfuse.
#6. The eval specialists: DeepEval, Promptfoo, Patronus
DeepEval, Promptfoo, and Patronus are excellent at the pre-ship half of the problem. DeepEval brings pytest-style LLM-as-judge testing and a real prompt optimizer, though one that tunes against your goldens rather than production traces. Promptfoo brings config-driven local evals and red-teaming. Patronus brings purpose-built judge models for hallucination detection and agent stress-testing. They are testing and evaluation frameworks, not production optimization platforms: they tell you what is wrong before or after a run, but they do not manage a live agent's configuration or ship a fix into production. Note the 2026 shakeout: Promptfoo is now owned by OpenAI, and OpenAI's own hosted Evals product is being shut down on November 30, though the open-source openai/evals repository is unaffected.
Best for: building a rigorous eval and red-teaming step into CI, upstream of whatever platform runs your loop.
#7. Elastic: a search company with an LLM dashboard
Elastic is a search and logging company that has bolted LLM observability onto its APM product. It will show you an agent's prompts, responses, token counts, and latency, and it markets AI agent optimization. Look at what backs that claim and it thins out fast. The optimization story is a short Elastic engineering blog series: engineers manually rewriting prompts to cut cost on internal workloads, not a product feature. The early-2026 Agent Builder builds retrieval agents over Elasticsearch data; it does not improve them. And the LLM instrumentation is a recent integration layer on a search engine, not tracing designed for the shape of an agent run. Elastic is also heavy to operate in the first place: between the query DSL, index lifecycle management, and cluster tuning, running it well can feel like it takes a PhD in Elastic, a reputation Elastic half-conceded by shipping AutoOps to catch misconfigurations.
So Elastic shows you prompts. It does not manage them. There is no versioning, no targeting, no rollout, no rollback, no optimizer that proposes a fix, and nothing that turns a bad run into a shipped change. It is an observability view of an agent sold next to the words "agent optimization," and it is last here because a blog post is not a product.
Best for: teams already deep in Elastic who want LLM calls on the same dashboards as the rest of their logs, and who will do any actual optimizing somewhere else.
#Comparison at a glance
Platform
Closes the loop (propose + ship)?
Trace scope
Open standards
Eval / score pricing
Best for
Pydantic Logfire
Yes: diagnoses fault, optimizer proposes, managed config ships
Whole distributed trace
OTel-native, independent
No score meter, flat spans
Optimizing agents in full production context
Braintrust
Yes, for prompts (Loop + registry)
LLM spans (OTel ingest)
Proprietary datastore (Brainstore)
Per-score ($1.50-2.50/1k) + per-GB
Eval-driven CI/CD teams
Arize AX
Yes: meta-prompt optimizer + Prompt Hub
LLM / agent spans
Phoenix ELv2; AX proprietary
Dual-axis: spans + $/GB payload
Enterprise ML + LLM teams
LangSmith
Partial: prompt hub + assistant
Agent / LLM traces
Closed
Seats + traces
LangChain / LangGraph shops
Langfuse
No built-in optimizer
LLM traces (OTel-compatible)
Open source (MIT)
Units include scores
Self-hosted, OSS-first teams
Eval specialists
No: test and score pre-ship
Test-time, not production
Mostly OSS
Varies
Rigorous CI eval and red-teaming
Elastic
No: shows prompts, manages nothing
LLM dashboards bolted on search
OTel; LLM layer bolted on
n/a
Existing Elastic customers
#How to choose
Start with your bottleneck. If you cannot see why the agent failed, or cannot tell whether the agent or your infrastructure is at fault, you need whole-trace observability first, and any full-stack option beats an LLM-only one. If you can see the failures but cannot turn them into shipped fixes without a deploy cycle, you need a platform that closes the loop, and that is where Logfire, Braintrust's Loop, and Arize AX actually compete. If you need to own your infrastructure, start open-source with Langfuse or self-hosted Phoenix. If your problem is pre-ship regression and red-teaming, add an eval specialist to CI regardless of which platform runs your production loop.
For most teams improving a real agent in production on an open, polyglot stack, Pydantic Logfire is the strongest starting point: it is the one that tells you what to fix, proposes the fix, and ships it, over the whole trace, without a per-score meter, and without asking you to bet on a single vendor's proprietary format in a year when everyone else is getting acquired.
#Frequently asked questions
What is AI agent optimization?
It is the loop that turns a failing production run into a shipped improvement: capture the trace, evaluate it, work out whether the agent or the surrounding system is at fault, propose a specific change to the agent's prompt or configuration, ship that change with a safe rollout, and measure the result. Optimization is distinct from observability, which stops at showing you what happened.
What is the best Braintrust alternative for agent optimization?
Pydantic Logfire. Braintrust's Loop is strong, and it does ship prompt changes, but it optimizes the LLM output inside Braintrust. Logfire works over the whole distributed trace, so it can tell you when the real problem was the retrieval or the database rather than the prompt, ships the agent's whole configuration through managed rollout, and does it on OpenTelemetry with no per-score meter.
Does Elastic do AI agent optimization?
Not really. Elastic shows you an agent's prompts and traces and markets AI agent optimization, but the optimization is a handful of engineering blog posts, not a product feature: engineers manually rewriting prompts to cut cost. Elastic displays prompts; it does not version, target, roll out, or roll back anything, and it has no optimizer that proposes a fix. It is an observability view of an agent, not a platform that improves one.
Are there open-source agent-optimization options?
Langfuse is MIT-licensed and self-hostable, and Arize Phoenix is source-available under the Elastic License. Both give you observability and evaluation to build a loop on, rather than a managed optimizer. Because Pydantic Logfire is OpenTelemetry-native, your instrumentation stays portable regardless of where you run it.
What happened to the AI eval tooling in 2026?
It consolidated fast. ClickHouse acquired Langfuse in January, OpenAI acquired Promptfoo in March and announced its own hosted Evals product will shut down on November 30, and Cisco moved to acquire Galileo. The open-source openai/evals framework on GitHub is unaffected, but the pattern is the point: build your loop on a platform that is OpenTelemetry-native and independent, so the loop and your data stay yours no matter who gets bought next.
#Try Pydantic Logfire free
You can have whole-stack agent traces, online and offline evaluation, and the optimization loop running in a few minutes. The free tier includes ten million spans a month.
Start free with Pydantic Logfire
AI is still just engineering.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み