2026 年主要 LLM 観測・評価プラットフォーム比較
本文の状態
日本語全文を表示中
詳細モードで約22分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
2026 年の LLM 観測・評価プラットフォーム市場が成熟し、LLM アプリケーションの複雑な挙動を捉えるための基盤インフラとして位置づけられ、主要ベンダーや導入実態に関する詳細な比較分析が行われる。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月10日 03:21
AI深層分析
キーポイント
市場規模と成長予測
Business Research Company は LLM 観測プラットフォーム市場が 2026 年に 26.9 億ドルに達し、2030 年には 92.6 億ドルに拡大すると予測している。
業界の構造的変化
LLM 観測と評価はオプションツールから、AI を本番環境で運用するチームにとってのコアインフラへと移行した。
主要プラットフォームのカテゴリ分類
市場は AI ネイティブ観測、オープンソース評価ライブラリ、AI ゲートウェイ、APM 拡張の 4 つのカテゴリに明確に分裂している。
エージェント開発の実態と課題
LangChain の調査によると、生産環境でエージェントを実行するエンジニアは過半数を占めるが、評価の実施率は依然として低い。
OpenTelemetry GenAI セマンティック規約の標準化
CNCF プロジェクトである OpenTelemetry が定義する gen_ai.* スパン属性が、主要クラウドプロバイダーやコーディングエージェント間で事実上の標準として採用されている。2026 年の購入判断においては、この OTel 互換性を必須要件と見なすべきである。
重要な引用
Standard application performance monitoring (APM) alone does not capture this semantic behavior
In 2026, this category has moved from optional tooling to core infrastructure for any team running AI in production.
Quality was cited by 32% as the top barrier to production deployment.
Buyers in 2026 should treat OTel compatibility as a hard requirement, not a nice-to-have.
編集コメントを表示
編集コメント
2026 年という未来の視点から、LLM 観測市場が成熟期に入った現状を分析した記事である。各プラットフォームの比較軸が明確に定義されており、本番環境での AI 運用責任者にとって重要な判断材料となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
LLM アプリケーションは、従来のソフトウェアとは異なる方法で失敗します。同じプロンプトでも出力が異なることがありますし、検索ステップで誤ったドキュメントを返しても HTTP ステータスはすべて 200 と表示されることがあります。また、エージェントが 14 回のツール呼び出しをループして数千トークンを消費し、自信満々に間違った回答を返すケースも珍しくありません。
従来のアプリケーションパフォーマンスモニタリング(APM)だけでは、こうした意味的な振る舞い——プロンプトと出力の品質、検索の関連性、あるいはエージェントレベルでの推論トレース——は捉えきれません。
LLM 観測・評価プラットフォームが埋めるのがこのギャップです。これらは LLM パイプラインのすべてのスパン(プロンプト、完成文、検索結果、ツール呼び出し、トークン数、レイテンシ、コスト)を記録し、自動評価器を用いて出力の品質をスコアリングします。2026 年現在、このカテゴリは「あれば便利なツール」から、本番環境で AI を運用するあらゆるチームにとって不可欠な基盤へと進化しています。
市場データは、このシフトを如実に示しています。Business Research Company によると、LLM 観測プラットフォーム市場規模は 2025 年の 19.7 億ドルから 2026 年には 26.9 億ドルに拡大し、2030 年までに 92.6 億ドルに達すると予測されています。これは年平均成長率(CAGR)36.2% のペースです。
Gartner は、2028 年までに LLM 観測への投資が GenAI デプロイ全体の 50% を占めると予測しています。これは 2026 年初頭の 15% から大幅な増加となります。一方、LangChain が実施した「エージェントエンジニアリングの現状」に関する調査(対象者 1,300 名以上)では、回答者の 57% がすでに本番環境でエージェントを稼働させていることが明らかになりました。さらに、ほぼ 89% の企業がエージェント向けの観測機能を導入済みです。
しかし、評価機能の普及は遅れています。オフラインでの評価を実行しているのは 52.4%、オンライン評価は 37.3% に留まり、なんと 29.5% の企業は何らかの評価も実施していません。品質が本番環境への展開における最大の障壁であるという回答が 32% を占めています。
本記事では、主要プラットフォームを「トレーシングの深さ」「評価機能」「本番モニタリング」の 3 つの軸で比較します。数値は 2026 年 8 月時点の一次情報(各社のドキュメント、プレスリリース、公式発表)に基づいて検証済みです。二次情報のみの場合は、その旨を明記してリンクを貼っています。なお、ランキングや「〜に最適」といった判断は、測定されたベンチマーク結果ではなく、編集部による評価です。
2026 年の市場構造
市場は現在、4 つの主要なカテゴリーに分化しています。個々の機能リストよりも、この市場構造の違いを理解することが重要です。
AI ネイティブな観測プラットフォーム:Langfuse、LangSmith、Braintrust、Arize、Opik は、LLM のトレースを主要な対象として扱います。これらはエージェント、検索機能(retrievers)、ツールにわたるネストされたスパンを捉え、本番環境のトラフィックに対して評価スコアを紐付けます。
オープンソースおよびソース利用可能な評価ライブラリとプラットフォーム:Arize Phoenix、DeepEval (Confident AI)、MLflow、RAGAS は、出力に対するスコアリングに焦点を当てています。具体的には、忠実度(faithfulness)、ハルシネーション、回答の関連性、タスク完了率などを評価しますが、これらは多くの場合「LLM を裁判官として使う」手法によって行われます。
AI ゲートウェイ:Helicone、Portkey、LiteLLM は、アプリケーションとモデルプロバイダーの間にプロキシとして機能します。最小限の変更でログ記録、キャッシュ、コスト追跡、ルーティングを追加できます。
APM 拡張機能:Datadog LLM Observability、New Relic、Dynatrace は、既存のインフラ監視に LLM のトレース機能を追加します。これにより、AI の信号と CPU、メモリ、ネットワークなどのメトリクスを相関させて分析することが可能になります。
現在、4 つの主要なプラットフォームを貫く共通規格が一つあります。それは OpenTelemetry GenAI セマンティック・コンベンションです。この規格は、モデル呼び出し、トークン使用量、エージェントのステップ、ツール実行などにおける、ベンダーに依存しない gen_ai.* スパン属性を定義しています。
OpenTelemetry は CNCF プロジェクトとしてこれらのコンベンションを維持しており、Google Cloud、AWS、Azure、Datadog といった主要プラットフォームが採用しています。この規格は専用のリポジトリで管理されており、2026 年 8 月現在、GenAI レジストリの開発も活発に進められています。
コーディングエージェント側でも、この標準への収束が進んでいます。GitHub Copilot のエージェントテレメトリーでは gen_ai.* スプントリーが公開され、Claude Code ではオプトイン形式で OpenTelemetry トレースを提供しています。また Codex もネイティブの OpenTelemetry エクスポートサポートを含んでいます。
一度 gen_ai.* に基づいて計測器(インストルメンテーション)を実装すれば、バックエンド間の移植性が向上し、ベンダー固有の実装コストを削減できます。ただし、実装の詳細にはまだばらつきがある点には注意が必要です。2026 年の購入判断においては、OpenTelemetry 互換性は「あれば便利」な機能ではなく、必須条件として扱うべきです。
3 つの軸:トレース、評価(Evals)、本番環境モニタリング
ベンダー各社がこれらの用語を曖昧に使用しているため、プラットフォームを比較する前に正確な定義を確認することが重要です。
トレーシングとは、LLM アプリケーションが実行したすべての操作の記録です。トレースにはネストされたスパンが含まれており、ユーザーの入力、各検索呼び出し、正確なプロンプトとパラメータを伴う各モデル呼び出し、各ツールの実行、そして最終的な出力などが含まれます。
深さが重要視されるのは、エージェントのトレースが非常に深くネストされており、大量のペイロードを含むためです。1 つの会話でも、数十回のランやツール呼び出しを通じてメガバイト規模のデータが発生することがあります。非決定性(Non-determinism)のため、トレーシングは必須となります。同じプロンプトでも異なる出力が生成されるため、問題の再現には、呼び出し時の正確な入力、モデルパラメータ、温度設定などをキャプチャしておく必要があります。
評価(Evals)は、トレーシングでは答えられない「出力は本当に良かったのか?」という問いに答えます。オフライン評価では、デプロイ前にキュレーションされたデータセットに対してスコアリングを行い、プロンプトやモデル、検索インデックスの変更による回帰(regressions)を検出します。一方、オンライン評価では、LLM をジャッジとして用いてライブの生産トラフィックを評価します。トレースをサンプリングし、忠実性、関連性、毒性、タスク完了度などの観点から採点を行います。
最も困難な失敗は、技術的には正しいがドメインにおいては誤りである出力です。例えば、事実と異なるポリシーの提示、トーンの drifting( drift )、あるいは自信満々だが間違った回答を導き出す検索ミスなどが挙げられます。従来のレイテンシ、エラーレート、可用性といった指標では、こうした意味的な品質の欠陥を検出することはできません。
プロダクションモニタリングは、ダッシュボード、モデルおよびユーザーごとのコスト配分、レイテンシのパーセンタイル、プロンプトやユースケース全体でのドリフト検知、品質スコアの低下時のアラート通知などを通じて、フィードバックループを完結させます。優れたプラットフォームは、実際の運用で発生したトレースを評価用データセットに自動反映し、すべての実世界での失敗が将来の回帰テストとして活用される仕組みを提供します。
各プラットフォームは得意分野と不得意分野が明確に分かれています。ゲートウェイ型ツールはモニタリングに優れていますが、詳細なトレーシング機能は不足しています。評価ライブラリは出力スコアリングには強みがありますが、プロダクション環境の監視までは行えません。以下に紹介するプラットフォームは、これら三つの機能をいかに網羅的にカバーしているかを基準にランキングされています。
- Langfuse (ClickHouse)
Langfuse は「最も広く採用されている LLM エンジニアリングプラットフォーム」を自称しており、そのオープンソースでの導入実績がその主張の正当性を裏付けています。
トレーシング機能:Langfuse は OpenTelemetry、LangChain、OpenAI SDK、LiteLLM などの統合を通じて、LLM の呼び出し、検索、埋め込み、エージェントのアクションにおけるネストされたトレースを捕捉します。特徴的なネストトレースビューでは、多段階の RAG やエージェント実行を、各スパンごとのレイテンシとトークン数を示すステップ可能なツリー構造に集約表示します。2026 年 3 月に導入された観測中心のデータモデルにより、ダッシュボードのパフォーマンスが 10 倍以上向上し、同社によると最大で 165 倍高速化する Langfuse v4 の基盤も整いました。
評価機能:このプラットフォームは、LLM をジャッジとして用いた自動評価、人間による注釈付けキュー、カスタムスコアリング、GitHub Actions を介した CI 実行型のデータセットベースの回帰テストをサポートしています。評価テンプレートには、幻覚(ハルシネーション)、毒性、関連性のチェック項目が含まれています。
プロダクションモニタリングでは、モデル別・ユーザー別・セッション別のコスト内訳を確認でき、会話型エージェントのセッション再生も可能です。
デプロイは MIT ライセンスのコアを提供し、Docker Compose で数分でセルフホスト可能。また、無料枠を備えた Langfuse Cloud での管理利用も選べます。このカテゴリでは、Langfuse はセルフホスティングのリーダーとして広く認知されています。
最も適しているのは、機能豊富なオープンソースプラットフォームを求め、かつデータ所在に厳格な制御を必要とするチームです。
- LangSmith (LangChain)
LangSmith は、エージェントの観測・評価・デプロイを担う LangChain の商用プラットフォームです。Python、TypeScript、Go、Java 向けの SDK と OpenTelemetry サポートを提供し、フレームワークに依存しません。ただし、LangChain 1.0 や LangGraph 1.0 ではデフォルトバックエンドとして採用されており、これらの環境ではほぼゼロの統合コードで連携できます。
トレーシング:会話全体やエージェント実行のトレースを詳細に表示し、各ステップ・ツール呼び出し・中間状態まで把握可能です。内蔵 AI アシスタント「Polly」が大規模なトレースを要約し、問題箇所を特定します。LangSmith Engine はプロダクションでの失敗事例を優先度順にクラスタリングし、トレースやコード上で根本原因を特定、修正案をレビュー用に提案します。
評価:LLM によるジャッジ、コードベースの評価、多ターン評価が、データセットまたはライブのプロダクショントレース上で実行可能です。ジャッジは人間の嗜好に合わせて調整でき、デプロイ前に回帰を防ぐためのサイドバイサイド比較も可能。アノテーションキューにより、ドメインの専門家がエージェントの出力をレビューできます。
運用監視:オンライン評価で実稼働トラフィックのスコアを計測し、自動的なトレースクラスタリングにより利用パターンや障害モードを検出します。2026 年現在、LangSmith は LLM の呼び出しに加え、検索やツール、外部 API にかかるカスタムコストを含め、エージェントワークフロー全体にわたる統一されたコストビューを提供しています。
デプロイメント:データ所在地の要件を持つチーム向けに、AWS または GCP 上のマネージドクラウド、ハイブリッド構成、オンプレミス(セルフホスト)構成に対応します。LangSmith のデプロイメント機能では、人間の承認を介する「人間-in-the-loop」機能を備えた永続的なエージェントランタイムを追加し、個々の実行試行に対して厳密に一度だけ実行されるセマンティクス(exactly-once semantics)を強制します。
最も適しているケース:LangChain や LangGraph をベースに開発を進めるチームや、観測性、評価、マネージド型エージェントデプロイメントを単一ベンダーで実現したい企業向けです。
- Braintrust
Braintrust は本リストの中で「評価(eval)ファースト」のプラットフォームであり、2026 年に AI 評価・観測性分野で最大規模の資金調達ラウンドの一つを背景に成長しています。
トレース:Python、TypeScript、その他の言語に対応するフレームワーク非依存 SDK を通じて、エージェントの完全なトレースを取得できます。専用データベースである Brainstore は、数百万件に及ぶ複雑なトレースに対するクエリを効率的に処理します。
評価:これが Braintrust の中核機能です。バージョン管理されたデータセット、自動化および人間によるスコアリング、モデルやプロンプトの実験、CI における回帰テストにより、デプロイ前に評価結果で回帰を防ぐことが可能です。リリース前にはプレイグラウンドで本番データを基にプロンプトの変更を検証できます。AI エージェントである Loop がトレースを分析し、より良いプロンプトの提案やスコアラーの生成、データセットの自動構築を行います。
プロダクションモニタリング:プロンプト、レスポンス、ツール呼び出し、レイテンシ、コスト、品質にわたるリアルタイムの観測性を実現し、ハルシネーション(幻覚)、ドリフト、回帰の監視も可能です。
最も適しているのは:評価をワークフローの中核に据え、CI/CD の品質ゲートとプロダクションフィードバックループを一つのシステムで完結させたい製品志向の AI チームです。
- Arize AX と Arize Phoenix
Arize AI は 2 つの層からなる戦略を採用しています。エンタープライズ向けに「Arize AX」を提供し、その下層にはオープンソース化されたセルフホスト可能なレイヤーとして「Phoenix」を配置しています。
トレーシング:Phoenix は OpenTelemetry ネイティブであり、Elastic License 2.0 のもとでセルフホスト可能です。OSI が認定するオープンソースライセンスではありませんが、LlamaIndex や OpenAI Agents SDK との統合は強力です。シリーズ C の発表時には月間ダウンロード数が 200 万を超え、最も広く採用されている評価ライブラリの一つとなりました。
評価:Arize の ML 観測性の歴史がこの機能に反映されています。その評価プリミティブは競合他社よりも深く、事前構築されたテンプレートや RAG(検索拡張生成)に特化した品質プロット、時間とともに静かに劣化する出力を検出するドリフト検知機能を備えています。また、音声アプリケーション向けのオーディオ評価機能も導入し、OpenEvals や AgentEvals の取り組みを通じてオープンリサーチを支援しています。
プロダクションモニタリング:埋め込みのクラスタリングやドリフト検知に加え、従来の ML モデルと生成ワークロードの両方をカバーする監視機能を備え、Azure AI Foundry との深い統合を実現しています。
規制が厳しく、精度が極めて重要なワークロードや、従来の機械学習モデルと大規模言語モデル(LLM)を並行して運用する組織に最適です。
- MLflow
Databricks が支援し、Linux Foundation のオープンソースプロジェクトとして展開される「MLflow」は、エージェントの観測性を提供する包括的なプラットフォームへと進化しました。
トレーシング:ユーザーが完全にデータを管理できるネイティブなトレーシング機能を提供します。データは OTel GenAI 形式のセマンティック・ビジュアルアクショントークナイザー(OTel GenAI semantic convention)でエクスポートされるため、独自スキーマに縛られることはありません。
評価(Evals):組み込みの LLM ジャッジや多ターン評価、人間フィードバックとの整合性チェックに対応しています。また、RAGAS、DeepEval、Phoenix、TruLens、Guardrails AI との連携も可能です。さらに、GEPA や MIPRO アルゴリズムを活用したプロンプト最適化機能も搭載されており、評価結果に基づいてプロンプトを自動的に改善できます。
本番環境モニタリング:AI ゲートウェイにより、OpenAI、Anthropic、Bedrock、Azure、Gemini といった主要プラットフォームへのアクセスを一元的に管理。ルーティング、レート制限、フォールバック処理、利用状況の追跡を実現します。
このツールは、トレーシングデータの所有権を重視するチームや、エンタープライズ版での有料制限を避けたい組織、あるいはすでに実験追跡に MLflow を導入しているチームに適しています。既存の MLflow 環境がない場合、より軽量な Langfuse などのツールの方が導入がスムーズなケースもあります。
- Weights & Bases Weave
W&B Weave は、Weights & Bases の実験管理プラットフォームを拡張し、LLM のトレーシングと評価機能を強化しました。マルチエージェントシステムの構造化された実行トレースを記録し、エージェント間の呼び出しにおける親子関係を保ちながら、各エージェントの入力・出力、レイテンシ、トークン使用量を詳細にキャプチャします。
差別化要因は「系譜(lineage)」です。エージェントの動作を、W&B で既に管理されているモデル、データセット、実験履歴と直接比較できます。料金は取り込み量に基づいて設定されており、無料プランでは月間 1 GB の Weave データが含まれます。Pro プランは月額 60 ドルからで 1.5 GB が利用可能、追加の取り込みは MB あたり 0.10 ドルです。そのため、プロンプトや取得したドキュメントが大きいとコストに大きく影響します。LLM の観測機能は、コアとなる実験追跡製品に比べるとまだ新しく、成熟度には差があります。
向いているのは:W&B にすでに投資している ML 研究チームで、プラットフォームを離れることなくプロダクション向けの LLM トレースを実現したいケースです。
- Helicone
Helicone はゲートウェイ分野のリーダー的存在です。オープンソースの AI ゲートウェイであり、1 行の proxy インテグレーションで導入可能です。トラフィックを Helicone にルーティングするだけで、各サービスのインストゥルメンテーション(計測)を行わずにコスト、トークン数、レイテンシのダッシュボードが利用できます。組み込みのレスポンスキャッシング機能により、シンプルなヘッダー設定だけで API コストとレイテンシを削減でき、非技術系のチームメンバーもアクセスできるプロンプト実験機能をサポートしています。
強みであると同時に限界でもあります。ここでの観測はリクエスト中心であり、深いエージェントグラフやスパンレベルの推論ステップ、豊富なプロダクション評価ループは主要な機能ではありません。多くのチームでは、Helicone のゲートウェイ機能と、専用のトレースまたは評価プラットフォームを併用しています。
向いているのは:ほぼゼロの設定工数で、複数プロバイダーのコスト可視化、キャッシング、ルーティングを即座に実現したいチームです。
- Datadog LLM Observability
Datadog の LLM 観測機能は、APM(アプリケーションパフォーマンス管理)の拡張カテゴリに位置します。トークン使用量やリクエストあたりのコスト、モデルのレイテンシ、プロンプトインジェクション試行といったセキュリティ信号を、既存のインフラメトリクス、APM、ログデータと統合して取り込みます。これにより、1,000 以上の組み込み連携を通じて、AI の挙動とシステム全体の健全性を相関分析できます。また、OTel GenAI セマンティック・コンベンション v1.37 以降にもネイティブ対応しています。
その後、Datadog は評価機能やエージェント監視、AI セキュリティ信号のサポートも強化しました。本質的な違いは「あるかないか」ではなく、「どこに重点を置くか」です。最大の強みは、AI のトレースをより広範な APM、インフラ、セキュリティスタックと結びつけて分析できる点にあり、一方、AI ネイティブプラットフォームの開発・評価ワークフローの中心にあります。すでに Datadog を標準化している企業にとって、LLM モジュールは導入障壁が最も低い選択肢です。CI(継続的インテグレーション)によるガバナンス付きの評価を必要とするチームは、専用の評価プラットフォームをその上に重ねて利用するのが一般的です。
向いているのは:すでに運用しているインフラやインシデント管理ワークフローと LLM トレースを相関させたい企業向けです。
一瞥での比較
- Platform:Camp / License / Model / Tracing Depth / Eval Strength / Self-Host
- Langfuse:AI-native OSS / MIT core; cloud / Deep, OTel-native / Strong (judge + datasets + CI) / Yes (leader)
- LangSmith:AI-native commercial / Proprietary / Deepest for LangChain/LangGraph / Strong (calibrated judges, clustering) / Yes (enterprise)
- Braintrust:Eval-first commercial / Proprietary / Deep (Brainstore) / Strongest workflow (CI gates, Loop) / Hybrid options
Arize AX / Phoenix
AI ネイティブかつソースコード公開型。Phoenix は ELv2 ライセンスでソースコード公開。OTel ネイティブ。深層データ、ドリフト検知、音声対応も可能(Phoenix 版)。
MLflow
OSS プラットフォーム。Apache 2.0 ライセンス。深層データと OTel GenAI エクスポートに対応。評価機能は強力(ジャッジや GEPA/MIPRO を活用)。
W&B Weave
ML-platform の拡張機能。Apache 2.0 SDK と商用クラウドを提供。マルチエージェントツリーでの評価に優れる。スコアラーやジャッジによる評価も可能。エンタープライズ向け。
Helicone
ゲートウェイ型。オープンソース。リクエストレベルの可観測性を提供。
Datadog LLM Obs.
APM 拡張機能。プロプライエタリ。インフラと相関した評価が可能。SaaS のみで、オンプレミス非対応。
※深さや強度の評価は、ベンダー資料や独立したレビューに基づく編集者の主観的な判断であり、測定されたベンチマーク結果ではありません。
主要なポイント
LLM 可観測性市場は 2026 年に 26.9 億ドルと推定され、2030 年には 92.6 億ドルに達し、年平均成長率(CAGR)は 36.2% と予測されています。
調査対象組織の 89% がエージェント可観測性を利用しており、52.4% がオフライン評価を、37.3% がオンライン評価を実行しています。
AI ネイティブ分野では Langfuse(現在は ClickHouse の一部)、LangSmith、Braintrust、Arize が先頭を走っています。ゲートウェイ分野は Helicone が、APM 拡張機能分野は Datadog がリードしています。
OpenTelemetry GenAI セマンティック・コンベンションが相互運用性の標準となっています。OTel サポートの対応状況は、購入判断における必須条件として扱うべきです。
自社の技術スタックとチーム構成に合わせて選択しましょう:LangChain や LangGraph を使うなら LangSmith、セルフホスティング重視なら Langfuse、評価の厳密さを求めるなら Arize、評価ファーストのワークフローを組むなら Braintrust がおすすめです。
原文を表示
LLM applications fail in ways traditional software does not. The same prompt can produce different outputs. A retrieval step can return the wrong document while every HTTP status reads 200. An agent can loop through fourteen tool calls, burn thousands of tokens, and deliver a confidently wrong answer. Standard application performance monitoring (APM) alone does not capture this semantic behavior — prompt and output quality, retrieval relevance, or agent-level reasoning traces.
This is the gap LLM observability and evaluation platforms fill. They record every span of an LLM pipeline — prompts, completions, retrievals, tool calls, token counts, latencies, and costs — and then score outputs for quality using automated evaluators. In 2026, this category has moved from optional tooling to core infrastructure for any team running AI in production.
The market data reflects the shift. The Business Research Company sizes the LLM observability platform market at $2.69 billion in 2026, up from $1.97 billion in 2025, and projects $9.26 billion by 2030 at a 36.2% forecast CAGR. Gartner predicts that by 2028, LLM observability investments will account for 50% of GenAI deployments, up from 15% in early 2026. LangChain’s State of Agent Engineering survey of 1,300+ professionals found that 57% of respondents now run agents in production. Nearly 89% have implemented observability for their agents. Evaluation lags behind: 52.4% run offline evaluations, 37.3% run online evaluations, and 29.5% report no evaluation at all. Quality was cited by 32% as the top barrier to production deployment.
This article compares the leading platforms across three axes: tracing depth, evaluation capability, and production monitoring. Figures were checked against primary sources (company documentation, press pages, and announcements) as of August 2026; where only secondary reporting exists, it is linked and identified as such. Rankings and “best for” judgments are editorial assessments, not measured benchmarks.
How the Category is Structured in 2026
The market has split into four camps, and understanding the split matters more than any individual feature list.
AI-native observability platforms: Langfuse, LangSmith, Braintrust, Arize, Opik — treat the LLM trace as the primary object. They capture nested spans across agents, retrievers, and tools, and attach evaluation scores to production traffic.
Open-source and source-available evaluation libraries and platforms: Arize Phoenix, DeepEval (Confident AI), MLflow, RAGAS — focus on scoring outputs: faithfulness, hallucination, answer relevance, and task completion, often via LLM-as-a-judge.
AI gateways: Helicone, Portkey, LiteLLM — sit as a proxy between the application and model providers. They add logging, caching, cost tracking, and routing with minimal code changes.
APM extensions: Datadog LLM Observability, New Relic, Dynatrace — bolt LLM tracing onto existing infrastructure monitoring so AI signals correlate with CPU, memory, and network metrics.
One standard now connects all four camps. The OpenTelemetry GenAI semantic conventions define vendor-neutral gen_ai.* span attributes for model calls, token usage, agent steps, and tool executions. OpenTelemetry, a CNCF project, maintains these conventions, which are adopted by platforms including Google Cloud, AWS, Azure, and Datadog. The conventions now live in a dedicated repository, with the GenAI registry under active development as of August 2026. Coding agents are converging on the standard too: GitHub Copilot’s agent telemetry exposes gen_ai.* span trees, Claude Code provides opt-in OpenTelemetry tracing, and Codex includes native OpenTelemetry export support. Instrumenting once against gen_ai.* improves backend portability and reduces vendor-specific instrumentation, even if implementations still differ. Buyers in 2026 should treat OTel compatibility as a hard requirement, not a nice-to-have.
The Three Axes: Tracing, Evals, and Production Monitoring
Because vendors use these terms loosely, precise definitions help before comparing platforms:
Tracing is the record of everything an LLM application did. A trace contains nested spans: the user input, each retrieval call, each model invocation with its exact prompt and parameters, each tool execution, and the final output. Depth matters because agent traces are deeply nested with heavy payloads — a single conversation can generate megabytes of data across dozens of runs and tool calls. Non-determinism makes tracing non-negotiable: the same prompt produces different outputs, so an issue cannot be reproduced without capturing the exact input, model parameters, and temperature at call time.
Evals answer the question tracing cannot: was the output any good? Offline evals score curated datasets before deployment, catching regressions when a prompt, model, or retrieval index changes. Online evals score live production traffic, typically via LLM-as-a-judge, sampling traces and grading them for faithfulness, relevance, toxicity, or task completion. The hardest failures are outputs that are technically valid but wrong for the domain — a hallucinated policy, a drifting tone, a retrieval miss that produces a confident but incorrect answer. Traditional latency, error-rate, and availability metrics do not detect these semantic quality failures.
Production monitoring closes the loop: dashboards, cost attribution per model and user, latency percentiles, drift detection across prompts and use cases, and alerting when quality scores fall. The best platforms feed production traces back into eval datasets, so every real-world failure becomes a future regression test.
A platform can be strong on one axis and weak on another. Gateways excel at monitoring but skip deep tracing. Eval libraries score outputs but do not watch production. The platforms below are ranked on how completely they cover all three.
- Langfuse (ClickHouse)
Langfuse describes itself as the most widely adopted LLM engineering platform, and its open-source adoption numbers back a strong claim.
Tracing: Langfuse captures nested traces for LLM calls, retrieval, embedding, and agent actions through OpenTelemetry, LangChain, OpenAI SDK, and LiteLLM integrations. Its signature nested trace view collapses a multi-step RAG or agent run into a stepable tree with per-span latencies and token counts. An observations-centric data model shipped in March 2026, delivering 10x+ dashboard performance gains and laying the groundwork for Langfuse v4, which the company says runs up to 165x faster.
Evals: The platform supports LLM-as-a-judge evaluators, human annotation queues, custom scores, and dataset-based regression testing that runs in CI via GitHub Actions. Evaluator templates cover hallucination, toxicity, and relevance.
Production monitoring: Cost breakdowns by model, user, or session, plus session replays for conversational agents.
Deployment: MIT-licensed core, self-hostable via Docker Compose in minutes, or managed on Langfuse Cloud with a free tier. Langfuse is widely regarded as the self-host leader in this category.
Best for: teams that want a full-featured, open-source, framework-agnostic platform with strict data-residency control.
- LangSmith (LangChain)
LangSmith is LangChain’s commercial platform for observing, evaluating, and deploying agents. It is framework-agnostic with Python, TypeScript, Go, and Java SDKs plus OpenTelemetry support, but it is the default backend for LangChain 1.0 and LangGraph 1.0, where integration requires near-zero glue code.
Tracing: Full conversation and agent-run traces expose every step, tool call, and intermediate state. Polly, a built-in AI assistant, summarizes large traces to pinpoint problems. LangSmith Engine clusters production failures into prioritized issues, locates root causes in traces and code, and proposes fixes for review.
Evals: LLM-as-judge, code-based, and multi-turn evaluators run on datasets or live production traces. Judges can be calibrated against human preferences, and side-by-side comparisons gate regressions before deployment. Annotation queues let domain experts review agent outputs.
Production monitoring: Online evals score live traffic, and automatic trace clustering detects usage patterns and failure modes. As of 2026, LangSmith provides a unified cost view across the full agent workflow — LLM calls plus custom costs for retrieval, tools, and external APIs.
Deployment: Managed cloud on AWS or GCP, hybrid, and self-hosted configurations for teams with data-residency requirements. LangSmith Deployment adds a durable agent runtime with human-in-the-loop approvals, and enforces exactly-once semantics for individual run attempts.
Best for: teams building on LangChain or LangGraph, and enterprises that want observability, evals, and managed agent deployment in one vendor.
- Braintrust
Braintrust is the eval-first platform in this list, behind one of 2026’s largest funding rounds in the AI evaluation and observability category.
Tracing: Framework-agnostic SDKs across Python, TypeScript, and other languages capture full agent traces. Brainstore, a purpose-built database, handles queries over millions of complex traces efficiently.
Evals: This is Braintrust’s core. Versioned datasets, automated and human scoring, model and prompt experiments, and CI regression testing let eval results block regressions before deployment. A playground tests prompt changes against real production data prior to release. Loop, an AI agent, analyzes traces to suggest better prompts, generate scorers, and build datasets automatically.
Production monitoring: Real-time observability across prompts, responses, tool calls, latency, cost, and quality, with monitoring for hallucination, drift, and regression.
Best for: product-focused AI teams that want evaluation as the center of the workflow, with CI/CD quality gates and production feedback loops in one system.
- Arize AX and Arize Phoenix
Arize AI runs a two-tier strategy: Arize AX for enterprises and Phoenix as its source-available, self-hostable layer.
Tracing: Phoenix is OpenTelemetry-native and self-hostable under the Elastic License 2.0 — source-available, though not an OSI-approved open-source license, with strong integrations for LlamaIndex and the OpenAI Agents SDK. At the Series C announcement, Phoenix had over two million monthly downloads, making it one of the most widely adopted eval libraries.
Evals: Arize’s ML-observability heritage shows here. Its eval primitives run deeper than most competitors, with pre-built templates, RAG-specific quality plots, and drift detection that catches outputs quietly degrading over time. Arize also introduced audio evaluation capabilities for voice applications and funds open research through its OpenEvals and AgentEvals initiatives.
Production monitoring: Embedding clustering, drift detection, and monitoring that spans both traditional ML models and generative workloads, with deep Azure AI Foundry integrations.
Best for: regulated or accuracy-critical workloads that need the deepest evaluation rigor, and organizations running classic ML and LLMs side by side.
- MLflow
MLflow, the Linux Foundation open-source project backed by Databricks, has evolved into a full agent observability platform.
Tracing: Native tracing for agents with trace data fully owned by the user, and export in OTel GenAI semantic convention format so nothing is locked into a proprietary schema.
Evals: Built-in LLM judges, multi-turn evaluation, judge alignment with human feedback, and integrations with RAGAS, DeepEval, Phoenix, TruLens, and Guardrails AI. MLflow also ships prompt optimization using GEPA and MIPRO algorithms that improve prompts automatically from eval results.
Production monitoring: An AI Gateway centralizes LLM access with routing, rate limiting, fallbacks, and usage tracking across OpenAI, Anthropic, Bedrock, Azure, and Gemini.
Best for: teams that prioritize trace-data ownership, want zero enterprise paywalls, or already run MLflow for experiment tracking. Teams without an existing MLflow footprint may find lighter tools like Langfuse faster to adopt.
- Weights & Biases Weave
W&B Weave extends the Weights & Biases experiment-tracking platform into LLM tracing and evaluation. It records structured execution traces for multi-agent systems, preserving parent-child relationships between agent calls, with inputs, outputs, latency, and token usage captured per agent.
The differentiator is lineage: agent behavior can be compared directly against model, dataset, and experiment history already managed in W&B. Pricing is ingestion-based: the free plan includes 1 GB of Weave data per month, Pro starts at $60/month with 1.5 GB, and additional ingestion runs $0.10 per MB — so large prompts and retrieved documents materially affect cost. The LLM observability layer is newer and less mature than the core experiment tracking product.
Best for: ML research teams already invested in W&B who want production LLM tracing without leaving the platform.
- Helicone
Helicone leads the gateway camp. It is an open-source AI gateway with one-line proxy integration: route traffic through Helicone and dashboards for cost, tokens, and latency appear without instrumenting every service. Built-in response caching cuts API costs and latency via simple headers, and the platform supports prompt experimentation accessible to non-technical team members.
The strength is also the boundary. Observability here is request-centric — deep agent graphs, span-level reasoning steps, and rich production eval loops are not the core story. Many teams pair Helicone’s gateway with a dedicated tracing or eval platform.
Best for: teams that want instant multi-provider cost visibility, caching, and routing with near-zero setup effort.
- Datadog LLM Observability
Datadog LLM Observability represents the APM-extension camp. It ingests token usage, cost per request, model latency, and security signals such as prompt-injection attempts alongside Datadog’s existing infrastructure metrics, APM, and logs, correlating AI behavior with system health across 1,000+ built-in integrations. Datadog also natively supports OTel GenAI Semantic Conventions v1.37+.
Datadog has since added evaluations, agent monitoring, and AI security signals, so the honest differentiation is emphasis rather than absence: its primary advantage is correlating AI traces with the broader APM, infrastructure, and security stack, while AI-native platforms center the development and eval workflow. For organizations already standardized on Datadog, the LLM module is the path of least resistance; teams needing CI-gated evals often layer a dedicated eval platform on top.
Best for: enterprises that want LLM traces correlated with infrastructure and incident-management workflows they already run.
Comparison at a Glance
PlatformCampLicense / ModelTracing DepthEval StrengthSelf-Host
LangfuseAI-native OSSMIT core; cloudDeep, OTel-nativeStrong (judge + datasets + CI)Yes (leader)
LangSmithAI-native commercialProprietaryDeepest for LangChain/LangGraphStrong (calibrated judges, clustering)Yes (enterprise)
BraintrustEval-first commercialProprietaryDeep (Brainstore)Strongest workflow (CI gates, Loop)Hybrid options
Arize AX / PhoenixAI-native + source-availablePhoenix: ELv2, source-availableDeep, OTel-nativeDeepest primitives, drift, audioYes (Phoenix)
MLflowOSS platformApache 2.0Deep, OTel GenAI exportStrong (judges, GEPA/MIPRO)Yes
W&B WeaveML-platform extensionApache 2.0 SDK; commercial cloudGood (multi-agent trees)Good (scorers, judge)Enterprise
HeliconeGatewayOpen sourceRequest-levelLightYes
Datadog LLM Obs.APM extensionProprietaryGood, infra-correlatedModerateNo (SaaS)
Depth and strength ratings are editorial assessments based on vendor documentation and independent reviews, not measured benchmarks.
Key Takeaways
The LLM observability market is estimated at $2.69B in 2026, heading to $9.26B by 2030 at a 36.2% CAGR.
89% of surveyed organizations use agent observability, while 52.4% run offline evals and 37.3% run online evals.
Langfuse (now part of ClickHouse), LangSmith, Braintrust, and Arize lead the AI-native camp; Helicone leads gateways; Datadog leads APM extensions.
OpenTelemetry GenAI semantic conventions are the portability standard — make OTel support a hard buying requirement.
Pick by stack and team shape: LangSmith for LangChain/LangGraph, Langfuse for self-hosting, Arize for eval rigor, Braintrust for eval-first workflows.
The post Top LLM Observability and Evaluation Platforms in 2026: Langfuse, LangSmith, Braintrust, Arize, and More Compared appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み