Artificial Analysis、カスタムベンチマークプラットフォーム「Optima」を発表
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
Artificial Analysis は、ユーザーが独自のデータやワークフローに基づいてベンチマークを構築・実行し、モデルの性能だけでなくコストと時間を比較できる新プラットフォーム「Optima」を発表した。
AI深層分析を開く2026年8月14日 15:56
AI深層分析
キーポイント
柔軟なベンチマーク構築機能
ユーザーは既存の評価データセットやエージェントのトレースデータをアップロードする、コーディング環境からコンテキストを取得する、あるいは使用例を記述するだけで独自のベンチマークを作成できる。
最新モデルとの一括比較
単一のクリックで主要な最新のモデルに対して同じベンチマークを実行でき、新モデルのリリースに合わせてリーダーボードを常に最新の状態に保つことが可能である。
客観的評価とペアワイズ判定
Artificial Analysis が採用している客観的なルブリック基準や、GDPval-AA などのベンチマークで用いられたペアワイズ判定アプローチを用いて、モデルの回答を評価する仕組みを提供する。
コストと効率性の可視化
性能スコアだけでなく、タスクあたりのコストや所要時間を追跡し、カテゴリ別結果やカスタムメトリクスをサポートすることで、特定のユースケースにおけるモデル間のトレードオフを明確にする。
カスタムベンチマークの具体例
コスト削減や品質維持、特定の業界の文体模倣、独自画像データセットの要素識別など、ユースケースに合わせたベンチマーク作成が可能である。
重要な引用
Optima lets you find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task.
Build benchmarks from your own data and workflows.
Cost per Task and Time per Task are tracked alongside benchmark scores
Which model can save me 10x the cost without a meaningful decrease in quality for my finance and accounting agent?
編集コメントを表示
編集コメント
この発表は、汎用的なベンチマークから自社の業務に特化した評価へシフトする動きを象徴しており、実務における AI モデル選定の精度向上に寄与すると考えられる。特にコストと時間の可視化機能は、大規模導入前の検証プロセスにおいて重要な役割を果たすだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Optima の紹介
ベンチマークの構築と実行は容易ではありません。Artificial Analysis がこれまで培ってきた研究とベンチマーク開発の知見を凝縮し、Optima という新しいプラットフォームを発表しました。これにより、独自のワークロード上でモデルを評価し、パフォーマンスや速度、コスト効率を比較することが可能になります。
Optma を利用すれば、特定のタスクに最適なモデルを見つけたり、現在の設定と同等のパフォーマンスを持ちながら、10 分の 1 のコストまたは時間で処理できる代替案を見出したりできます。
自社のデータとワークフローに基づいてベンチマークを構築できます。Optima では、ベンチマーク作成に 3 つの方法があります。
既存の評価データを自分のファイルや Hugging Face からアップロードする、Arize や Braintrust、Langfuse などのプラットフォームからエージェントのトレースをインポートする、あるいは Optima スキルをインストールしてコーディング環境のコンテキストや過去のセッション情報を活用してベンチマークを作成する方法です。また、使用ケースの説明と入力・出力の例を提供するだけで、Optima が自動的にベンチマークを構築することも可能です。
最新モデルへの対応も充実しています。主要なモデルに対してワンクリックで同一のベンチマークを実行可能で、新モデルがリリースされたらすぐにリーダーボードを更新できます。
Artificial Analysis の評価基準を独自のベンチマークにも適用できます。客観的なルブリック基準に基づく評価や、GDPval-AA や AA-Briefcase といった Artificial Analysis ベンチマークで採用されているペアワイズ比較手法による評価が可能です。ペアワイズ比較では、サンプルから好みの回答を選択するだけで、Optima がその選好に基づいてテストセット全体でのモデル順位を算出します。
性能だけでなく、コストと時間効率も比較できます。Optima のベンチマークはモデルの性能のみならず、「タスクあたりのコスト」や「タスクあたりの所要時間」もスコアと同時に追跡します。カテゴリ別結果のカスタムメトリクス対応により、特定のユースケースにおけるモデル間のトレードオフを明確に比較可能です。
事前リリーステスターが構築した事例
ローンチ前に、一部の事前リリーステスターに Optima を提供しました。彼らが作成したベンチマークの例は以下の通りです:
私の財務・会計エージェントで、品質の低下を最小限に抑えつつコストを 10 分の 1 に削減できるモデルはどれか?
私の法律家向けエージェントに、弁護士のような文章スタイルを最もよく反映しているのはどのモデルか?
カスタム画像データセット内の異なる要素を最も正確に識別できるのはどのモデルか?
はじめに
Optima は本日利用可能です。あなた自身でベンチマークを作成し、成果物を Twitter で @ArtificialAnlys さんにタグ付けしてください。
原文を表示
Introducing Optima
Building and running benchmarks is difficult. We have distilled Artificial Analysis' research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency.
Optima lets you find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task.
How Optima works
We have applied the benchmarking approaches and infrastructure we use at Artificial Analysis across each part of Optima.
- Build benchmarks from your own data and workflows. There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize, Braintrust and Langfuse. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you.
- Run across the latest models. Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released.
- Bring Artificial Analysis grading to your own benchmark. Evaluate responses against objective rubric criteria, or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set.
- Compare performance, cost and time efficiency. Optima benchmarks more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, so you can compare the tradeoffs between models for your specific use case.
What pre-release testers built
Ahead of launch, we gave a group of pre-release testers access to Optima. Examples of benchmarks they created include:
- Which model can save me 10x the cost without a meaningful decrease in quality for my finance and accounting agent?
- Which model best matches the writing style of lawyers for my legal agent?
- Which model can best identify different elements in my custom image dataset?
Get started
Optima is available today. Build your own benchmark, and tag @ArtificialAnlys with what you create.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み