OpenRouter、エージェントの行動とコストを追跡する機能追加
OpenRouter は、AI エージェントの実行内容とコストを詳細に記録・追跡できる「Classifiers」機能を新設し、利用状況の可視化を実現した。
AIニュース価値スコアβ
主要ニュースAI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 75
- 情報源の信頼性
- 25
- 新規性
- 75
- 検索具体性
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
OpenRouter が AI エージェントの行動追跡とコスト管理を可能にする新機能を発表したことは、エージェント運用における重要な実装知見であり、新規性が高い。ただし、日本固有の導入事例や規制情報がないため、日本の関連性は限定的となる。
キーポイント
エージェント行動の追跡機能
OpenRouter が新たに発表した Classifiers は、AI エージェントが具体的に何を実行したかを記録する機能を備えている。
コスト管理の可視化
各エージェントの実行に伴う費用を明確に追跡・表示することで、利用コストの透明性を高める。
運用効率の向上
利用状況の詳細な可視化により、開発者や企業が AI エージェントの運用最適化とリソース配分を効果的に行えるようになる。
重要な引用
OpenRouter は、AI エージェントが実行した作業内容とその費用を記録・追跡できる「Classifiers」機能を新たに発表した
これにより、利用状況の可視化が可能となる
影響分析・編集コメントを表示
影響分析
本発表は、複雑化する AI エージェント環境におけるコスト管理と運用監視の課題に対する重要な解決策となる。特に、エージェントが自律的に行動する現代において、その挙動と費用を可視化できることは、企業による大規模導入や信頼性確保に不可欠な基盤を提供する。
編集コメント
AI エージェントが自律的に行動する時代において、その挙動とコストを可視化できる機能は運用管理の必須要件となりつつあります。OpenRouter のこの取り組みは、企業によるエージェント導入の障壁を下げる重要な一歩と言えるでしょう。
OpenRouter の生成結果を、構造化されたメタデータ付きで自動的に分類できるようになりました。これにより、AI の利用状況レポートが容易になります。
すべてのリクエストには、作業の種類や複雑さのレベル、発生源となる部署、機密データを誤って含んでいないかなどの情報が含まれています。ベータ版として提供されている「クラスフィアー」機能を使えば、こうした情報を可視化できます。タスクタイプ、エージェントの複雑度、コンプライアンスカテゴリ、コストセンターなど、独自の基準を設定するだけで、任意のモデルが各生成結果(またはサンプリングされたサブセット)を定義した分類体系に照らしてタグ付けし、その結果をログに記録します。これにより、エージェントやユーザーが何を行っているか、どのタスクにどのモデルを使用しているか、コストがどこに発生しているかを継続的に把握できます。
ワークスペース設定 からクラスフィアーを作成するか、事前に ドキュメント をご覧ください。
テンプレートを選択するか、独自の分類体系を定義する
クラスフィアーは、4 つの要素からなる小さな設定ファイルです。1 つ目は分類体系(タクソノミー)で、最大 8 つの次元を持ち、各次元には任意の値を設定できます。2 つ目は分類プロンプトで、システムメッセージとして分類モデルに送信される指示文です。3 つ目は、各プロンプトを読み込んで適用するモデルです。4 つ目はサンプリングレートです。
クラスフィアーは、リクエスト完了後に非同期で実行されるため、推論パスのレイテンシには一切影響しません。
6 つの事前設定されたテンプレートから選択するか、既存のテンプレートをカスタマイズするか、ゼロから独自に構築することも可能です。
| テンプレート | タグ付け内容 |
|---|
部署
どの業務部門からリクエストが来たか。エンジニアリング、営業、マーケティング、法務などです。組織のどの部分が推論コストを牽引しているかを把握するのに役立ちます。
対象者
出力先は誰か。社内利用、顧客向け、規制当局、一般公開などです。モデルの出力を誰が読むかに依存するコンプライアンスワークフローにフィードバックされます。
タスクタイプ
モデルは何をしているのか。コーディング、エージェントワークフロー、データ処理、コンテンツ作成などです。各タスクに対して適切なティアのモデルが使われているか確認するのに役立ちます。
エンジニアリング作業
機能開発、バグ修正、ドキュメント作成、リファクタリング、コードレビューなどです。AI がどこで役立っているか、またどのモデルがどの種類の作業に使われているかを追跡するのに適しています。
エージェントの複雑さ
難易度ティア(単純なツール呼び出しから最前線の専門業務まで)とタスクファミリーです。エージェントを運用しているチームにとって重要なのは、「どのモデルが難しいタスクをうまく処理したか」という問いに答えることです。
資本化可能なソフトウェア費用
AI 支援によるエンジニアリング作業が、開発(資本化可能)なのか、保守・運用・サポート(費用化)なのかの区別です。
分類モデルを選択してください。最もコストパフォーマンスが良いのは Gemini 3.5 Flash Lite です。安価で構造化された出力に対する精度が高く、ほとんどの分類体系に十分対応しています。必要に応じていつでもモデルを変更できます。
高スループット環境では、すべてのリクエストを分類するコストが累積します。コストを抑えるにはサンプリングレートを活用しましょう。100% のトラフィックに対して高精度なコンプライアンス分類器を実行しつつ、より広範なコスト帰属分析用分類器は同じトラフィックの 10% をサンプリングすることで、必要な監督レベルに見合った比例したコストで運用できます。
分類器の出力は、定義した次元と値に制約された構造化フォーマットに変換されます。各生成結果は ログ にタグ付けされるため、分類に基づいてリクエストをフィルタリング可能です。例えば、「department: legal」や「agent_complexity_difficulty_tier: complex_multistep」というタグが付いたリクエストを一括で抽出できます。各タグ付き生成の詳細パネルでは、分類された次元と値の内訳を確認できます。
また、新しいタクソノミーの妥当性を検証するために、過去の任意の生成結果に対してオンデマンドで分類器を実行することもできます。ログから対象の生成を開き、分類器を選択して、どのようにタグ付けされるかを確認しましょう。

Activity で集約する
個々の生成結果に付与されたタグは「このリクエストは何だったのか」を答えます。一方、Activity Explorer は集計的な問いに応えるツールです。任意の分類器次元でトラフィックをグループ化すれば、各タスクタイプやエージェント複雑度レベルにどのモデルが使用されているか、またどの部署やタスクが最も支出を押し上げているかを把握できます。
結果は時間経過とともに集約され、データ内のパターン変化を追跡できます。これにより、ステークホルダーに対して AI 利用がどのように管理されているかを明確に示すことが可能です。
Classifier(分類器)のフィルタリング機能は Activity タブにも反映されるため、任意の Classifier の値に基づいた トレンド や ガードレール の適用状況を直接確認できます。

使い始め方
Classifier は現在ベータ版として利用可能です。ワークスペースで Classifier を作成するか、ドキュメント を参照して、タクソノミーの設計方法や課金ルール、内部での分類処理の仕組みについて詳しく学んでください。なお、入力・出力ログを無効化している場合でも Classifier は正常に動作します。
Discord の #feedback チャンネルでのご意見をお聞かせください。
原文を表示
You can now automatically classify your OpenRouter generations with structured metadata for AI usage reporting.
Every request carries information: the type of work, the level of complexity, which department it came from, whether it contains internal data it shouldn’t. Classifiers, now available in beta, give you that visibility. Define your criteria (task type, agent complexity, compliance category, cost center). A model of your choice tags each generation, or a sampled subset, against your taxonomy and write the results to your logs. You get continuous visibility into what your agents and users are doing, which models they’re using for different tasks, and where the costs go.
Create a classifier in your workspace settings, or read the docs first.
Pick a template or define your own taxonomy
A classifier is a small config with four parts: a taxonomy (up to eight dimensions, each with the values you choose), a classification prompt (instructions sent to the classifier model as a system message), a model to read each prompt and apply it, and a sampling rate. Classification runs asynchronously after each request completes, so it never adds latency to your inference path.
Choose from six preset templates, customize a template, or build your own from scratch.
TemplateWhat it tags
DepartmentWhich business function originated the request: engineering, sales, marketing, legal, and so on. Useful for seeing which parts of the org drive inference cost
AudienceWho the output is for: internal use, client-facing, regulators, or the public. Feeds compliance workflows that depend on who reads a model’s output
Task typeWhat the model is doing: coding, agent workflows, data processing, content writing. Useful to check whether the right tier of model is being used for each task
Engineering workFeature development, bug fixing, documentation, refactoring, code review. Good for tracking where AI is helping and which models are used for each type of work
Agent complexityDifficulty tier (from trivial tool calls to frontier-expert work) plus task family. For teams running agents, where “which model handled a hard task well” is the question that matters
Capitalizable software expenseWhether AI-assisted engineering work is potentially capitalizable development versus maintenance, operations, or support
Select your classification model. We recommend Gemini 3.5 Flash Lite for the best value: cheap, strong accuracy on structured output, good enough for most taxonomies. You can change the model at any time.
At high throughput, the cost of classifying every request adds up. Use the sampling rate to keep costs down. Run a high-fidelity compliance classifier at 100% while a broader cost-attribution classifier samples 10% of the same traffic, keeping costs proportional to the oversight you need.
Classifier outputs are coerced into structured formats, constrained to the dimensions and values you define. Every classified generation is tagged in your logs, so you can filter requests by classification. For example, you can pull every request tagged department: legal or agent_complexity_difficulty_tier: complex_multistep. Each tagged generation’s detail panel breaks down classified dimensions and values.
You can also run a classifier on demand against any past generation to sanity-check a new taxonomy. Open it in your logs, pick a classifier, and see how it gets tagged.

Roll it up in Activity
Individual tags on generations answer “what was this request?” The Activity Explorer answers the aggregate questions: group your traffic by any classifier dimension to see which models are being used for each task type or level of agent complexity, and which departments or tasks drive the most spend.
Results are aggregated over time; watch patterns shift in your data and show stakeholders how your AI usage is governed. Classifier filters carry across the Activity tabs so you can see trends and guardrail enforcement by any classifier value.

Get started
Classifiers are available now in beta. Create a classifier in your workspace or read the docs to learn more about taxonomy design, billing, and how classification works under the hood. Classifiers work even with input & output logging disabled.
Tell us what you think in #feedback on Discord.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み