マイクロソフト、テキスト記述から AI の動作テストを構築できる新ツールを発表
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TechCrunch AI
マイクロソフトは開発者がテキスト記述を用いて AI の動作テストを迅速に構築・実行できる新しいツールの提供を開始した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI 研究者やラボは、安全性やコンプライアンスから、迎合性やアライメントに至るまで、AI モデルの評価において飛躍的な進歩を遂げてきました。しかし、企業や開発者には新たな特定のニーズが生じています:自社の製品やサービスに対して AI システムが意図した通りに動作していることを確認することです。
そのテストプロセスをよりシンプルにするため、マイクロソフトは火曜日に ASSERT(Adaptive Spec-driven Scoring for Evaluation and Regression Testing の略)を発表しました。
マイクロソフトによると、このオープンソースフレームワークは、AI を活用して目標、ポリシー、または意図された動作に関する高レベルの自然言語記述を、調査可能な包括的で採点済みのテストに変換することで、アプリケーション固有の AI 動作の評価を容易にします。
ASSERT は、AI モデルの期待される動作やポリシーに関する平易な記述を取り込み、それらを許容される動作と許容されない動作の構造化されたセットに変換し、問題シナリオやテストケースを生成して対象システムに対して実行し、結果にスコアを付けます。また、AI システムがたどったパス(中間アクションやツール呼び出しを含む)も記録するため、開発者は失敗が発生した箇所を検査することができます。
開発者は、評価範囲をさらにカスタマイズしたい場合、システムコンテキスト、ツール、制約も提供できます。
例えば、ドキュメント調査用 AI エージェントが社外の人へメールを送信しないこと、機密情報を役員レベルに限定すること、事前の文脈を考慮した簡潔な要約を提供することを指定できます。ASSERT はこれらのルールを使用して、システムが継続的にこれらのルールに従っているかを確認するテストケースを生成します。
image画像クレジット: Microsoft
Microsoft によると、このフレームワークは、AI モデルがアプリケーションや製品のコンテキスト、ポリシー、ツールによって形成された振る舞いをするように意図されている場合に、より広範で一般的な評価では埋められないギャップを埋めるものです。
「私たちが学んだことのひとつに、評価は良い意思決定を行うために絶対に不可欠であるということがあります」と、Microsoft の責任ある AI 担当チーフプロダクトオフィサーである Sarah Bird は述べています。「AI システムの振る舞いを理解していなければ、それが組織の基準を満たしているかどうかを知ることは本当に難しいからです […] 私たちが発見したのは、信頼性の高いシステムを本当に構築したいのであれば、アプリケーション固有の多くの次元について評価を行うべきだということです。」
Bird 氏は、ASSERT はシステムの構築中、デプロイ後、さらには継続的なモニタリング時にも使用できると述べています。
この発表は、AI業界における緩やかだが広範な変化の最中に行われたものです。モデルの能力が高まるにつれ、研究者らは反復可能なテストや回帰チェックに注力しており、Stanford's HELM や MLCommons' AILuminate、また異なる条件下でのモデルの挙動を測定するためのベンチマークを公開している評価グループである METR などがその例です。
*当記事内のリンクを通じてご購入いただいた場合、私たちは少額のコミッションを獲得する可能性があります。これは当社の編集の独立性には影響しません。
Ram への連絡や、彼からのアウトリーチの確認は、ram.iyer@techcrunch.com までメールを送信してください。
原文を表示
AI researchers and labs have advanced by leaps and bounds in evaluating AI models for everything from safety and compliance to sycophancy and alignment. But it appears companies and developers are faced with a new, specific need: making sure that their AI system behaves as intended for their specific product or service.
In a bid to make that testing process simpler, Microsoft on Tuesday took the wraps off ASSERT, short for Adaptive Spec-driven Scoring for Evaluation and Regression Testing.
The open-source framework, Microsoft says, makes evaluating application-specific AI behavior easy by using AI to turn high-level, natural-language descriptions of goals, policies, or intended behaviors into thorough, scored tests that can be investigated.
ASSERT takes plain-language descriptions of an AI model’s expected behavior and policies, turns them into a structured set of acceptable and unacceptable behaviors, generates problem scenarios and test cases, runs them against the target system, and scores the results. It can also record the paths the AI system takes, including intermediate actions and tool calls, so developers can inspect where failures happen.
Devs can provide system context, tools, and constraints, too, if they want to further customize what the evaluations cover.
For example, a developer could specify that a document research AI agent shouldn’t send emails to people outside the company, limit confidential information to C-level executives, and provide concise summaries with prior context in mind. ASSERT will use those rules to generate test cases that check whether the system follows those rules on an ongoing basis.

The framework, according to Microsoft, fills a gap that broader, more general evaluations cannot when AI models are intended to behave in a manner that is shaped by an application or product’s context, policies, and tools.
“One of the things we’ve learned is that evaluations are absolutely critical to making good decisions,” said Sarah Bird, chief product officer of Responsible AI at Microsoft. “Because if you don’t understand the behavior of the AI system, it’s really hard to know if it’s meeting your organization’s bar […] What we found is that if you really want to have a trustworthy system, you should evaluate many more dimensions that are application-specific.”
Bird said ASSERT can be used to evaluate systems when they’re being built, after deployment, and even for continuous monitoring.
The release comes amidst a gradual but broader shift in the AI industry. As models grow more capable, researchers are focusing on repeatable testing and regression checks, with Stanford’s HELM, MLCommons’ AILuminate, and evaluation groups like METR rolling out benchmarks to measure how models behave under different conditions.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Ram is a financial and tech reporter and editor. He covered North American and European M&A, equity, regulatory news and debt markets at Reuters and Acuris Global, and has also written about travel, tourism, entertainment and books.
You can contact or verify outreach from Ram by emailing ram.iyer@techcrunch.com.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み