Azure Content Understanding、GPT-5シリーズ対応と精度向上を発表
本文の状態
日本語全文を表示中
詳細モードで約15分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Microsoft Foundry Blog
Microsoft は Azure Content Understanding のサポート範囲を GPT-5 シリーズ全体に拡大し、グラウンディングと信頼性スコアの精度向上によりコスト削減と処理効率の両立を実現した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 04:11
AI深層分析
キーポイント
GPT-5 シリーズの包括的サポート開始
Azure Content Understanding が GPT-5 から GPT-5.5 の全シリーズ(標準、mini、nano)およびドキュメント、画像、動画、音声分析に対応するよう拡張された。
グラウンディングと信頼性スコアの刷新
新しい評価方法により推論トークンが最大 28% 削減され、LLM コストは最大 25% 低下した一方で、信頼性とグラウンディングの精度はそれぞれ最大 14%、3% 向上した。
ワークロード別モデル選定ガイドライン
品質とコストを調整するミキシングボードのように機能し、タスクに応じた最適なモデル選択(例:ドキュメント処理では GPT-5.1/5.2 がバランス型)を推奨している。
レガシー移行の円滑化とパイプライン最適化
既存のコンテンツ理解ワークフローを再設計することなく、古い基盤モデルから新世代モデルへ移行できる道筋と、大規模ファイル処理に対応した事前処理パイプラインを提供する。
モデル選択の最適化とコストパフォーマンス
ドキュメントや音声タスクではGPT-5.1またはGPT-5.2がバランス型として推奨され、画像分類にはGPT-5.1が適している。高品質なGPT-5.5はコストが約100%増加し、ミニモデルは安価だが精度が低下するため、用途に応じた選択が必要である。
重要な引用
This expanded model catalog enables organizations to choose the right level of intelligence for each workload, helping reduce costs for high-volume processing while preserving access to advanced reasoning capabilities where needed.
In our tested configurations, it consumed up to 28% fewer total inference tokens and the full-inference LLM cost decrease by up to 25% while improving confidence scores accuracy and grounding accuracy by up to 14% and 3% respectively
Think of model selection as a mixing board, with quality on your content and end-to-end cost as the two faders to adjust.
We crafted a new confidence scoring method that is applicable to a broader set of models.
編集コメントを表示
編集コメント
GPT-5 シリーズの段階的なラインナップと、コスト削減を伴う精度向上の数値は、実務におけるモデル選定の判断材料として極めて有用である。特にグラウンディング精度の改善は、AI の出力信頼性を担保する上で重要な指標となるため、導入検討者は注目すべき内容だ。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
エンタープライズ向けのコンテンツは、もはや単に人が消費するだけのものではありません。組織が文書、画像、音声、動画から情報を抽出し、それに基づいて行動を起こすために AI をますます活用するようになり、Azure Content Understanding は GPT-5 シリーズのサポートを拡大するとともに、グラウンディングと信頼性の向上を図り、より高い柔軟性、効率性、品質を実現します。
この拡張されたモデルカタログにより、組織は各ワークロードに最適な知能レベルを選択できるようになります。これにより、大量処理のコスト削減が可能になる一方、必要な場面では高度な推論能力へのアクセスも維持できます。また、各モデル向けに最適化されたパイプラインも提供されます。前処理機能により、単純な LLM ドキュメントパイプラインよりも大きなファイルや高品質なデータを扱えるようになります。グラウンディングスコアと信頼性スコアの生成により、自動検証が可能になり、スルー処理率の向上が図れます。さらに、古い基盤モデルが段階的に廃止される中で、コンテンツ理解ワークフローを再設計することなく、新しい世代のモデルへ移行するための明確な道筋も示します。
新機能
今回のリリースでは、ドキュメント、画像、動画、音声分析における標準、mini、nano モデルを含む GPT-5、GPT-5.1、GPT-5.2、GPT-5.4、GPT-5.5 シリーズ全体への Content Understanding のサポートが拡大されました。
同時に、今回のリリースでは、より高品質な出力を生成し、全体のコストを削減するための、更新されたグラウンディングと信頼性スコアリング手法が導入されました。テストした構成では、総推論トークン使用量が最大 28% 減少し、フル推論時の LLM コストは最大 25% 低下しました。その一方で、信頼性スコアの精度は最大 14%(AUROC で測定)、グラウンディングの精度は最大 3%(グラウンディング完全一致で測定)それぞれ向上しています。
タスクに最適なモデルを選ぶ
モデル選択は、コンテンツの品質とエンドツーエンドのコストという 2 つのフェーダーを調整するミキシングコンソールのようなものです。フォーム処理に優れたモデルが、動画セグメンテーションや音声分類、画像生成といった他のタスクでも優れているとは限りません。組織が AI を活用してコンテンツから情報を抽出するケースが増える中、精度、コスト、レイテンシ、スループット、コンプライアンス要件、地域ごとの展開要件のバランスを取るために、適切なモデルを選定することが極めて重要です。
入力タイプ、スキーマ定義、展開トポロジーのすべての組み合わせをベンチマークすることはできませんが、最も一般的なシナリオに対する私たちのテスト結果に基づき、以下のスタートポイントと推奨事項を示します。
顧客の皆様には、上記の表を活用して候補となるモデルを絞り込み、自社のデータで評価した上で最終決定を下すことをお勧めします。その際、予算、品質基準、地域での利用可否、スループットの目標、既存のリソースといった他の要因も併せて考慮してください。
表 1: 一般的なモデル選定ガイドライン
モダリティ別推奨モデル
文書処理の場合、バランス型は GPT-5.1 または GPT-5.2 が最適です。最高品質を求めるなら GPT-5.5 を選べば、コストが約 101% 高くなるものの品質は約 2% 向上します。一方、低コスト志向なら GPT-5.4 Mini が推奨され、バランス型である GPT-5.2 と比較して回答一致率のスコアが平均 2% 低下するものの、コストは約 50% 削減できます。
動画処理では GPT-5 または GPT-5.1 がバランス型の選択肢となります。この用途においては最高品質モデルも同様の性能を発揮します。低コスト版の GPT-5 Mini は GPT-5 よりも約 28% 安価ですが、全体の生成 F1 スコアが約 7% 低下します。
音声処理(注 1)では文書抽出と同様に扱われるため、バランス型には GPT-5.1 または GPT-5.2 を推奨しています。最高品質モデルの GPT-5.5 は GPT-5.2 よりも約 2% 高品質ですが、コストは約 101% 増加します。低コスト版の GPT-5.4 Mini はバランス型より約 48% 安価で利用できますが、回答一致スコアは約 2 ポイント低下します。
画像処理では GPT-5.1 がバランス型の推奨モデルです。最高品質を求めるなら GPT-5.5 を選択すると、コストが約 130% 高くなるものの品質は約 3% 向上します。分類タスクに限定すれば、GPT-5 Mini は GPT-5.1 よりも約 52% 安価ですが、分類 F1 スコアは約 12% 低下します。
注 1: 音声処理と文書抽出は類似したタスクであるため、現在は両者に同じモデルを推奨しています。
詳細なモデル品質とコスト分析については付録をご覧ください。
グラウンディングと信頼性スコアの改善
今回のリリースでは、グラウンディングの効率化を図り、基盤となる信頼性スコアリング手法を更新しました。これにより、複雑な後処理システムを構築しなくても、より有用な根拠データやランキング信号を取得できるようになりました。
グラウンディング:消費トークンの削減
更新されたグラウンディング・システムは、抽出データのソースをより効率的に特定します。すべてのモデルタイプにおいて、文書あたりの入力トークン数が 20〜30% 減少し、推論に要する総トークン数も 18〜28% 削減されました。
image 最新のアプローチであるグラウンディングの改善により、GPT-4.1 ではゼロショット推論に必要なトークン数が 28% 削減され、GPT-5.2 でも 23% の削減が実現しました。
これらの削減効果は、テストされたモデルとラベル付きサンプルの組み合わせ全体で、LLM の完全推論にかかるコストを 11〜25% 引き下げると同時に、P50 および P95 のレイテンシも短縮されました。グラウンディングの精度については、従来の手法と同等の水準を維持しています。
信頼性スコアの精度と汎用性の向上
より広範なモデル群に対応可能な新しい信頼性スコアリング手法を開発しました。AUROC(曲線下面積)を用いて評価したところ、この新システムは、さまざまな閾値において正解フィールドを不正解フィールドよりも上位にランク付けする信頼性を測定する指標として、GPT-4.1 では約 9%、GPT-5.2 では約 14% の向上を示しました。
image 信頼性スコアに依存するワークロードへの注意: GPT-5 Mini および GPT-5 Nano では、信頼性とモデル品質の間に顕著な低下が見られました。信頼性の質がワークフローにおいて重要である場合は、これらのモデルの使用は避けてください。
各フィールドのベースラインとなる信頼性スコアは複数の入力値を組み合わせるものであり、フィールドの種類によってスコアの分布が異なる場合があります。ストリーム処理を行う際は、フィールドごとに受容閾値を設定し、モデルを変更するたびに再調整を行う必要があります。
付録:詳細なモデル品質とコスト分析
本付録では、社内で行ったモデルの品質とコスト分析の詳細を提示します。これらの結果はすべての生産環境ワークロードに当てはまるわけではありませんが、モデル選定の指針としてお役に立てるはずです。
ドキュメント/音声:最適なバランスを見つける
データセット
評価用データセットは、構造化文書から半構造化文書、非構造化文書まで幅広く網羅しており、合計 71 のドキュメントタイプが含まれています。Content Understanding は、ラベル付きサンプルの有無に関わらずフィールド抽出をサポートしているため、ゼロ、一つ、五个、利用可能なすべてのラベル付きサンプルを使用する構成をテストし、それらの平均をとって各モデルの平均精度を算出しました。
品質は、配列やオブジェクトなどのコンテナフィールドを除くリーフフィールドとアナライザー全体でのマクロ平均として報告されます。評価対象となったドキュメントの平均ページ数は 3.35 ページでした。なお、ファイルが大幅に長い場合、スキーマが異なる場合、あるいはトレーニング例の戦略が異なる場合は、コストと品質のトレードオフ関係が異なる可能性があります。
両タスクともソースコンテンツから構造化フィールドを抽出する点で共通しているため、これらのドキュメント結果に基づいて現在の音声に関する推奨事項を導き出しています。別途、音声精度のグラフは公開していません。
以下の図において、GPT-5.1 と GPT-5.2 は、AI の品質とコストのバランスが最も優れたモデルとして際立っており、両者の Answer Match スコアは約 1% の差しかありません。一方、品質範囲の最上位にある GPT-5.5 は、GPT-5.2 よりも Answer Match が約 2% 向上しますが、推定平均コストは約 101% 増加します。
GPT-5.2 と比較すると、GPT-5.4 Mini は Answer Match を約 2% 下げる代わりに、推定コストを約 48% 削減しています。まずは GPT-5.1/5.2 または GPT-5.4 Mini から開始し、品質向上がプレミアム価格に見合う場合にのみ GPT-5.5 を追加することをお勧めします。
imageGPT-5.1 と GPT-5.2 は、文書コストと品質のバランスを示すフロンティアを形成しています
動画:最大のモデルが最良とは限らない
データセット
今回の動画評価では、2 つの異なるワークロード形状を組み合わせています。1 つはセグメンテーションに焦点を当てた 60 分間の動画データセットで、もう 1 つは平均して 1 分にも満たない短尺の全編動画データセットです。生成 F1 スコアは、スカラー値の回答フィールドとタイムスタンプ付きのカスタムセグメント(例:ロゴが画面に表示されるタイミングなど意味的に重要な時間窓や、ニュースセグメントの長さ、広告ブレイクの期間など)の両方をカバーして計算しています。
このベンチマークは、長尺メディアから動画レベルの事実情報と特定のセグメントを抽出するワークロードにおいて最も関連性が高いものです。短編クリップや異なるフレーム密度、あるいはセグメンテーションを伴わないスキーマを使用する場合、ランキング結果が異なる可能性があります。
結果
図に示す通り、GPT-5 の全体生成 F1 スコアは 89.0%、GPT-5.1 は 88.6%、GPT-5.4 は 87.5% を記録しました。このユースケースにおいては、GPT-5 と GPT-5.1 がバランスと品質の両面で最適なペアとなります。また、リリース概要によると、GPT-5 のコストは GPT-4.1 ベースラインと比較して約 18% 削減されています。
より低コストな動画オプションをお求めの場合は、GPT-5 Mini から始めることを推奨します。セグメンテーションに特化したベンチマークでは GPT-5 よりも約 28% コストが抑えられますが、その分、全体の生成 F1 スコアは約 7% 低下します。
image GPT-5 は、ベンチマークコストを 18% 削減しながらも、ベースラインの動画セグメンテーション品質と同等の性能を発揮します。
画像:分類(Classify)と生成(Generate)タスクには異なるモデルを選択してください
データセット
画像評価では、5 回の反復を含む 17 のゼロショットデータセットが対象となりました。これらは製品、衣類、産業欠陥、シーン、オブジェクト指向コンテンツなど、多岐にわたる分類および生成タスクを網羅しています。
分類の評価指標には F1 スコアを採用しました。一方、生成分野では 1~7 の評価基準(ルブリック)スコアを使用しているため、これら 2 つの品質軸は異なる問いに答えるものであり、単一の勝者を決めるために統合してはいけません。
結果
図に示す通り、GPT-5.1 は分類タスクで 69.5% の Classify F1 を記録し、首位となりました。このため、分類用途には魅力的な選択肢と言えます。生成タスクにおいては GPT-5.5 が 7 点満点中 5.0 で首位に立ちましたが、これは GPT-5 Mini よりも約 3% 高いスコアです。ただし、その性能向上にはコストが約 375% 増えるという代償を伴います。

コストを抑えた分類オプションとして、GPT-5 Mini は GPT-5.1 よりも約 52% 安価ですが、Classify F1 スコアは約 12% 低下します。分類を主目的とするワークロードには GPT-5.1 を、生成が中心のワークロードには GPT-5 Mini を使用することをお勧めします。必要に応じて、GPT-5.5 の品質向上がプレミアム価格に見合うかどうかを検証してください。
image GPT-5.1 は画像分類で首位に立ちながら、GPT-4.1 ベースラインよりも低コストを実現しています
では、チームは次に何をすべきでしょうか。
カタログが拡張されたことで、ミキシングボードの各チャンネルがより広い選択肢を得ました。次のステップとして、ワークロードに適した候補を 2〜3 選定し、品質・コスト・可用性・スループットのどの組み合わせが最適かをテストして決定してください。
比較はシンプルに保ちましょう。すべての実行で同じアナライザー、スキーマ、代表となる入力セット、ラベル付きサンプルを使用します。変更するのはモデルのデプロイのみとし、出力品質、レイテンシ、トークン使用量、失敗率を比較します。テスト前に必ず supportedModels 応答を確認し、各候補があなたのリージョンで利用可能であることを確認してください。
Content Understanding Studio で始める
Content Understanding Studio の「設定」を開き、Foundry リソースを追加して、デフォルトのモデルデプロイメントを構成します。Studio は、適切なデフォルトが存在しない場合に必要なモデルを自動的にデプロイできます。
Content Understanding Studio のクイックスタートガイドに従って、解析器を選択し、代表的なコンテンツで実行してください。
候補となるモデルはすべて同一のファイルに対してテストします。抽出されたフィールドと生レスポンスを確認し、ワークロードにとって重要な品質、レイテンシ、使用量を記録しましょう。
あるいは REST API を通じてデプロイメントを比較することも可能です。
再現性のある評価を行うには、各 analyze リクエストで異なる modelDeployments マッピングを渡します。リクエストレベルでのマッピングはリソースのデフォルトを上書きするため、解析器と入力を変更せずに完了用デプロイメントだけを差し替えることができます。
例えば、以下のように呼び出すことができます。
POST /contentunderstanding/analyzers/myInvoice:analyze
{
"inputs": [
{
"url": ""
}
],
"modelDeployments": {
"prebuilt-analyzer-completion": "",
"prebuilt-analyzer-embedding": ""
}
}
各候補に対して同じリクエストを1回ずつ実行し、抽出結果とレスポンスの使用量データを比較します。まずはバランスの取れた推奨モデルから始め、コストを抑えたオプションを追加し、ワークフローにおいて精度の差が重要になる場合のみ最高品質のモデルを含めるようにしてください。
関連リンク
最新の Content Understanding プレビューリリースにおける新機能(同期 API の導入や GPT-5 モデルシリーズ、エージェント機能への対応など)についてはこちらをご覧ください。
CU Studio で Content Understanding を試すには – クイックスタート
Content Understanding Studio または Foundry ポータルで試すには – Foundry ツール | Microsoft Learn
CU API と SDK の利用については、以下のクイックスタートガイドをご覧ください。
Azure Content Understanding in Foundry Tools – Foundry Tools | Microsoft Learn
原文を表示
Enterprise content is no longer just something people consume. As organizations increasingly rely on AI to extract and act on information from documents, images, audio, and video, Azure Content Understanding is expanding support for the GPT-5 series and improving grounding and confidence to deliver greater flexibility, efficiency, and quality.
This expanded model catalog enables organizations to choose the right level of intelligence for each workload, helping reduce costs for high-volume processing while preserving access to advanced reasoning capabilities where needed. It also provides optimized pipelines tuned for each model. Preprocessing allows the models to support larger files and higher quality than the simple LLM document pipelines. Generating grounding and confidence scores enables automated validation and higher straight-through processing rates. It also provides a clear path forward as older foundation models retire, allowing customers to transition to newer generations of models without redesigning their Content Understanding workflows.
What’s New
This release expands Content Understanding support to the GPT-5, GPT-5.1, GPT-5.2, GPT-5.4, and GPT-5.5 series including standard, mini, and nano models across document, image, video, and speech analysis.
Just as important, the release introduces an updated grounding and confidence scoring method that generates higher quality outputs and reduces overall cost. In our tested configurations, it consumed up to 28% fewer total inference tokens and the full-inference LLM cost decrease by up to 25% while improving confidence scores accuracy and grounding accuracy by up to 14% and 3% respectively (measured by AUROC and grounding exact match).
Choosing the Best Model for Your Task
Think of model selection as a mixing board, with quality on your content and end-to-end cost as the two faders to adjust. A model that excels on forms may not lead on other tasks such as video segmentation, speech classification, or image generation. As organizations increasingly leverage AI to extract information from content selecting the right model is critical for balancing accuracy, cost, latency, throughput, compliance, and regional deployment requirements.
While we cannot benchmark every combination of input types, schema definitions, and deployment topologies, below is a set of starting points and recommendations based on our testing of the most common scenarios.
We encourage customers to leverage the table above to choose a model short list, and then evaluate on your own data to make the decision. Meanwhile, you should consider other factors such as your budget, quality bar, regional availability, throughput target, and existing capacity.
Table 1: General Model Selection Guidelines
Modality
Balanced recommendation
Best quality
Lower-cost choice
Document
GPT-5.1 or GPT-5.2
GPT-5.5 about +2% better quality at about 101% higher cost
GPT-5.4 Mini costs about 50% less with an average –2% lower quality on our answer match metric than balanced GPT-5.2
Video
GPT-5 or GPT-5.1
Same as balanced for this use case
GPT-5 Mini costs about 28% less than GPT-5 with about –7% lower overall Generation F1
Speech1
GPT-5.1 or GPT-5.2
GPT-5.5 about +2% better quality than GPT-5.2 at about 101% higher cost
GPT-5.4 Mini costs about 48% less with about –2 Answer Match points lower quality than balanced GPT-5.2
Image
GPT-5.1
GPT-5.5 delivers about +3% better quality for about 130% higher cost
For classification, GPT-5 Mini costs about 52% less than GPT-5.1 with about –12% lower Classify F1
1 Note: Speech and document extraction are similar tasks, so we currently recommend the same models for both.
Detailed model quality and cost analysis is included in the Appendix.
Grounding and Confidence improvements
In this release, we also improved grounding efficiency and refreshed the underlying confidence scoring method, so customers can get more useful evidence and ranking signals without building complex post-processing systems.
Grounding: Fewer Tokens Consumed
The updated grounding system more efficiently identifies the source for the extracted data. Across all model types, we observed 20-30% fewer input tokens and 18-28% fewer total inference tokens per document.
imageUpdated grounding reduced zero-shot inference tokens by 28% for GPT-4.1 and 23% for GPT-5.2
Those reductions lowered the full-inference LLM bill by 11-25% across the tested model-and-labeled-sample configurations and reduced P50 and P95 latency. Grounding accuracy remains similar to the previous method.
Confidence: Improved Accuracy and Generalizability
We crafted a new confidence scoring method that is applicable to a broader set of models. Measured by AUROC, which measures how reliable the confidence scores rank correct fields above incorrect fields across various threshold, the new confidence scoring system improved by about 9% for GPT-4.1 and 14% for GPT-5.2 against the previous method.
imageConfidence-sensitive workloads: We observed a significant confidence-model quality drop with GPT-5 Mini and GPT-5 Nano. Avoid these models when confidence quality is important to your workflow.
A field’s baseline confidence score combines several inputs, and score distributions can differ by field type. For straight-through processing, set acceptance thresholds field by field and recalibrate them whenever you switch models.
Appendix: Detailed Model Quality/Cost Analysis
In this Appendix, we present more details on the model quality and cost analysis conducted in-house. These results may not be representative for every production workload, but it should be helpful to the reader as a guidance for model selection.
Documents/Speech: Finding the Right Balance
Dataset
The evaluation dataset span from structured to semi-structured to unstructured documents, with a total of 71 document types. Since Content Understanding supports field extraction with and without labeled samples, we tested configurations where there are zero, one, five, and all available labeled samples, and average across them to calculate the average accuracy of a given model.
Quality is reported as a macro average across leaf fields and analyzers, excluding container fields such as arrays and objects. The evaluated documents averaged 3.35 pages. Note workloads with substantially longer files, different schemas, or different training-example strategies may see a different cost and quality frontier.
We use these document results to guide the current speech recommendation because both tasks extract structured fields from source content. We are not publishing a separate speech accuracy graph.
Results
In the figure below, GPT-5.1 and GPT-5.2 stand out as well-balanced models between AI quality and cost, and they differ by about 1% in Answer Match score. At the top of the quality range, GPT-5.5 delivers about 2% more Answer Match than GPT-5.2, but increases estimated average cost by about 101%.
Compared with GPT-5.2, GPT-5.4 Mini reduces estimated cost by about 48% while giving up about 2% in Answer Match. We recommend customers to start with GPT-5.1/5.2 or GPT-5.4 mini, then add GPT-5.5 when its quality improvement can justify the premium.
imageGPT-5.1 and GPT-5.2 form the balanced document cost-quality frontier
Video: the Biggest Model is not the Best Model
Dataset
The video evaluation combined two distinct workload shapes: a segmentation-focused, 60-minute video dataset, and a short whole-video dataset averaging just under one minute. We compute generation F1 score covering both scalar answer fields and timestamped custom segments, e.g., semantically relevant time windows such as when a logo appears on screen, the duration of a news segment, or an ad break.
This benchmark is most relevant to workloads that extract both video-level facts and segments from long media. Short clips, different frame density, or schemas without segmentation may produce a different ranking.
Results
As shown in the figure below, GPT-5 reached 89.0% overall Generation F1, GPT-5.1 reached 88.6%, and GPT-5.4 reached 87.5%. GPT-5 and GPT-5.1 form both the balanced and best-quality pair for this use case; GPT-5 also cost about 18% less than the GPT-4.1 baseline in the release summary.
For a lower-cost video option, we recommend customers start with GPT-5 Mini. It cost about 28% less than GPT-5 on the segmentation-focused benchmark while giving up about 7% overall Generation F1.
GPT-5 matches baseline video segmentation quality at 18% lower benchmark cost
imageGPT-5 matches baseline video segmentation quality at 18% lower benchmark cost
Images: Select Different Models for Classify and Generate tasks
Dataset
The image evaluation covered 17 zero-shot datasets with five repeats. The set spanned classification and generative tasks across product, apparel, industrial-defect, scene, and object-oriented content.
Classification used F1 as metrics. Generative fields used a 1-7 rubric score, so the two quality axes answer different questions and should not be collapsed into one winner.
Results
As shown in the figure below, GPT-5.1 led classification at 69.5% Classify F1, making it an attractive choice for that use case. For generation, GPT-5.5 led at 5.0 out of 7, about 3% higher than GPT-5 Mini at about 375% higher cost.

For a lower-cost classification option, GPT-5 Mini cost about 52% less than GPT-5.1 while giving up about 12% Classify F1. We recommend customers to start with GPT-5.1 for classification-led workloads, and GPT-5 Mini for generation-led workloads. If necessary, test if GPT-5.5’s quality gain could justify the premium.
imageGPT-5.1 leads image classification while costing less than the GPT-4.1 baseline
So what should teams do next?
The expanded catalog gives every channel on the mixing board a wider range. The next step is to choose two or three candidates for your workload and test which combination of quality, cost, availability, and throughput deserves the final setting.
Keep the comparison simple: use the same analyzer, schema, representative input set, and labeled examples for every run. Change only the model deployment, then compare output quality, latency, token usage, and failure rate. Before testing, check the analyzer’s supportedModels response and confirm that each candidate is available in your region.
Start in Content Understanding Studio
In Content Understanding Studio, open Settings, add your Foundry resource, and configure its default model deployments. Studio can deploy required models automatically when no suitable default exists.
Follow the Content Understanding Studio quickstart to select an analyzer and run it on your own representative content.
Test each candidate against the same files. Review the extracted fields and raw response, and record the quality, latency, and usage that matter to your workload.
Or compare deployments through the REST API
For a repeatable evaluation, pass a different modelDeployments mapping in each analyze request. A request-level mapping overrides the resource defaults, so you can keep the analyzer and inputs unchanged while swapping the completion deployment.
For example, you can call:
POST /contentunderstanding/analyzers/myInvoice:analyze
{
"inputs": [
{
"url": "<representative-input-url>"
}
],
"modelDeployments": {
"prebuilt-analyzer-completion": "<candidate-deployment-name>",
"prebuilt-analyzer-embedding": "<embedding-deployment-name>"
}
}
Run the same request once per candidate, then compare the extracted results and the response’s usage data. Start with the balanced recommendation, add the lower-cost option, and include the best-quality model only when the remaining accuracy gap matters to the workflow.
Additional Links
For more on new capabilities in the latest Content Understanding Preview release – From Sync APIs to support for the GPT-5 model series and agentic
To try out Content Understanding in CU Studio – Quickstart Try out Content Understanding Studio or Foundry portal – Foundry Tools | Microsoft Learn
To use the CU APIs and SDK see – Quickstart: Azure Content Understanding in Foundry Tools – Foundry Tools | Microsoft Learn
The post Azure Content Understanding GPT-5 Series Guide: Model Selection, Grounding Improvements, and Confidence Enhancements appeared first on Microsoft Foundry Blog.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み