Azure Content Understanding、GPT-5 シリーズとエージェントワークフロー対応へ
本文の状態
日本語全文を表示中
詳細モードで約16分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Microsoft Foundry Blog
Microsoft は Azure Content Understanding の更新を発表し、GPT-5 シリーズの包括的サポート、同期 API の導入、およびエージェント型ドキュメント推論機能を追加した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 04:11
AI深層分析
キーポイント
GPT-5 シリーズの包括的サポート拡大
CU 1.0 API が GPT-5.5 から GPT-5 シリーズ全体、および標準・ミニ・ナノバリアントをサポートするよう拡張され、開発者は精度、レイテンシ、コストのトレードオフを柔軟に選択できる。
同期 API と新機能の導入
CU 2.0 のパブリックプレビューにより、Read および Layout 用の同期 API が追加され、高度な文脈化、セマンティックチャンキング、およびエージェント型ドキュメント推論が可能になる。
生産ワークロードの効率化
CU 1.0 の刷新によりグラウンディング効率が向上し、トークン使用量が削減されると同時に、信頼性スコアリングモデルが改善されたことで既存の生産環境での利用が強化される。
抽出とグラウンディングの統合による効率化
処理フロー内で抽出とグラウンディングを効果的に統合することで、GPT-4.1およびGPT-5.2で推論トークン使用量が最大28%削減され、精度は最大3%向上した。
信頼性モデルの刷新
内部評価においてAUROCによる測定精度がGPT-4.1およびGPT-5.2で最大14%改善し、自動化パイプラインの効率性と監査可能性を高める。
重要な引用
Content Understanding analyzers can now use the GPT-5 model series, including GPT-5.5... across standard, mini, and nano variants.
CU 2.0 adds synchronous Read and Layout APIs, advanced contextualization, semantic chunking, improved classification, new prebuilt analyzers, and agentic document reasoning.
Confidence scores are a critical part of production automation. They help applications decide when to move straight through, when to route to human review, and when to apply additional validation.
Customers can evaluate accuracy, latency, and cost on representative data before deciding which model deployment and API version best fit their production needs.
編集コメントを表示
編集コメント
GPT-5 シリーズへの対応拡大は、開発者がコストと精度のバランスを最適化する上で極めて重要なステップである。特に同期 API の導入により、リアルタイム性が求められるアプリケーションにおけるドキュメント処理の実現可能性が高まっている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
エンタープライズコンテンツはもはや単に人が読むだけのものではありません。AI アプリやエージェントの有用性は、それらが理解できる情報の質にかかっています。しかし、世界中の企業知識の多くはまだ文書、フォーム、表、画像、音声、動画の中に閉じ込められています。
最新の Azure Content Understanding のアップデートでは、カスタム処理を減らしながら、こうしたコンテンツを構造化された信頼性の高いデータに変換できるようになります。今回のリリースでは、GPT-5 モデルシリーズへの対応拡大やトークン使用量の削減、精度の高い自信スコアリングの導入、Read API と Layout API の同期化、高品質な抽出のための高度な文脈理解機能、税務関連の事前構築済み解析器、検索ワークフロー向けのセマンティックチャンキング、そして複雑な抽出シナリオに対応するエージェント型ドキュメント推論機能が追加されました。
具体的には、Foundry Tools 内の Azure Content Understanding について以下の 2 つのアップデートを発表します。
- CU 1.0 API の刷新版が一般提供(GA)となりました。CU 1.0 では、GPT-5 シリーズへの対応範囲拡大、トークン使用量の削減、およびグラウンディングと自信スコアリングの精度向上が実現されています。
- 次世代ドキュメント理解を探求する開発者向けに、CU 2.0 の公開プレビューを開始しました。CU 2.0 では、同期型の Read API と Layout API、高度な文脈化機能、セマンティックチャンキング、分類性能の向上、新規事前構築済み解析器、そしてエージェント型ドキュメント推論機能が追加されています。
今回のアップデートにより、Azure Content Understanding (CU) は、現在の生産環境での効率性が向上するだけでなく、開発者が今後構築できるドキュメント自動化、検索、推論のシナリオの範囲も拡大します。
imageAzure Content Understanding の新機能概要
CU 1.0:生産ワークロードにおけるコスト効率と信頼性の向上
CU 1.0 GA API (2025-11-01) に大幅な改良を加え、既存の生産ワークフローで GPT-5 シリーズのより広範なサポート、グラウンディング(根拠付け)の効率化、そして更新された信頼度モデルを活用できるようになりました。以下の 3 つの改善点は、それぞれ生産プロセスの異なる側面——実行するモデル、発生するトークンコスト、結果への信頼性——に焦点を当てています。
GPT-5.5 および GPT-5 シリーズのサポート拡大
Content Understanding の分析機能は now、標準版、ミニ版、ナノ版を含む、GPT-5.5、GPT-5.4、GPT-5.3、GPT-5.2、GPT-5.1、および GPT-5 シリーズ全体をサポートするようになりました。
これにより、開発者はワークロードに最適なモデル展開を選択できるようになります。シナリオによっては最大精度を最優先する場合もあれば、レイテンシーやコストを最優先するケースもあります。GPT-5 シリーズのサポート拡大によって、CU は顧客が独自のデータでこれらのトレードオフを検証しながら、同じ分析器モデルと API パターンを使い続けることを可能にします。既存の PTU 契約もそのまま利用できます。
接地精度の向上
GA リフレッシュでは、抽出処理と接地処理をフロー内でより効果的に統合することで、接地効率も改善されました。内部評価によると、GPT-4.1 および GPT-5.2 で平均推論トークン使用量が最大 28% 削減され、同時に平均精度が最大 3% 向上しました。
高ボリュームの抽出ワークフローを構築する顧客にとって、この改善は二つの点で重要です。第一に、トークン使用量の低下により、LLM を活用した抽出処理の実行コストを抑えられます。第二に、接地効率の向上により、追加のポストプロセッシングロジックを追加しなくても、ソースコンテンツへの追跡可能性を維持できます。
信頼度モデルの刷新
GA リフレッシュには、刷新された信頼度モデルも含まれています。内部評価では、AUROC で測定した精度が GPT-4.1 および GPT-5.2 で最大 14% 向上しました。
信頼スコアは、生産自動化において極めて重要な要素です。これによりアプリケーションは、処理をそのまま進めるべきか、人間によるレビューへルーティングすべきか、追加の検証が必要かを判断できます。より強力な信頼モデルを持つことで、チームは効率的かつ監査可能な自動化パイプラインを構築しやすくなります。
特定のタスクに最適なモデルの選定方法や、品質ベンチマークの詳細については、「Azure Content Understanding GPT-5 シリーズガイド:モデル選択、グラウンディングの改善、信頼性の向上」をご覧ください。
CU 2.0 プレビュー:AI アプリとエージェントのための新ビルドブロック
刷新された GA API で利用可能な機能に加え、新しい CU 2.0 プレビュー(2026-06-01-preview)も導入します。
CU 2.0 プレビューは、開発者の二つの主要な優先事項を前進させます:
コンテンツパイプラインの品質と効率性の向上、および
低遅延かつ推論集約的なシナリオの実現です。
以下の機能は、刷新された GA API と同じ GPT-5 シリーズ基盤の上に構築されています。
GPT-5.5 およびそれより下のモデルがプレビューでサポートされる
CU 2.0 パブリックプレビューでは、GPT-5 モデルシリーズもサポートします。つまり、顧客は刷新された GA パスと同じモデルシリーズの方向性を用いて、新しいプレビュー機能をテストできます。
これは移行計画において重要です。顧客は、どのモデルデプロイメントと API バージョンが生産環境に最も適しているかを決定する前に、代表的なデータ上で精度、遅延、コストを評価できます。
生産ワークロードの品質と効率を向上させる
新しい CU 2.0 プレビュー版の第一の目標は、すでに本番環境で運用されているチームの業務改善です。抽出精度の向上、トークンコストの削減、そして検索とルーティングの信頼性強化がその柱となります。この方向性を支える 5 つのプレビュー機能のうち、まずは「Advanced Contextualization(高度な文脈化)」から紹介します。
高度な文脈化
高度な文脈化は、カスタム分析器がラベル付きサンプルやドキュメント知識をより効率的に活用できるよう支援します。社内評価では、平均精度が最大 3.5%向上し、同時に LLM のトークン使用量が最大 22%削減されました。トレーニングデータは顧客の Azure Storage アカウント内に留まり、分析器へコピーされるのではなく、知識源として参照されるため、顧客によるストレージ管理権限が維持されます。その結果、より高精度な構造化抽出が可能になりつつ、必要なトークン数とデータ管理の手間を大幅に減らすことができます。
トレーニングデータの管理に関する詳細は、「分析器の改善」セクションをご覧ください。
このアプローチは、新たに導入された 5 つのプリビルト分析器でも採用されています。これらのプリビルト分析器では、LLM のトークン消費量を最大 99%削減できます。これにより、品質とコストの両方が重要な大規模シナリオにおいて、新しいプリビルト分析器の実用性がさらに高まります。
要約すると、高度な文脈理解機能は単なる品質向上の要素ではなく、生産効率を高める機能です。これにより、顧客はより構造化されたデータ抽出が可能になり、LLM のトークン使用量を削減しながら、トレーニング用の入力データを自社のストレージガバナンス内で管理できます。
税務およびその他の文書タイプ向けの新しい事前構築型アナライザー
今回のプレビューでは、高度な文脈理解機能を基盤とした新たな事前構築型税務用アナライザーが導入されました。これにより、個別の税務申告書のサポートから、企業レベルや州レベルの税務ワークフローまで対応範囲が拡大します。対象となる書類には、1065 号書、1120-S 号書、8865 号書、1041 Schedule K-1、およびミネソタ州 M1 申告書が含まれます。
これらの新アナライザーは複雑な多ページレイアウトに対応し、高度な文脈理解を活用することで、高い抽出精度と低遅延、競争力のある価格を実現します。これにより LLM のトークン消費量が劇的に削減され、一部の抽出シナリオでは LLM トークンを全く使用しないことも可能です。これらの事前構築型アナライザーは、Content Understanding Studio や Microsoft Foundry でも利用でき、開発者は本番環境への展開前に、値の視覚的確認や文書内の位置情報(バウンディングボックス)の確認、信頼度スコアのチェックを行うことができます。利用可能な事前構築型アナライザーの完全なリストおよび、それらの使用方法とカスタマイズに関するガイダンスについては、事前構築型アナライザーのドキュメントをご覧ください。
image 事前構築された分析器「Schedule K-1 prebuilt analyzer (prebuilt-tax.us.1065ScheduleK1)」は、複雑な税務書類を構造化された JSON 形式に変換し、後続の処理を可能にします。
事前構築された文書検索におけるセマンティックチャンキング
検索精度は、コンテンツをどのようにチャンク(断片)化するかに大きく依存します。固定サイズのチャンキングでは、表が見出しから切り離されたり、関連する段落が分断されたり、複数ページにわたる構造が検索システムにとって扱いにくい断片化されてしまう恐れがあります。
CU 2.0 のパブリックプレビューでは、事前構築された文書検索機能に「セマンティックチャンキング」が追加されました。文字数やページの境界線だけで分割するのではなく、ドキュメントの構造を活用してより意味のある検索単位を作成します。これにより、文や段落をまたぐ文脈や意味的な関係性が保持され、特に RAG(検索拡張生成)やエージェント型検索において、取得した情報の質が推論結果の質に直結するシナリオで威力を発揮します。
分類:ページ内での分割と信頼度
実際の企業向け提出書類は、必ずしも物理的なページの境界線と一致するとは限りません。ローンパッケージや税務申告書、ケースファイル、スキャンされた資料などには複数の論理的な文書が含まれていることがあり、1 ページの中に別のセクションの末尾と次のセクションの冒頭が混在していることも珍しくありません。
CU 2.0 のパブリックプレビューでは、ページ内分割機能による分類精度が向上しました。これにより、文書全体ではなく、より細かなセグメント単位で分類を識別できるようになります。
また、分割と分類の信頼度スコアも追加されました。これらの指標を活用すれば、アプリケーションは自動的にセグメントをルーティングすべきタイミングや、下流の分析ツールを呼び出すべきタイミング、そして人間による確認が必要な結果を送信するタイミングを判断できます。
これにより、文書自動化パイプラインにおける分類機能は、単なるベストエフォートのルーティングステップから、実務に役立つ制御ポイントへと進化します。詳細については「分類機能の強化」をご覧ください。
レイアウト解析における署名検出とメタデータ抽出
分類機能は「文書が何であるか」を判断する一方、レイアウト解析器は「その内部に何が含まれているか」をより深く捉えます。新たに追加されたのは、署名の検出と文書メタデータの抽出です。
署名検出機能により、契約書やフォーム、請求書、署名付き提出書類などにおいて、署名領域とその位置を特定できます。一方、メタデータ抽出では、作成者、タイトル、作成日、コンテンツタイプ、言語といった利用可能な文書プロパティを把握できます。
これら新機能は、レイアウト解析器が既に備えているセクション、見出し、書式設定、表、図、ハイパーリンク、注釈、その他のレイアウト要素の抽出能力と組み合わさることで、単一の解析器から文書のコンテンツとコンテキストをより包括的に表現できるようになります。これにより、インテリジェントな文書処理やエージェント型ワークフローの構築が容易になります。
詳細は、レイアウト解析器のドキュメントをご覧ください。
Content Understanding Studio で署名検出とメタデータ抽出を試すことができます。
imageAzure Content Understanding Studio における署名検出。元の文書と併せて、検出された署名領域が表示されています。
Content Understanding の解決範囲の拡大
このプレビュー API の 2 つ目の目標は、標準的な抽出を超えた新たなユースケースへの対応です。その鍵となるのが、同期処理と、難易度の高い抽出タスクに対するエージェント型推論機能の 2 つです。
読み取りおよびレイアウト API
顧客とのやり取り中に AI エージェントを補完する、本人確認文書の検証を行う、ファイル提出時にワークフローを開始するなど、即座にレスポンスが必要なドキュメント処理ワークフローがあります。CU 2.0 プレビューでは、これらの低遅延シナリオに対応するため、読み取り (prebuilt-read) およびレイアウト (prebuilt-layout) 解析器に対して同期オペレーションが追加されました。これにより、構造化された結果をレスポンス内に直接返すことが可能になります。
ドキュメントはバイナリコンテンツまたは URL で提出でき、一時的なサーバー側のストレージを使用せずに処理されます。詳細については、「Azure Content Understanding が同期オペレーションを発表」をご覧ください。
複雑なフィールド抽出のためのエージェントモード
同期 API が「速い応答」を目指すなら、エージェントモードは「深い推論」を実現するものです。文書抽出タスクの中には、単なる一度の読み込みでは対応できないケースがあります。答えが長い文書のあちこちに散在する証拠に基づいて導かれる場合や、数値を比較・検証する必要がある場合、あるいは最終出力を出す前に中間結果に対する推論が必要な場合があります。
こうしたシナリオに対応するため、CU 2.0 プレビューではエージェントモードを導入しました。
エージェントモードは、複雑な文書理解のための反復的な抽出ワークフローを採用しています。これは、標準的な抽出手法では不十分なケース、つまり長文の法務契約や財務報告書、保険記録、あるいは関連する証拠が複数のセクションに分散しているその他のドキュメントなど、より難易度の高い抽出シナリオ向けに設計されています。
エージェントモードは分析スキーマと連携し、追加的な推論によって関連コンテンツを特定し、値の抽出を行い、中間結果を評価して最終出力を精緻化します。追加の推論処理を行うため、標準的な抽出と比較してレイテンシやトークン消費量が増加する可能性があります。顧客は、代表的なドキュメントでエージェントモードの評価を行い、品質向上が見込める場合にのみ、追加のコストと処理時間を考慮して利用を検討してください。エージェントモードの詳細についてはこちらをご覧ください。
はじめ方
今回の 2 つの更新は併用を想定しています。一つは今日から本番環境で活用するもの、もう一つは次なるステップを検証するためのものです。ご自身のワークロードに最適な道筋を選択してください。
刷新された CU 1.0 GA API を利用して、既存のワークロードを改善しましょう。GPT-5 シリーズへの対応やグラウンディング効率の向上、そして刷新された信頼性モデルなどが含まれます。
CU 2.0 のパブリックプレビューでは、次世代の CU キャパビリティを評価できます。同期型の Read API と Layout API、高度な文脈化(Advanced Contextualization)、セマンティックチャンキング、新しいプリビルトアナライザー、改善された分類機能、そしてエージェントモードなどが対象です。
プリビルトアナライザーやカスタムアナライザー、構造化出力について詳しく知りたい場合は、Content Understanding Studio または Microsoft Foundry から始めてください。
Quickstart Content Understanding Studio or Foundry | Microsoft Learn
Quickstart: Content Understanding SDK and Rest| Microsoft Learn
原文を表示
Enterprise content is no longer just something people read. AI apps and agents are only as useful as the information they can understand, yet much of the world’s enterprise knowledge is locked in documents, forms, tables, images, audio, and video. The latest Azure Content Understanding updates help developers turn that content into structured, grounded data with less custom processing. This release brings broader support for the GPT-5 model series, lower token usage, improved confidence scoring, new synchronous APIs for Read and Layout, advanced contextualization for higher-quality extraction, new tax-focused prebuilt analyzers, semantic chunking for retrieval workflows, and agentic document reasoning for more complex extraction scenarios.
Specifically, we’re announcing two updates to Azure Content Understanding in Foundry Tools:
A refreshed CU 1.0 API, now generally available for production workloads. CU 1.0 adds broader GPT-5 series support, lower token use, and improved grounding and confidence scoring.
A new CU 2.0 public preview for developers exploring next-generation document understanding. CU 2.0 adds synchronous Read and Layout APIs, advanced contextualization, semantic chunking, improved classification, new prebuilt analyzers, and agentic document reasoning.
Together, these updates make Content Understanding (CU) more efficient for production workloads today while expanding the range of document automation, retrieval, and reasoning scenarios developers can build tomorrow.
imageOverview of new Azure Content Understanding features.
CU 1.0: Better economics and reliability for production workloads
We’ve made significant enhancements to the CU 1.0 GA API (2025-11-01), allowing customers to take advantage of broader GPT-5 series support, improved grounding efficiency, and a refreshed confidence model in their existing production workflows. Each of the three improvements below targets a different part of that production path: the model you run, the tokens it costs, and the confidence you can place in the result.
Expanded support for GPT-5.5 and lower GPT-5 series model support
Content Understanding analyzers can now use the GPT-5 model series, including GPT-5.5, GPT-5.4, GPT-5.3, GPT-5.2, GPT-5.1, and GPT-5 series models, across standard, mini, and nano variants.
This gives developers more flexibility to choose the right model deployment for their workload. Some scenarios prioritize maximum accuracy. Others prioritize latency or cost. By expanding GPT-5 series support, CU enables customers to evaluate these tradeoffs on their own data while continuing to use the same analyzer model and API patterns. Customers can use existing PTU commitments.
Improved grounding efficiency
The GA refresh also improves grounding efficiency by merging extraction and grounding more effectively in the processing flow. In internal evaluations, this reduced average inference token usage by up to 28 percent for GPT-4.1 and GPT-5.2, while also improving average accuracy by up to 3 percent.
For customers building high-volume extraction workflows, this matters in two ways. First, lower token usage can reduce the cost of running LLM-backed extraction. Second, better grounding efficiency helps preserve traceability to the source content without requiring developers to add extra post-processing logic.
Refreshed confidence model
The GA refresh includes a refreshed confidence model. In internal evaluations, accuracy measured by AUROC improved by up to 14 percent for GPT-4.1 and GPT-5.2.
Confidence scores are a critical part of production automation. They help applications decide when to move straight through, when to route to human review, and when to apply additional validation. A stronger confidence model makes it easier for teams to build automated pipelines that are both efficient and auditable.
For guidance on how to choose the right model for specific tasks and more details on our quality benchmarks see Azure Content Understanding GPT-5 Series Guide: Model Selection, Grounding Improvements, and Confidence Enhancements.
CU 2.0 preview: New building blocks for AI apps and agents
In addition to the enhancements available in the refreshed GA API, we are introducing the new CU 2.0 Preview (2026-06-01-preview).
CU 2.0 preview advances two developer priorities:
Improving the quality and efficiency of content pipelines, and;
Enabling new low-latency and reasoning-intensive scenarios.
The capabilities below build on the same GPT-5 series foundation as the refreshed GA API.
GPT-5.5 and lower models supported in preview
The CU 2.0 public preview also supports the GPT-5 model series. This means customers can test new preview capabilities using the same model series direction as the refreshed GA path.
This is important for migration planning. Customers can evaluate accuracy, latency, and cost on representative data before deciding which model deployment and API version best fit their production needs.
Improving quality and efficiency for production workloads
The first goal of the new CU 2.0 preview version is to improve the work teams already run in production: raising extraction quality, lowering token cost, and making retrieval and routing more dependable. Five preview capabilities move in that direction, starting with Advanced Contextualization, which underpins several of the new features.
Advanced Contextualization
Advanced Contextualization helps custom analyzers use labeled examples and document knowledge more efficiently. In internal evaluations, it improved average accuracy by up to 3.5 percent while reducing average LLM token usage by up to 22 percent. Training data remains in the customer’s Azure Storage account and is used as a knowledge source rather than copied into the analyzer, preserving customer-controlled storage. The result is higher-quality structured extraction with fewer tokens and less data-management overhead.
For more details, see analyzer improvements for more information on training data management.
The same approach is used in five new prebuilt analyzers. For these prebuilt analyzers, Advanced Contextualization reduces LLM token consumption by up to 99 percent. That makes the new prebuilt analyzers more practical for high-volume scenarios where both quality and cost matter.
The takeaway is simple: Advanced Contextualization is not just a quality feature. It is a production efficiency feature. It helps customers get better structured extraction while using fewer LLM tokens and keeping training inputs under their own storage governance.
New prebuilt analyzers for tax and other document types
This preview introduces new prebuilt tax analyzers powered by Advanced Contextualization, extending support beyond individual tax forms to enterprise and state-level tax workflows: 1065, 1120-S, 8865, 1041 Schedule K-1 and Minnesota State M1.
The new analyzers support complex, multi-page layouts and use advanced contextualization to achieve higher extraction quality, lower latency, and competitive pricing, dramatically reducing LLM token consumption, with some extraction scenarios requiring no LLM tokens at all. These prebuilt analyzers are also available in the Content Understanding Studio and Microsoft Foundry, allowing developers to visually verify values, review bounding boxes for in-document grounding, and inspect confidence scores before deploying to production. For the complete list of available prebuilt analyzers and guidance on how to use and customize them, see prebuilt analyzers documentation.
imageThe Schedule K-1 prebuilt analyzer (prebuilt-tax.us.1065ScheduleK1) converts complex tax forms into structured JSON for downstream processing.
Semantic chunking in prebuilt-documentSearch
Retrieval quality often depends on how content is chunked. Fixed-size chunking can split a table from its heading, separate related paragraphs, or break a multi-page structure into fragments that are hard for retrieval systems to use.
The CU 2.0 public preview adds semantic chunking in prebuilt-documentSearch. Instead of splitting only by character count or page boundary, semantic chunking uses document structure to create more meaningful retrieval units. Semantic chunking preserves context and semantic relationships across sentences and paragraphs, and it is especially useful for RAG and agentic retrieval scenarios where the quality of what gets retrieved directly impacts the quality of reasoning outcome.
Classification: in-page splitting and confidence
Real enterprise submissions do not always align cleanly with page boundaries. A loan package, tax submission, case file, or scanned packet may contain multiple logical documents, and a single physical page can include the end of one section and the beginning of another.
The CU 2.0 public preview improves classification with in-page splitting. This allows classification to identify document segments at finer granularity than whole pages.
The preview also adds confidence for splitting and classification. Applications can use those signals to decide when to route a segment automatically, when to invoke a downstream analyzer, and when to send a result for human review.
This moves classification from a best-effort routing step toward a more operationally useful control point in document automation pipelines. For details, see classification enhancements.
Signature detection and metadata extraction in Layout
Classification decides what a document is; the Layout analyzer captures more of what is inside it. It now adds signature detection and document metadata extraction.
Signature detection helps identify signature regions and their locations in documents such as contracts, forms, invoices, and signed submissions. Metadata extraction surfaces available document properties such as author, title, creation date, content type, and language.
Combined with Layout analyzer’s existing ability to extract document structure and visual elements, including sections, headings, formatting, tables, figures, hyperlinks, annotations, and other layout elements, these new capabilities provide a more complete representation of document content and context from a single analyzer, making it easier to build intelligent document processing and agentic workflows.
To learn more, see the Layout analyzer documentation.
Try out signature detection and metadata extraction in the Content Understanding Studio.
imageSignature detection in Azure Content Understanding Studio, showing detected signature regions alongside the source document.
Expanding what Content Understanding can solve
The second goal of this preview API is to reach new scenarios beyond the standard extraction. Two capabilities open that door: synchronous processing and agentic reasoning for the hardest extractions.
Read and Layout APIs
Document workflows such as grounding an AI agent during a customer interaction, validating an identity document, or triggering a workflow when a file is submitted require an immediate response. CU 2.0 preview adds synchronous operations for the Read (prebuilt-read) and Layout (prebuilt-layout) analyzers to support these low-latency scenarios. The operations return structured results directly in the response. Documents can be submitted as binary content or by URL and are processed without temporary service-side storage. To learn more, see Azure Content Understanding announces Synchronous Operations.
Agentic mode for complex field extraction
If synchronous APIs are about responding faster, agentic mode is about reasoning harder. Some document extraction tasks require more than a single pass over the content. The answer may depend on evidence spread across a long document, values may need to be compared or validated, or a field may require reasoning over intermediate results before producing a final output.
For these scenarios, the CU 2.0 preview introduces agentic mode.
Agentic mode applies an iterative extraction workflow for complex document understanding. It is designed for harder extraction scenarios where standard extraction may not be sufficient, such as long legal agreements, financial filings, insurance records, or other documents where the relevant evidence is distributed across multiple sections.
Agentic mode works with an analyzer schema and uses additional reasoning to identify relevant content, extract values, evaluate intermediate results, and refine the final output. Because it performs additional reasoning, it can increase latency and token consumption compared with standard extraction. Customers should evaluate agentic mode on representative documents and use it when the expected quality gain justifies the additional cost and processing time. Learn more about agentic mode.
How to get started
The two updates are designed to be used together: one for production today, and one for evaluating what comes next. Choose the path that matches your workload:
Use the refreshed CU 1.0 GA API to improve existing workloads, including GPT-5 series support, grounding efficiency improvements, and the refreshed confidence model.
Use the CU 2.0 public preview to evaluate the next generation of CU capabilities, including synchronous Read and Layout APIs, Advanced Contextualization, semantic chunking, new prebuilt analyzers, improved classification, and agentic mode.
To explore prebuilt analyzers, custom analyzers, and structured outputs, start in Content Understanding Studio or Microsoft Foundry.
Quickstart Content Understanding Studio or Foundry | Microsoft Learn
Quickstart: Content Understanding SDK and Rest| Microsoft Learn
The post From Sync APIs to support for the GPT-5 model series and agentic workflows: What’s new in Azure Content Understanding – August 2026 appeared first on Microsoft Foundry Blog.
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み