Box、Gemini Embeddings 2 で多モーダルエンタープライズエージェントを推進
本文の状態
日本語全文を表示中
詳細モードで約10分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Google Cloud AI
Box はコンテンツ管理システムでテキスト検索や RAG に加え、Gemini Embeddings 2 を活用した多モーダルなエンタープライズエージェントの実現に向けた転換を進めている。
AI深層分析を開く2026年8月19日 01:53
AI深層分析
キーポイント
エンタープライズ管理システムの構造的転換
クラウド移行期以来最大のアーキテクチャシフトとして、従来のテキストベース検索からマルチモーダル要素を扱う方向へ進化している。
Gemini Embeddings 2 の統合と機能拡張
Box の Agentic Platform に Google Cloud の Gemini Multimodal Embeddings 2 を組み込み、表の列構造やフローチャートの論理を空間レイアウトのまま解釈可能にする。
視覚的・空間的幾何学の保持
複雑な文書要素を単なる文字列に変換するのではなく、人間が認識する方法で空間関係の整合性を維持したまま処理を行う能力を提供する。
視覚的要素の可視化
技術チャートやフロー図などの文書内の視覚的指標が検索システムで無視されることがなくなる。ユーザーは画像とテキストを同時にクエリできるようになる。
ハイブリッドファイル形式の統合
PDF、スプレッドシート、プレゼンテーションなど異なるフォーマット間でのクロス参照が可能となる。マルチモーダル埋め込みにより、これらの多様な形式にわたる統一された理解が実現する。
重要な引用
Enterprise content management is experiencing its biggest architectural shift since the cloud migration era.
Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.
Multimodal capabilities ensure that these elements are no longer invisible to search systems, allowing users to query images and text simultaneously.
Extending RAG with multimodal embeddings creates a unified understanding across these varied formats.
編集コメントを表示
編集コメント
この連携は、単なるテキスト検索の延長ではなく、文書の視覚的・構造的意味を AI が理解する新たな段階を示している。特に金融や医療分野のように表や図が重要な役割を果たす業界において、AI の実用範囲を大きく広げる可能性を秘めている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
エンタープライズ向けコンテンツ管理は、クラウド移行期以来最大のアーキテクチャ転換期を迎えています。
長年にわたり、企業は Box に財務モデルや臨床試験プロトコル、M&A のデューデリジェンスルーム、エンジニアリング図面、法務コンプライアンスの手引など、重要なデータを数兆ギガバイト規模で蓄積してきました。これまでテキストベースの検索と検索拡張生成(RAG)がこれらのリポジトリ内に眠る膨大なナレッジを解き放ち、エンタープライズ AI にとって強力かつ極めて効果的な基盤を築いてきました。
従来の RAG アーキテクチャはテキスト処理において確立された実績がありますが、エージェント時代にはさらに高度な機能が求められます。次なる自然な進化として、このフレームワークを拡張し、テキストと共存する「本質的に多様なモダリティを持ち、深く空間的であり、かつ高度に構造化された要素」を取り込む必要があります。テキスト埋め込みは文章のインデックス化には優れていますが、マルチモーダルアーキテクチャは新たな大きな可能性を開きます。具体的には、財務表の厳密な行と列の意味構造を保持し、臨床データのような視覚的証拠を解釈し、多ページにわたるフローチャートの論理関係を空間レイアウトを損なうことなくマッピングすることが可能になります。
膨大なデジタルコンテンツを処理し、次世代の機能を実現するために、Google Cloud と Box は「Gemini Multimodal Embeddings 2」Gemini Multimodal Embeddings 2 を活用した、Box の Agentic Platform に高度なマルチモーダル機能を統合しました。これにより、業界をリードするインテリジェント・コンテンツ管理プラットフォームである Box と、Google Cloud の先進的な AI 埋め込み技術が融合します。
埋め込みの向上による恩恵:ドキュメント内容の次元拡大
- 視覚的・空間的幾何学の維持: 多段表や財務行列といった複雑なドキュメント要素は、その配置によって意味を伝えます。これらを単なるテキストの羅列に変換すると、見出しと対応するデータポイントが切り離されてしまいます。マルチモーダル埋め込みを使用すれば、システムは人間と同じように文書を解釈でき、空間的な関係性の完全性を保つことができます。
- 視覚的モダリティの可視化: 企業ドキュメントには、技術チャート、工程フロー図、ブランディング素材、製品写真など、多くの視覚的指標が盛り込まれています。マルチモーダル機能により、これらの要素は検索システムにとって「見えない存在」ではなくなり、ユーザーは画像とテキストを同時にクエリできるようになります。
ハイブリッドなファイル形式の連携:現実のビジネスワークフローは単一のドキュメントフォーマットに収まることは稀です。エージェントは、PDF のポリシー文書、スプレッドシートで管理される追跡ログ、そしてプレゼンテーション資料を相互参照する必要がある場合があります。マルチモーダル埋め込み(embeddings)を活用して RAG(検索拡張生成)を拡張することで、これらの多様なフォーマットにまたがる統合的な理解が可能になります。
アーキテクチャの解決策:Gemini Multimodal Embeddings 2
Google Cloud の Gemini Multimodal Embeddings 2 は、テキスト、ラスタ画像、ドキュメントページ、レンダリングされたスプレッドシートの表、視覚的なチャートをすべて同一のセマンティック表現空間に埋め込むことのできる、統合型マルチモーダルベクトル空間を導入しました。

gemini-embeddings-2 で実現される主要な機能:
- クロスモーダル検索(テキストから視覚へ/視覚からテキストへ):自然言語によるクエリで、特定の図表やダイアグラムを正確に検索できます。膨大なスライドライブラリの中から目的の要素を見つける際にも、手動でのタグ付けは不要です。
- レイアウト認識型ドキュメント埋め込み:ファイルを任意のテキストブロックに分割するのではなく、システムはドキュメントページのレンダリング画像をそのまま埋め込むことで、視覚的な階層構造、吹き出しボックス、構造的な文脈を保持します。
異種フォーマット間の橋渡し
.docx、.xlsx、.pdf、.pptx、.png、.csv など、さまざまな形式のコンテンツをシームレスに連携させるネイティブサポート。モダリティ固有の構造化情報を失うことなく、データの変換が可能です。
マルチモーダルエンタープライズエージェントの 3 つのコアパターン
Box でマルチモーダル埋め込みを活用することで、従来の RAG(検索拡張生成)を拡張し、複雑なビジュアルワークフローに対応する方法を示す、3 つのユニークな設計パターンが明らかになりました。
パターン 1:複雑な財務・分析レポート
課題
企業の財務、リサーチ、監査チームは、重要なデータが埋め込まれた表や成長チャート、脚注注釈を含む、非常に構造化された文書を分析しています。テキストベースのインデックス化だけでは、これらの数値から文脈が切り離されてしまい、自動化された分析が困難になります。
マルチモーダルの優位性
- 構造的整合性の確保: 埋め込みモデルは表やチャートの物理的な構造を捉えるため、財務エージェントは「列見出しが特定の行の指標に適用される」といった関係を正しく理解できます。
- 視覚的トレンド分析: エージェントは記述された要約と、添付の棒グラフや折れ線グラフにおける視覚的な傾向を相互参照できます。これにより、記述された主張とソースデータとの間の不一致を特定し、指摘することが可能になります。
- 文脈に基づく情報源の提示: ユーザーは複雑なポートフォリオに対してクエリを実行するだけで、特定の指標を支える正確なページ、表、またはチャートを即座に取得できます。

パターン2:多機能的臨床意思決定支援と診断補助
課題
医療や臨床現場では、重要な患者データが非常に異なる非構造化の視覚的・テキスト形式に断片的に散在しています。具体的には、外部から撮影された物理写真(視覚的証拠)や顕微鏡下の病理スライド(検査報告書)、そして構造化されたリスクマトリクス(トリアージグリッド)などです。従来のテキストベースのシステムや孤立した分析ツールでは、これらの異種モダリティ間の関係を同時に統合することができず、結果として重要な診断が遅れたり、即座に生命を脅かす手技上の合併症を見逃したりするリスクがあります。
多機能的な優位性
- 異種モダリティによる臨床的統合: 臨床写真、組織病理画像、トリアージグリッドを単一の空間にインデックス化することで、身体的症状と細胞レベルの検査証拠を同時に評価します。
- 詳細な異常検出: 顕微鏡下のニッチな視覚パターン(寄生虫嚢壁など)を医学知識と結びつけることで、稀な疾患を迅速に特定・隔離できます。
- リスク認識型の意思決定支援: 発見された事象をトリアージフレームワークと照合し、生命を脅かすアナフィラキシーショックなどの即時的な患者リスクに対する即時警告を提供します。

パターン3:文書間多機能的統合とデータ整合性
課題
企業の情報は、バラバラのファイルや形式に散らばっています。例えば、PDF の議事録、Excel のチャート、PNG のフライヤー、メールのスレッドなどです。従来のツールはこれらのファイルを個別に分析するだけで、独立した文書間で詳細を確認したりデータの不整合を解決したりする際に、関連性を結びつけることができません。
マルチモーダルな優位性
- ファイル横断の統合: PDF、スプレッドシート、画像、メールなど全く異なる形式間の情報を同時に結びつけ、複雑なビジネス問い合わせに回答します。
- 矛盾の解決: 資産間の不整合を指摘し解決します。例えば、最新の財務スプレッドシートと照合することで、画像上の古い価格情報を検出できます。
- 視覚からテキストへの監査: 署名付き PDF 契約書と法務レビューメールなど、視覚的またはスキャンされたファイルとテキストベースの記録を照合し、条項の欠落や変更を検出します。

エンタープライズコンテンツ管理におけるエージェントの未来
Gemini Embeddings 2 を Box の Agentic Platform に統合したことは、次世代のコンテンツインテリジェンスを強化する重要な新機能です。マルチモーダルな埋め込み(embeddings)を活用することで、Box は単なる基本的な検索を超え、能動的で知的なコラボレーションへと進化します。
Box のインテリジェント・コンテンツ・マネジメントプラットフォームは、エンタープライズ AI インフラにおける根本的な転換点を示しています。受動的な文書保存から脱却し、AI エージェントが完全なコンプライアンスとセキュリティ制御のもとでコンテンツを照会し、相互参照し、行動を起こせるよう、統制された意味論的インデックスを持つ推論層を提供するものです。
検索、メタデータ抽出、調査、分析、構成にわたるネイティブ AI エージェント群とマルチモーダルな埋め込みによって支えられた Box は、組織がビジネスリスクとなる前に、期限切れの価格データや契約条項の有効期限、あるいは文書間での矛盾などを事前に検知し、洞察を提示することを可能にします。金融サービス、ライフサイエンス、法務運営といった高複雑度の業界において、テキスト、表、グラフ、画像を横断して推論を行う Box の能力は、マルチモーダルな理解力を競争上の必須要件へと押し上げています。
より広範なエンタープライズ AI エコシステムとの相互運用性を意識して設計された Box は、すべての AI 駆動型ワークフローが承認され監査可能な企業データに基づいていることを保証する、唯一の統制されたコンテンツ基盤として機能します。
企業データの風景は、もともとマルチモーダルでした。今や、その価値を最大限に引き出す技術が揃っています。
Gemini Embeddings 2 を統合することで、Box はユーザーが非構造化されたエンタープライズコンテンツから前例のない価値を引き出せるよう支援します。マルチモーダルファーストのアーキテクチャを採用し、厳格な精度ベンチマークと監査対応可能なグラウンディングを徹底する製品リーダーこそが、次なる企業生産性とイノベーションの波を牽引することになるでしょう。
本プロジェクトへの貢献に感謝いたします:Ken Ikeda 氏、Afshaan Mazagonwalla 氏、Samip Thakkar 氏。
原文を表示
Enterprise content management is experiencing its biggest architectural shift since the cloud migration era.
For years, enterprises have stored trillions of gigabytes of critical data in Box: financial models, clinical trial protocols, M&A due diligence rooms, engineering schematics, and legal compliance playbooks. Up to this point, text-based search and retrieval-augmented generation (RAG) have successfully unlocked the vast narrative knowledge within these repositories, establishing a powerful and highly effective baseline for enterprise AI intelligence.
Traditional RAG architectures have mastered text processing, but the agentic era demands more. The next logical evolution is to extend this framework to capture the inherently multimodal, deeply spatial, and highly structured elements that exist alongside text. While text embeddings excel at indexing prose, multimodal architectures unlock a major new capability: For example, they preserve the strict row-column semantics of financial tables, interpret visual evidence like clinical data, and map the logic of multi-page flowcharts without losing their spatial layout.
To deliver next-generation capabilities that can handle the vast universe of digital content, Google Cloud and Box are integrating advanced multimodal capabilities into Box's Agentic Platform, powered by Gemini Multimodal Embeddings 2 merging Box’s industry-leading Intelligent Content Management platform with Google Cloud’s advanced AI embeddings.
Benefits of improved embedding: Extending the dimensions of document content
- Preserving visual and spatial geometry: Complex document elements like multi-column tables or financial matrices rely on their spatial layout to convey meaning. Converting these elements into a flat string of text can disassociate column headers from their corresponding data points. Multimodal embeddings allow systems to interpret the document exactly as a human does, maintaining the integrity of spatial relationships.
- Illuminating the visual modality: Enterprise documents are filled with visual indicators: technical charts, process flowcharts, branding assets, and product photography. Multimodal capabilities ensure that these elements are no longer invisible to search systems, allowing users to query images and text simultaneously.
- Connecting hybrid file formats: Real-world business workflows rarely live in a single document format. An agent may need to cross-reference a PDF policy, a spreadsheet tracking log, and a presentation deck. Extending RAG with multimodal embeddings creates a unified understanding across these varied formats.
The Architectural Solution: Gemini Multimodal Embeddings 2
Google Cloud’s Gemini Multimodal Embeddings 2 introduces a unified, multimodal vector space capable of embedding text, raster images, document pages, rendered spreadsheet tables, and visual charts into the same semantic representation space.

Key product capabilities unlocked by gemini-embeddings-2:
- Crossmodal retrieval (text-to-visual / visual-to-text): Enables natural language queries to retrieve highly specific visual components, such as locating a target chart or diagram within a massive library of slides, without requiring manual tagging.
- Layout-aware document embedding: Rather than breaking files into arbitrary text blocks, the system can embed document page renderings directly, preserving visual hierarchies, callout boxes, and structural context.
- Heterogeneous format bridging: Native support for seamlessly bridging content across .docx, .xlsx, .pdf, .pptx, .png, and .csv without losing modality-specific structural information.
Three core patterns of multimodal enterprise agents
By leveraging multimodal embeddings within Box, we have identified three uniqueprimary design patterns that illustrate how organizations can extend traditional RAG to support complex, visual workflows.
Pattern 1: Complex financial & analytical reporting
The challenge
Corporate finance, research, and audit teams analyze highly structured documents where vital data resides in embedded tables, growth charts, and footnote annotations. Text-only indexing can separate these numbers from their context, making automated analysis challenging.
The multimodal advantage
- Structural alignment: The embedding model captures the physical structure of tables and charts, allowing financial agents to understand that a column header applies to a specific row of metrics.
- Visual trend analysis: Agents can cross-reference written summaries with visual trends in accompanying bar or line charts, identifying and pointing out discrepancies between written claims and source data.
- Contextual sourcing: Users can query complex portfolios and instantly retrieve the exact page, table, or chart supporting a specific metric.

Pattern 2: Multimodal clinical decision support & assisted diagnosis
The challenge
In healthcare and clinical environments, critical patient data is fragmented across vastly different, unstructured visual and textual formats — ranging from external physical photos (visual evidence) and microscopic pathology slides (lab reports) to structured risk matrices (triage grids). Traditional text-based systems or isolated analysis tools cannot synthesize these cross-modal relationships simultaneously, which can delay critical diagnoses or risk missing immediate, life-threatening procedural complications.
The multimodal advantage
- Cross-modal clinical synthesis: Evaluates physical symptoms alongside cellular-level laboratory evidence simultaneously by indexing clinical photos, histopathology imagery, and triage grids into a single space.
- Granular anomaly identification: Connects niche visual patterns under a microscope (like parasitic cyst walls) with medical knowledge to rapidly isolate rare conditions.
- Risk-aware decision support: Cross-references findings against triage frameworks to deliver instant warnings about immediate patient risks, such as life-threatening anaphylactic shock.

Pattern 3: Cross-document multimodal synthesis & data reconciliation
The challenge
Enterprise information is fragmented across disconnected files and formats (e.g., PDF minutes, Excel charts, PNG flyers, and email threads). Traditional tools analyze these files in isolation, failing to connect the dots when verifying details or resolving data contradictions across independent documents.
The multimodal advantage
- Cross-file synthesis: Connects information across entirely different formats (PDFs, spreadsheets, images, emails) simultaneously to answer complex business queries.
- Conflict resolution: Flags and resolves contradictions between assets, such as catching outdated pricing on an image by cross-checking it against the latest financial spreadsheets.
- Visual-to-text auditing: Audits visual or scanned files against text-based records (e.g., verifying a signed PDF contract against a legal review email) to catch missing clauses or changes.

The future of agentic enterprise content management
The integration of gemini-embeddings-2 into Box’s Agentic Platform is an important new capability to improve the next era of content intelligence. Multimodal embeddings help Box to move beyond basic search to active, intelligent collaboration.Box's Intelligent Content Management platform represents a fundamental shift in enterprise AI infrastructure — moving beyond passive document storage to deliver a governed, semantically indexed reasoning layer where AI agents can interrogate, cross-reference, and act on content with full compliance and security controls already in place.
Powered by multimodal embeddings and a suite of native AI agents spanning search, metadata extraction, research, analysis, and composition, Box enables organizations to proactively surface insights such as flagging stale pricing data, expiring contract clauses, or cross-document contradictions before they become business risks. For high-complexity industries like financial services, life sciences, and legal operations, Box's ability to reason across text, tables, charts, and images makes multimodal understanding a competitive requirement.
Designed to interoperate with the broader enterprise AI ecosystem, Box serves as the single governed content foundation that ensures every AI-driven workflow is grounded in authorized, auditable enterprise data.
When you think about it, the enterprise data landscape was always multimodal. Now we have the technology to make the most of it. By integrating gemini-embeddings-2, Box helps its users unlock unprecedented value from unstructured enterprise content. Product leaders who embrace multimodal-first architectures, rigorous precision benchmarking, and audit-ready grounding will lead the next wave of enterprise productivity and innovation.
The team would like to thank Ken Ikeda, Afshaan Mazagonwalla, and Samip Thakkar for their work on this project.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み