AWS、データ所在地で構築するエージェント型 AI のベクトルソリューションを発表
本文の状態
日本語全文を表示中
詳細モードで約20分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
AWS Machine Learning Blog
AWS は Agentic AI の基盤となるベクトル検索技術の重要性を解説し、既存データソースを活用したインテリジェントな検索・取得機能を提供する 6 つの専用ソリューションを発表した。
AI深層分析を開く2026年8月21日 09:58
AI深層分析
キーポイント
Agentic AI とベクトルの役割
エージェントが計画や推論を行う際、リアルタイムで正確な文脈にアクセスするためにベクトル検索が不可欠であり、これが AI の精度と信頼性を支える。
データ移動なしでの統合
AWS はデータベースやオブジェクトストアなど既存の場所にデータを保持したまま、PDF や動画などの非構造化データからも検索・取得を可能にするソリューションを提供する。
6 つの専用ソリューション
既存のデータストアが適用できない新規ワークロード向けに、用途に応じた 6 つの目的別ベクトルソリューションを選択できる判断モデルを示した。
主要ユースケースの多様化
RAG、セマンティック検索、ハイブリッド検索、GraphRAG、知識グラフなど、組織のナレッジを統合しハルシネーションを減らすための具体的な活用事例を提示した。
既存データストアへのベクトル機能追加
新しいデータベースを導入せず、既存のAWSデータストアにベクトル検索機能を追加するアプローチが推奨される。これによりデータ移行やクロスサービス間の移動が不要となり、コスト削減とパフォーマンス向上が実現できる。
重要な引用
Agentic AI is changing how you work, and vector search powers the retrieval layer that makes agents accurate, contextual, and grounded in real data.
AWS vector solutions bring intelligent search and retrieval to your data where it already lives, helping agents find and use the right context without requiring you to move or duplicate your data.
Vectors are the language of AI. They bridge frontier models and the scattered organizational knowledge accumulated over decades.
The right vector search solution follows the data, not the other way around.
編集コメントを表示
編集コメント
このブログ記事は、単なる技術紹介にとどまらず、企業が抱えるデータ分散の課題を解決し、実用的な Agentic AI を実現するための具体的なアーキテクチャ指針を示している。特に「データを移動しない」という制約下でのベクトル検索の実現性は、多くの組織にとって即戦力となる重要な知見である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
エージェント型 AI は働き方を変えつつあり、ベクトル検索はその中核となる情報取得層を支えています。これにより、AI エージェントは正確で文脈に即した、実データに基づいた判断が可能になります。エージェントは複数のステップを跨ぐワークフローにおいて計画・推論・実行を行いますが、組織の知識への高速かつ関連性の高いアクセスが不可欠です。
その知識はすでに、データベースやオブジェクトストレージ、検索エンジン、PDF や録画されたビデオ通話、チームが毎日使用するシステムといった多様な場所に存在しています。AWS のベクトルソリューションは、データが存在する場所でインテリジェントな検索と取得を実現し、データを移動または複製することなく、エージェントが必要な文脈を見つけ活用できるように支援します。
既存のデータストアを利用できない新規ワークロードについては、6 つの専用ソリューションに基づく明確な意思決定モデルを提供しています。これにより、エージェント型 AI や分析ワークロードに適したベクトルソリューションを適切に選択できます。
なぜベクトルが重要なのか:エージェント型 AI の主要ユースケース
ベクトルは AI の言語です。最先端のモデルと、数十年にわたって蓄積された散在する組織知識をつなぐ架け橋となります。データを高次元ベクトルとして表現することで、アプリケーションは意味を理解し、テキスト・画像・音声・動画間の関係を特定し、セッションを跨いで文脈を維持できます。製品説明やセキュリティログ、メディアライブラリなどどのようなデータであっても、ベクトル化によって共通の数学的空間に変換されるため、異なるモダリティ間での比較や検索が可能になります。
最先端の AI モデルと組み合わせることで、ベクトルはインテリジェントで文脈を理解し、パーソナライズされたユーザー固有の体験を実現する基盤となります。
- 検索拡張生成 (RAG) とナレッジベース: 実行時に信頼できるデータを取得することで最先端モデルの回答を裏付け、精度向上やハルシネーション(幻覚)の低減を図り、組織の知見に合致した回答を生成します。
- セマンティック検索: 正確なキーワード一致ではなく意味や意図に基づいて情報を検索するため、用語が異なっていても関連するコンテンツを発見できます。
- ハイブリッド検索: 語彙ベースの検索とセマンティック検索を組み合わせることで、構造化データおよび非構造化データの両方に対して包括的な結果を提供します。
- GraphRAG: セマンティック検索と知識グラフを統合し、多段階の推論が必要なエンタープライズ向けシナリオにおいて、正確で文脈に富み、追跡可能な回答を実現します。
- 知識グラフ: 人、製品、ドキュメント、概念などのエンティティを明示的な関係性で結びつけ、よりインテリジェントな検索や発見、AI を活用した推論をサポートします。
前節で概説した選択肢により、ベクトルは以下のようなユースケースを支えることができます。
- リアルタイム推薦システム: 小売、メディア、エンターテインメント分野において、ユーザーの興味に合致する製品やコンテンツ、体験をベクトルの類似性を通じて特定し、パーソナライズを実現します。
異常検知と不正検出では、高次元データ内の不規則なパターンを特定し、脅威やサイバーセキュリティ上の問題、設備の故障などを早期に発見する支援を行います。
マルチモーダルコンテンツ検索では、ファイルタイプやメタデータではなく意味に基づいた単一のクエリを用いて、テキスト、画像、音声、動画全体を検索できます。
こうしたユースケースの多くには、新しいデータベースを構築する必要はありません。AWS では、既存のサービスとデータベースにベクトル機能を追加できるため、データを移動させることなく、すでに利用している環境でベクトル処理が可能です。
以下の図は、AWS のベクトル機能の広範な範囲と、エージェント型 AI における主要ユースケースを示しています。

図 1: エージェント型 AI のユースケースにおける AWS ベクトル機能の広範な範囲
第一原則:データが既に存在する場所にベクトルを追加する
これが、私たちのアプローチ全体を導く基本原則です。すでに AWS データストアを持っている場合は、そこにベクトル検索機能を追加します。明確な理由がない限り、新しいサービスを導入する必要はありません。ベクトルはソースデータと共にあり続けるため、サービス間でのデータ移動が不要になり、ネイティブのクエリ機能とベクトル検索を組み合わせることができます。
既存のデータストアを活用すれば、新しいプログラミングツールや API、SDK などを習得するための学習コストをゼロにできます。また、すでに信頼されているデータストアが要件を満たしていることも安心材料になります。例えば、本番環境でスケーラビリティ、可用性、パフォーマンスが証明されたデータベースも、ベクトル検索機能を追加することで引き続き高い性能を発揮します。さらに、ベクトルとデータを同じ場所に保存すれば、アプリケーションの動作が高速化されます。データの同期や移動に関する心配も不要です。すでに投資した基盤の上に構築できるため、コスト削減にもつながります。
もしデータが既に Amazon OpenSearch Service、Amazon Simple Storage Service (Amazon S3)、Amazon Aurora PostgreSQL、Amazon DynamoDB、Amazon ElastiCache for Valkey、あるいは Amazon Neptune にあるなら、その場所にそのままベクトルを追加してください。最適なベクトル検索ソリューションは、データを動かすのではなく、データに追従するものです。
新しいワークロードに取り組む際は、レイテンシ、コスト、アクセスパターンのうち最も重要な要件を特定し、それに最適化されたエンジンを選びましょう。多くのワークロードでは、検索機能、スケーラビリティ、そしてアジェンティック AI の統合のバランスが求められます。そのようなケースには、Amazon OpenSearch Service をデフォルトとして推奨します。同サービスは、レキシカル検索、ベクトル検索、ハイブリッド検索、アジェンティック検索を単一のシステムで統合しており、大規模なデータでも高スループット、低レイテンシ、そして関連性の高い結果を提供できます。
以下の意思決定モデルは、ワークロードの要件に基づいて最適なベクトルソリューションを選定する際の参考になります。

図 2: ベクトルエンジン選択の意思決定モデル
Amazon OpenSearch Service:新規ワークロードのデフォルト
Amazon OpenSearch Service は、高スループットと低遅延を実現し、大規模なデータでも関連性の高い結果を返す管理型検索エンジンです。単一のシステム内で単語ベース検索、ベクトル検索、ハイブリッド検索を統合しており、単純な RAG(Retrieval-Augmented Generation)アプリケーションから高度なマルチシグナル検索まで幅広く対応しています。
機械学習(ML)を活用した自動最適化機能により、手動でのチューニングが不要となり、最適な設定が自動的に選択されます。GPU によるアクセラレーションを活用すれば、大規模データセットのインデックス作成速度を最大 10 倍に向上させながら、コストは従来の四分の一で抑えることが可能です。また、UltraWarm や Writable Warm のストレージ階層を採用することで、頻繁にはアクセスされないデータの保存コストを大幅に削減できます。
新しいワークロードには単一の支配的な要件がない場合がほとんどであるため、OpenSearch Service をデフォルトとして選択することをお勧めします。検索、スケーラビリティ、そしてアジェンティック AI への統合のバランスが必要とされるからです。OpenSearch Service は、レイテンシ、ベクトル量、1 秒間のクエリ数(QPS)、コスト効率、ハイブリッド検索、導入の容易さなど、あらゆる面で最も高い柔軟性を提供します。RAG(検索拡張生成)、異常検知、マルチモーダルコンテンツの発見、そしてハイブリッド検索を必要とするあらゆるワークロードを含む、広範なユースケースに対応しています。
数十億規模のベクトルボリュームをサポートし、1 秒間に数千件のクエリを処理可能で、月間 10 兆件以上のリクエストを処理する 10 万人以上のアクティブ顧客にサービスを提供しています。次世代の Amazon OpenSearch Serverless は、アジェンティック AI と動的なワークロードのために設計されています。前世代と比較して 20 倍高速で自動スケーリングし、数秒でプロビジョニングが完了します。また、アイドル時にはゼロまでスケールバックするため、ピーク容量用に常にリソースを確保する場合と比べて最大 60% のコスト削減を実現します。消費した容量に対してのみ課金されるため、エージェントが稼働していない場合は費用が発生しません。
Adobe は、Acrobat AI アシスタントを数億人のユーザーに提供するために OpenSearch Service を採用しました。これは Adobe のドキュメントエコシステムに直接統合された会話型生成 AI エンジンです。
Amazon S3 Vectors: 任意の規模でコスト最適化されたベクトルストレージ
Amazon S3 の機能である Amazon S3 Vectors は、ベクトルの保存とクエリをネイティブにサポートする初のクラウドオブジェクトストアです。S3 が持つコスト構造、スケーラビリティ、そしてシンプルさをベクトルストレージにもたらすことで、専用ベクトルデータベースと比較して、ベクトルのアップロード・保存・検索にかかるコストを最大 90% 削減します。これにより、Amazon S3 に保存されたコンテンツ全体で AI エージェントの記憶力や文脈理解、意味検索を強化する数十億規模のベクトルインデックスを構築・維持することが、管理すべきインフラゼロで実現可能になります。
S3 Vectors の一般提供開始以来、顧客は平均して 1 日あたり数千万回のクエリを実行しており、これはプレビュー期間中の 5 倍以上の増加です。この度、2 つの主要な機能強化により、クエリの体験と料金体系がさらに改善されました。
まず、S3 Vectors は 1 クエリあたり最大 10,000 件の検索結果をサポートするようになりました [1]。これは 100 倍もの増強であり、再ランク付けや集約、重複排除などを適用してより関連性の高い結果セットを生成する多段階の検索パイプラインにおいて特に価値があります。
また、ベクトルインデックスに 1,000 万本以上のベクトルが含まれる場合のクエリ料金が最大 80% 引き下げられました [2]。これにより、大規模な AI、RAG(検索拡張生成)、およびセマンティック検索ワークロードにおける類似度検索の実行コストが大幅に削減されます。
[1]: https://aws.amazon.com/about-aws/whats-new/2026/06/s3-vectors-supports-10000-search-results-per-query/
[2]: https://aws.amazon.com/about-aws/whats-new/2026/06/s3-vectors-reduces-query-charges-80-percent-large-indexes/
大規模なベクトルデータセットを頻繁に検索しない場合や、コスト効率の高いベクトルストレージが必要で、単純なベクトル検索とメタデータフィルタリングが求められる場合は、Amazon S3 Vectors を選択してください。レイテンシが約 100 ミリ秒以上許容できる場合や、QPS(1 秒間のクエリ数)は中程度だがベクトルデータセットの成長速度が速い場合に推奨されます。S3 Vectors はベクトルインデックスあたり最大 20 億個のベクトルをサポートし、クエリ課金モデルを採用しています。つまり、保存されたベクトルに対してのみストレージコストが発生し、検索を実行した際のみにクエリ費用が加算されます。
主なユースケースとしては、データレイクス内での意味的検索、RAG(Retrieval-Augmented Generation)に基づくナレッジ検索、大規模なベクトルストレージ、バッチ取得などが挙げられます。また、顧客は S3 API を介して直接ベクトルやインデックスを操作することも可能です。
BMW Group は、Amazon Bedrock AgentCore で構築されたインテリジェント検索エージェントを中核とするハイブリッド検索ソリューションの基盤として S3 Vectors を活用しています。エンジニアたちは、S3 Vectors を用いた意味的類似度検索と Amazon Athena による SQL クエリを組み合わせて、20 ペタバイトに及ぶデータを自然言語で直接クエリすることができます。
Amazon DynamoDB: スケールに関わらず単一桁ミリ秒のベクトル検索
Amazon DynamoDB は、サーバーレスで完全管理型の分散 NoSQL データベースであり、あらゆるスケールでミリ秒単位の低遅延を実現します。DynamoDB でのベクトル検索は、99% 以上の再現率を維持しながらミリ秒単位の応答性を提供し、数兆件のベクトル規模でも対応可能です。サーバーの用意やパッチ適用、管理が不要で、ソフトウェアのインストールや保守も不要な完全サーバーレスアーキテクチャです。
DynamoDB が提供するインフラ管理ゼロのメリットは、ベクトル検索機能においてもそのまま享受できます。バージョン管理の必要やメンテナンスウィンドウ、ダウンタイムを伴うメンテナンスもありません。DynamoDB のベクトル検索では、ベクトル埋め込み値を格納する属性に対して新しいインデックスを作成します。この機能は最大 4,096 次元まで対応し、ユークリッド距離、コサイン類似度、内積距離関数をサポートしています。また、インラインフィルタリングも利用可能です。
DynamoDB のベクトル検索は、グローバルテーブルとも連携して動作します。マルチリージョンでの最終整合性モデルおよび強整合性モデルの両方に対応しています。
DynamoDB は 100 万社以上の顧客に利用されています。すでに多くの企業で、会話中のセッションコンテキストの保持や、多段階タスクにおける状態追跡など、エージェントワークロードにも活用されています。
ベクトル検索機能を追加することで、DynamoDB はエージェントアプリケーションにおける長期記憶のための意味的検索もサポート可能になります。また、RAG(検索拡張生成)、マルチモーダルな類似度検索、レコメンデーションエンジンなどのユースケースにも対応しています。これらすべてを、サーバーレスのデータベース 1 つで完結できます。ベクトルストアを別途用意してデータを同期する必要はなく、習得すべき新しい API もありません。
運用データがすでに DynamoDB に存在する場合や、インフラ管理ゼロであらゆるスケールにおいて単一桁ミリ秒レベルのベクトル検索が必要であれば、DynamoDB のベクトル検索を選択してください。
Globant は、世界中の企業向けに AI 搭載製品やデジタルトランスフォーメーションソリューションを開発する、デジタルネイティブな企業です。
「私たちはすでに DynamoDB を基盤としたクライアントソリューションを構築しています。そのため、同じデータベース内でネイティブベクトル検索が利用可能になることは極めて価値が高いです。データを別々のベクトルストアに複製する必要も、第 2 のシステムを管理する必要もありません。完全にサーバーレスで、ほぼあらゆるスケールで自動的にスケーリングするため、各クライアントの利用分に対してのみ課金されます。また、ユーザー向け AI エクスペリエンスが求めるリアルタイムかつ低遅延の検索を実現します。これにより、エンジニアはインフラ管理から解放され、AI アプリの開発に集中できます。」
— Gastón Milano 氏(Globant エンタープライズ AI CEO)
Amazon ElastiCache for Valkey:意味的キャッシングのためのマイクロ秒レイテンシ
AWS のサーバーレス・フルマネージド型キャッシュサービス「Amazon ElastiCache」は、マイクロ秒レベルの低遅延性能を実現し、Valkey、Memcached、Redis OSS との完全な互換性を提供します。
ElastiCache for Valkey を利用すれば、大規模言語モデル(LLM)に対してセッションをまたぐ会話履歴を即座に提示するメモリ機構を実装することで、よりパーソナライズされた文脈に応じた回答を生成できます。キャッシュがデータベースコストの削減やアプリケーションパフォーマンスの向上に寄与するのと同様に、セマンティック・キャッシングは類似したプロンプトに対するキャッシュ済みレスポンスを提供することで、LLM の利用コストと遅延を低減します。また、ベクトル検索を活用して大規模データセット上で RAG(Retrieval-Augmented Generation)を実現すれば、回答の関連性を高め、実世界データを根拠として出力に反映させることでハルシネーション(幻覚)を抑制できます。
マイクロ秒レベルの遅延要件が求められるワークロードには、ElastiCache for Valkey を選択してください。具体的には、リアルタイム推薦エンジン、セッションベースのパーソナライゼーション、遅延がクリティカルな RAG パイプライン、セマンティック・キャッシングなどが該当します。同サービスは、マイクロ秒レベルのベクトル検索に対応する最大 10 億本のベクトルを処理可能です。
サンマ(Sanoma)社は、ElastiCache for Valkey を活用し、人間のモデレーターによる判断をベクトル化してリアルタイムで将来の AI モデレーションの判断に役立てています。再学習は不要です。現在、コメントの 30% が過去の判断と照合され、そのうち 6.5% はより正確な別の結果として処理されています。
Amazon Neptune: For GraphRAG
Amazon Neptune は、つながったデータを扱い、AI の精度を高めるためのサーバーレス型グラフデータベースサービスです。
ワークロードに高度につながったデータや、多段推論が必要となる場合、Amazon Neptune は単一のクエリ内でグラフのトラバーサルとベクトル類似性を組み合わせることで、独自にこの課題を解決します。
Neptune は低遅延なベクトル検索を提供し、20 億〜30 億件のベクトルを処理する容量を持ち、多段推論や追跡可能性も標準で備えています。特に後者の「追跡可能性」が重要です。規制の厳しい業界では、「何が返されたか」だけでなく、「なぜその結果が返されたのか」を示す必要があります。この透明性が、コンプライアンス、リスク管理、セキュリティ関連のワークロードにおいて Neptune を不可欠なものにしています。
グラフのトラバーサルとベクトル類似性を単一のクエリで組み合わせる必要がある場合、Neptune の選択が最適です。Neptune を使えば、「このアラートに影響を受けるサービスを担当しているチームはどれか?」といった接続関係をたどると同時に、その関係性の文脈をベクトル類似性と組み合わせて検索できます。多くの顧客が、金融やコンプライアンスのリスク管理、医薬品研究、創薬、セキュリティインテリジェンスなど、ナレッジグラフが必要なワークロードで Neptune を活用しています。
デロイトは、Amazon Neptune と AWS GraphRAG ツールキットを活用し、グラフベースの知識検索と生成 AI を融合させたセキュリティインテリジェンスセンターを構築しました。GraphRAG によってポリシー解釈、運用上の強制執行、リアルタイムメトリクスが連携することで、組織の最新状況を踏まえた予測的なセキュリティガイダンスを提供しています。
Amazon Aurora PostgreSQL: SQL ネイティブなベクトル検索
Amazon Aurora は、PostgreSQL の高パフォーマンスと可用性をグローバルスケールで実現します。pgvector 拡張機能を有効化し、最適化された読み取り処理を活用することで、高速なベクトル検索が可能になります。
AWS ブログ記事 で紹介されている Amazon Aurora PostgreSQL with pgvector 0.8.0 は、インデックス作成が 9 倍高速化され、フィルタリング後の検索結果の関連性が 100 倍向上しました。これは大きな飛躍です。ベクトル検索と SQL の全機能(結合、集計、WHERE 句、ACID トランザクション)を単一のエンジンで統合しています。
ソースデータが既に Aurora に存在する場合は Aurora を選択するのが最適です。SQL を活用したい場合や、ベクトル処理とリレーショナル処理を単一データベースに統合したい場合に推奨されます。低遅延を実現し、数百億個のベクトルをサポートするため、マルチテナント型の SaaS アプリケーション、エージェントのメモリストア、構造化データコンテキストを活用した RAG などに理想的です。
SaaS CRM の LeadSquared は、Amazon Bedrock と Amazon Aurora PostgreSQL を活用した生成 AI を用いることで、チャットボットの展開を加速させました。
原文を表示
Agentic AI is changing how you work, and vector search powers the retrieval layer that makes agents accurate, contextual, and grounded in real data. Agents plan, reason, and take action across multi-step workflows, making fast, relevant access to your organization’s knowledge essential.
That knowledge already has a home across databases, object stores, search engines, and unstructured sources such as PDFs, recorded video calls, and the systems your teams use every day. AWS vector solutions bring intelligent search and retrieval to your data where it already lives, helping agents find and use the right context without requiring you to move or duplicate your data.
For new workloads where no existing data store applies, we offer a clear decision model across six purpose-built solutions so you can choose the right vector solution for your agentic AI and analytics workloads.
Why vectors matter and top use cases for agentic AI
Vectors are the language of AI. They bridge frontier models and the scattered organizational knowledge accumulated over decades. By representing data as high-dimensional vectors, applications can understand semantic meaning, identify relationships across text, images, audio, and video, and maintain context across sessions. Whether you’re working with product descriptions, security logs, or media libraries, vectors convert everything into a shared mathematical space so you can compare and search across modalities.
Combined with frontier AI models, vectors are the foundation for intelligent, context-aware, personalized, and user-specific experiences:
- Retrieval Augmented Generation (RAG) and knowledge bases ground frontier model responses with trusted data retrieved at runtime, improving accuracy, reducing hallucinations, and generating responses aligned with organizational knowledge.
- Semantic search retrieves information based on meaning and intent rather than exact keyword matches, allowing users to discover relevant content even when different terminology is used.
- Hybrid search combines lexical search with semantic search to deliver comprehensive results across structured and unstructured data.
- GraphRAG combines semantic search with knowledge graphs to deliver accurate, context-rich, and traceable responses for enterprise scenarios requiring multi-step reasoning.
- Knowledge graphs connect entities, such as people, products, documents, and concepts through explicit relationships, supporting more intelligent search, discovery, and AI-powered reasoning.
With the options outlined in the previous section, vectors support several use cases, including:
- Real-time recommendation systems identify products, content, or experiences that align with user interests for personalization through vector similarity across retail, media, and entertainment.
- Anomaly and fraud detection identifies unusual patterns in high-dimensional data, supporting earlier detection of threats, cyber security issues, or equipment failures.
- Multimodal content discovery searches across text, images, audio, and video using a single query based on semantic meaning rather than file type or metadata.
The good news is you don’t need a new database for most of these use cases. With AWS, you get vector capabilities where your data already lives, across the services and databases that you already know and use, with no data migration required.
The following figure shows the breadth of AWS vector capabilities and top use cases for agentic AI.

Figure 1: Breadth of AWS vector capabilities for your agentic AI use cases
First principle: Add vectors where your data already lives
This is the principle that guides our entire approach. If you already have an AWS data store, add vector search to that data store. Don’t introduce a new service unless there’s a compelling reason to do so. Vectors stay with the source data, removing cross-service hops and combining vector search with native query capabilities.
When you use your existing data store, you remove the learning curve for new programming tools, APIs, SDKs, and more. You also can be confident that your existing data stores meet your requirements. As an example, your databases proven in production for scalability, availability, and performance will continue to deliver now with vector search. Finally, when your vectors and data are stored in the same place, your applications run faster. There’s no data sync or data movement to worry about. You also realize cost savings by building on investments that you’ve already made.
If your data is already in Amazon OpenSearch Service, Amazon Simple Storage Service (Amazon S3), Amazon Aurora PostgreSQL, Amazon DynamoDB, Amazon ElastiCache for Valkey, or Amazon Neptune, add vectors where the data already is. The right vector search solution follows the data, not the other way around.
For new workloads, identify your dominant requirement: latency, cost, or access pattern, and choose the engine optimized for it. Many workloads need a balance of search, scale, and agentic AI integration. For those, default to Amazon OpenSearch Service, which combines lexical, vector, hybrid, and agentic search in a single system with high throughput, low latency, and relevant results at scale.
The following decision model can help you select the right vector solution based on your workload requirements.

Figure 2: Decision model for vector engines
Amazon OpenSearch Service: The default for new workloads
Amazon OpenSearch Service is a managed retrieval engine that combines lexical, vector, and hybrid search in a single system with high throughput, low latency, and relevant results at scale. It supports multiple indexing strategies, vector quantization, and metadata filtering, scaling from simple RAG applications to advanced multi-signal retrieval. Machine learning (ML)-powered auto-optimization removes manual tuning by selecting the right configurations automatically. GPU acceleration indexes massive datasets up to 10x faster at a quarter of the cost, while UltraWarm and Writable Warm tiers reduce storage costs for less frequently accessed data.
Choose OpenSearch Service as the default for new workloads because most new workloads don’t have a single dominant requirement. They need a balance of search, scale, and agentic AI integration. OpenSearch Service provides the most flexibility across latency, vector volume, queries per second (QPS), cost efficiency, hybrid search, and ease of adoption. It covers the broadest set of use cases including RAG, anomaly detection, multimodal content discovery, and any workload requiring hybrid search. It supports multi-billion vector volumes, handles thousands of QPS, and serves more than 100,000 monthly active customers processing over 10 trillion requests per month.
The next generation of Amazon OpenSearch Serverless is built for agentic AI and dynamic workloads. It autoscales 20x faster than its previous generation, provisions in seconds, and ramps from zero to thousands of requests per second. It also scales back to zero when idle, delivering up to 60% cost savings compared to provisioning for peak capacity. You only pay for consumed capacity. If your agents aren’t running, you pay nothing.
Adobe adopted OpenSearch Service to scale its Acrobat AI Assistant to serve hundreds of millions of users. This is a conversational generative AI engine integrated directly into Adobe’s document ecosystem.
Amazon S3 Vectors: Cost-optimized vector storage at any scale
Amazon S3 Vectors, a capability of Amazon S3, is the first cloud object store with native support to store and query vectors. It brings the cost structure, scale, and simplicity of S3 to vector storage, reducing the cost of uploading, storing, and querying vectors by up to 90 percent compared to specialized vector databases. This makes it cost-effective to build and maintain billion-scale vector indexes that improve AI agent memory, context, and semantic search across content stored in Amazon S3 with zero infrastructure to manage.
Since the general availability of S3 Vectors, customers have been performing tens of millions of queries per day on average, a more than 5x increase over the preview period. Two recent enhancements improve the query experience and pricing. First, S3 Vectors now supports up to 10,000 search results per query, a 100x increase that’s especially valuable for multi-stage retrieval pipelines that apply reranking, aggregation, or deduplication to produce a more relevant result set. Second, query charges on vector indexes with over 10 million vectors are now reduced by up to 80%, significantly lowering costs for running similarity search across large-scale AI, RAG, and semantic search workloads.
Choose Amazon S3 Vectors when you need cost-effective vector storage with simple vector search and metadata filtering for infrequent queries of large vector datasets. It’s recommended when latency can be approximately 100 ms or more, or for fast-growing vector datasets at moderate QPS. It supports up to two billion vectors per vector index with pay-per-query pricing, so you pay for stored vectors while query costs accrue only when you search. Common use cases include semantic search over data lakes, RAG-based knowledge retrieval, large-scale vector storage, and batch retrieval. Customers also work directly with vectors and indexes through the S3 API.
BMW Group uses S3 Vectors as a building block for its hybrid search solution that is powered by an intelligent search agent built with Amazon Bedrock AgentCore. Engineers can query 20 petabytes of data in plain natural language, combining S3 Vectors for semantic similarity searches and Amazon Athena for SQL queries.
Amazon DynamoDB: Single-digit millisecond vector search at any scale
Amazon DynamoDB is a serverless, fully managed, distributed NoSQL database with single-digit millisecond performance at any scale. Vector search with DynamoDB delivers single-digit millisecond latency at over 99 percent recall designed for any scale, even trillions of vectors. It’s fully serverless with no servers to provision, patch, or manage, and no software to install, maintain, or operate. The zero infrastructure management that you love with DynamoDB, including no versions, no maintenance windows, and no downtime maintenance, is also available on vector search. Vector search in DynamoDB introduces a new index you create on the attribute that stores vector embeddings. It supports up to 4,096 dimensions, Euclidean, cosine, and dot-product distance functions, and inline filtering. DynamoDB vector search works with DynamoDB global tables, both with multi-Region eventual consistency and strong consistency.
DynamoDB serves over a million customers. Customers already use DynamoDB today for agentic workloads, such as holding session context during conversations and tracking state across multi-step tasks. With vector search, DynamoDB can support semantic retrieval for long-term memory in agentic applications. Other use cases include RAG, multimodal similarity search, and recommendation engines. And you can do all of this in one serverless database, no separate vector store to maintain and synchronize data, and no new API to learn.
Choose DynamoDB vector search when your operational data already lives in DynamoDB, or you need single-digit millisecond vector search at any scale with zero infrastructure management.
Globant is a digitally native company that builds AI-powered products and digital transformation solutions for enterprises around the world.
“We already build client solutions on DynamoDB, so having native vector search in the same database is extremely valuable — no need to replicate data into a separate vector store or manage a second system. It’s fully serverless, scaling automatically across virtually any scale, so we only pay for what each client uses, and it delivers the real-time, low-latency search that user-facing AI experiences demand. It lets our engineers focus on building AI apps instead of managing infrastructure.”
— Gastón Milano, CEO Enterprise AI, Globant
Amazon ElastiCache for Valkey: Microsecond latency for semantic caching
Amazon ElastiCache is a serverless, fully managed caching service delivering microsecond latency performance with full Valkey, Memcached, and Redis OSS compatibility.
With ElastiCache for Valkey, you can build more personalized, context-aware responses by implementing memory mechanisms that surface cross-session conversation history to large language models (LLMs). Similar to how a cache reduces database costs and improves application performance, semantic caching reduces the cost and latency of using LLMs by serving cached responses for semantically similar prompts. You can also use vector search to power RAG on large datasets to improve response relevance and reduce hallucinations by grounding outputs with real-world data.
Choose ElastiCache for Valkey when the workload has a microsecond latency requirement. This includes real-time recommendation engines, session-based personalization, latency-critical RAG pipelines, and semantic caching. It supports up to one billion vectors for microsecond latency vector search.
Sanoma uses ElastiCache for Valkey to turn human moderator decisions into vectors that inform future AI moderation calls in real time, with no retraining required. Today, 30 percent of comments are matched against past decisions, and 6.5 percent receive a different, more accurate outcome as a result.
Amazon Neptune: For GraphRAG
Amazon Neptune is a serverless graph database service for connected data and improved AI accuracy.
When your workload involves highly connected data, or multi-hop reasoning, Amazon Neptune uniquely solves this by combining graph traversal with vector similarity in a single query.
Neptune delivers low-latency vector search with 2–3 billion vector capacity, and built-in multi-hop reasoning and traceability. That last point is critical. Regulated industries must show why a result was returned, not just what was returned. This transparency is what makes Neptune essential for compliance, risk, and security workloads.
Choose Neptune when the workload needs to combine graph traversal with vector similarity in a single query. With Neptune, you can traverse connections (for example, “which teams own the services affected by this alert?”) and combine that relationship context with vector similarity in a single query. Many of our customers use Neptune for workloads that need a knowledge graph, such as financial and compliance risk, pharmaceutical research, drug discovery, and security intelligence.
Deloitte uses Amazon Neptune with the AWS GraphRAG Toolkit to power a Security Intelligence Center that combines graph-based knowledge retrieval with generative AI. By connecting policy interpretation, operational enforcement, and real-time metrics through GraphRAG, Deloitte delivers predictive security guidance grounded in timely organizational context.
Amazon Aurora PostgreSQL: SQL-native vector search
Amazon Aurora delivers high performance and availability at global scale for PostgreSQL. You can turn on the pgvector extension and use optimized reads for high-performance vector search.
Amazon Aurora PostgreSQL with pgvector 0.8.0 delivers 9x faster indexing and 100x more relevant filtered results. This is a major leap. It combines vector search with the full SQL query surface: joins, aggregations, WHERE clauses, and ACID transactions in a single engine.
Choose Aurora when your source data already lives in Aurora. It’s recommended when you want to use SQL, or when you need to consolidate vector and relational workloads into a single database. It delivers low latency and supports hundreds of billions of vectors, and is ideal for multi-tenant software as a service (SaaS) applications, agentic memory stores, and RAG with structured data context.
LeadSquared, a SaaS CRM
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み