Fireworks、MongoDB傘下Voyage AIと連携しAI推論スタックを統合
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Fireworks AI Blog
Fireworks AI と MongoDB が運営する Voyage AI の連携により、埋め込み・検索・再ランク付け・生成を単一プラットフォームで完結させる統合インフラが実現し、データ特化型AIの構築コストと複雑さを削減する。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 10:46
AI深層分析
キーポイント
単一プラットフォームでの完全統合
Voyage AI の全ラインナップ(Voyage 4 ファミリー、voyage-multimodal-3.5、rerank-2.5)が Fireworks でネイティブに実行可能となり、検索から生成までのパイプラインを単一APIとレイテンシドメインで完結させる。
データ特化型AIの構築基盤
一般モデルの限界を超え、自社のデータを基盤とした後学習(post-training)と高品質な検索機能(retrieval)を組み合わせることで、競合他社との差を広げる「専門知能」の実現を目指す。
アーキテクチャの複雑さ解消
従来の別ベンダーへの依存による二重請求や追加ネットワークホップ、セキュリティ面のリスクを排除し、開発チームが単一プラットフォームで運用コストと管理負荷を削減できる。
セキュリティとコンプライアンスの強化
Voyage AI を Fireworks に統合することで、データが通過する境界と監査対象を減らし、セキュリティリスクを低減できる。これにより、独自ドキュメントやクエリを fewer boundaries で管理しつつ、最先端の検索品質を維持できる。
ワークロードに応じた柔軟なモデル選択
精度、速度、コスト、ローカル開発、マルチモーダルデータなど、各ユースケースに最適な Voyage AI モデルシリーズを選択してチューニング可能である。既存の埋め込み・reranking モデルと併用し、自社データでベンチマークを行える。
重要な引用
Your entire retrieval-to-response pipeline (embed, retrieve, rerank, generate) now runs on one platform, one API, one latency domain.
Retrieval quality, not model size, is what limits AI built on your data
A stronger generalist does not rescue a weak retrieval layer.
Bringing the Voyage AI lineup to Fireworks means you consolidate retrieval and generation on one platform, keeping proprietary data inside fewer boundaries and under one review, without giving up frontier retrieval quality.
編集コメントを表示
編集コメント
検索と生成の分離という従来の課題に対し、両者を単一ドメインで統合するアプローチは実運用における遅延削減に寄与する。ただし、後学習による性能向上効果については、各社の具体的なユースケースにおける検証結果が今後の評価基準となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Fireworks は、MongoDB が運営する Voyage AI と提携した、唯一の専用推論プラットフォームです。Voyage の全ラインナップ(Voyage 4 ファミリー、voyage-multimodal-3.5、rerank-2.5)が Fireworks でネイティブに動作します。これにより、埋め込み・検索・再ランク付け・生成に至るまでの一連の処理を、単一のプラットフォームと API、そして同じレイテンシドメインで完結させることができます。さらに、Fireworks では独自のデータでポストトレーニング可能なオープンモデルを幅広く選択でき、そのモデルを完全に支配下に置くことも可能です。
自社データを基盤に AI を構築するチームにとって、この組み合わせは同時に 2 つの課題を解決します。
- ベストインクラスの検索機能と、単一プラットフォームによる簡素化の間のギャップ
- 汎用的な知能を借り受けることと、自社のビジネスに特化した知能を所有することの間のギャップ
AI の性能はモデルのサイズではなく、データに基づく検索の質で決まる
最先端の研究機関や企業は、より大きく、より汎用性の高いモデルの開発競争を繰り広げています。しかし、自社データを基盤に AI を構築する場合、モデルが推論できるのは検索によって提示された情報に限られます。強力な一般化モデルであっても、検索層が脆弱であれば性能は向上しません。精度は埋め込みと再ランク付けの段階で決まり、これが「会議室でのデモでは驚かれる」レベルか、「本番環境でも生き残る」システムかの分岐点となります。
これは、Fireworks が構築する大きなアイデアの半分です。その核心は「専門特化型知能」にあります。
汎用モデルには容量に限界があり、その多くが「決して実行しないタスク」で適切に動作するために消費されてしまいます。専門特化型知能を実現するには、オープンなベースモデルを自社のデータとユースケースに合わせてポストトレーニング(後方学習)します。これにより、モデルの能力を実際に必要な業務に集中させることが可能になります。
その結果、特定の領域内ではクローズド型の汎用モデルに匹敵し、あるいは凌駕する、より高速でコスト効率の高いモデルが生まれます。
このモデルとその記憶装置について:
- 半分は「モデル」です。ポストトレーニングして自社で所有するオープンなベースモデルです。
- もう半分は、「その知能を貴社独自のデータに根付かせる」ことです。これが検索(Retrieval)の役割です。
Voyage AI は、この検索レイヤーの最前線です。これら 2 つを一つのプラットフォーム上で統合すれば、ループが完成します。検索によってシステムは今日の貴社のデータに基づいて動作し、ポストトレーニングによってモデルは時間とともに洗練されます。そして、プロダクトの利用サイクル一つひとつがデータとなり、競合他社との差を広げることになります。
トレーニング、検索、配信を一つのプラットフォームで
これまで、チームは以下のトレードオフに直面していました。
- 検索機能を専門のベンダーに任せ、配信プロバイダーと手動で連携させる
- 一つのプラットフォームに統一し、そのプラットフォームが提供する検索機能に頼る
最初の選択肢には、2 つの請求書と 2 つのレイテンシプロファイル、そしてすべての呼び出しで追加のネットワークホップが発生するという現実的なコストがかかります。さらに、セキュリティやコンプライアンスのリスクも広がります。各ベンダーが増えるほど、クエリや機密文書が通過する場所が増え、信頼境界が拡大し、監査すべきデータ処理契約も増えるからです。
一方、2 つ目の選択肢はオーバーヘッドを排除しますが、検索精度には上限が生じます。
Voyage AI のラインナップを Fireworks に統合すれば、検索と生成を 1 つのプラットフォーム上で完結させられます。機密データを通過する境界を減らし、レビューも一元化できますが、最先端レベルの検索品質は維持されます。
Fireworks で Voyage を利用すると、ワークロードごとに検索を最適化できます:
- voyage-4-large:精度が最優先される場合
- voyage-4:精度と速度のバランス
- voyage-4-lite:レイテンシとコストに最適化
- voyage-4-nano:ローカル開発に適したモデル
- voyage-multimodal-3.5:テキストと視覚データが混在するコーパスの場合
- rerank-2.5:検索結果の精度向上用
これらは、Fireworks 上に既に存在する 既存の埋め込み・再ランクモデル と並列して利用可能です。自社のデータでベンチマークを行い、ユースケースに最適なモデルを選定できます。
以下の棒グラフは、Voyage 4 シリーズのモデルと Gemini Embedding 001、Cohere Embed v4、OpenAI v3 Large の平均検索品質を比較したものです。全体として「voyage-4-large」が最も高い性能を示し、他のモデル(voyage-4、voyage-4-lite、Gemini Embedding 001、Cohere Embed v4、OpenAI v3 Large)を上回る結果となりました。平均でそれぞれ 1.87%、4.80%、3.87%、8.20%、14.05% の差をつけています。
詳細は Voyage 4 シリーズの発表 をご覧ください。

Voyage を活用する
顧客対応やドキュメント作成エージェント。
これらの回答は直接顧客に届くため、精度が最も重要です。最適な検索機能として、埋め込みには「voyage-4-large」を、再ランク付けには「rerank-2.5」を使用し、LLM には Fireworks を活用してください。
社内知識アシスタント。
リスクは低く処理量が多いため、コスト最適化が優先されます。この場合は Voyage 4 Lite に「rerank-2.5」を組み合わせるのがバランスの良い選択です。
エージェントシステム内での根拠ある検索。 長い時間軸を持つエージェントはループ中にコンテキストを取得する際、多様で指示に富んだクエリを発行します。これはまさに rerank-2.5 が指向している「指示の遵守」そのものです。Voyage AI の検索機能をエージェントが呼び出すツールとして活用し、推論モデルを Fireworks で動かすことで、複雑なシステム全体を単一の場所で完結させられます。各ステップで他社ベンダーへ移動する必要はありません。
テキストや従来の RAG を超えて。 voyage-multimodal-3.5 は、図表、スクリーンショット、画像、動画といったコンテキストも検索対象に含めることで、テキストのみでは対応が難しいコーパスの活用を可能にします。ビジョン機能を備えた LLM と組み合わせることで、以前は不可能だったユースケースを実現するマルチモーダル RAG が誕生し、新たなアプリケーションのカテゴリーが開拓されます。
また、すべてのユースケースで生成が必要とは限りません。大規模なセマンティック検索、レコメンデーション、重複排除などは、埋め込みベクトルと再ランク付けだけで十分です。低次元のベクトルを使用することで、数百万アイテム規模でもストレージと検索コストを効率的に保つことができます。
すぐに始めよう
このパイプラインは、1 つの統合で完了します。Voyage AI の埋め込みモデルを設定し、rerank-2.5 を追加し、生成モデルを接続するだけで、検索とレスポンスが単一の API の背後で稼働します。製品が信号を生成する過程で、事後学習や独自化も可能です。
- • 構築: Voyage と Fireworks のクックブック を参照して進めてください。
- • Fireworks が初めての方へ: アカウントを作成 すれば、数分でフルパイプラインを稼働させることができます。
原文を表示
Fireworks is the first and only dedicated inference platform Voyage AI by MongoDB has partnered with. The full Voyage lineup now runs natively on Fireworks: the Voyage 4 family, voyage-multimodal-3.5, and rerank-2.5. Your entire retrieval-to-response pipeline (embed, retrieve, rerank, generate) now runs on one platform, one API, one latency domain. And it runs next to the broadest choice of open models you can post-train and own.
For teams building AI on their own data, that combination closes two gaps at once:
- The gap between best-in-class retrieval and single-platform simplicity
- The gap between renting generic intelligence and owning intelligence specialized to your business.
Retrieval quality, not model size, is what limits AI built on your data
The frontier-lab race optimizes for a bigger, more general model. But when you build on your own data, the model can only reason over what retrieval puts in front of it. A stronger generalist does not rescue a weak retrieval layer. Accuracy is won or lost at the embedding and reranking stage, and that is usually the line between a demo that impresses in a meeting and a system that survives production.
This is one half of a larger idea we build Fireworks around: specialized intelligence. A general model has finite capacity, most of it spent being adequate at tasks you will never run. Unlocking specialized intelligence, you post-train an open base model on your data and your use case, concentrating model capabilities on the work you actually do. The result is a faster and more cost-effective model that can match or beat a closed generalist model inside your lane.
The model and its memory:
- •One half is the model, an open base you post-train and own.
- •The other is grounding that intelligence in data only you have, which is the job of retrieval.
Voyage AI is the frontier of that retrieval layer. Bring the two together on one platform and a loop closes: retrieval grounds the system in your data today, post-training sharpens the model over time, and each cycle of product usage becomes data that widens the distance between you and your competitors.
One platform for training, retrieval, and serving
Until now, teams faced a tradeoff:
- Route retrieval to a separate specialist vendor and stitch it to your serving provider
- Consolidate on one platform and work with whatever retrieval it happened to offer.
The first path carries real costs: two bills, two latency profiles, and an extra network hop on every call. It also widens your security and compliance surface, since each additional vendor is another place your queries and proprietary documents travel, another trust boundary, and another data-processing agreement to audit. The second path removes that overhead but caps your retrieval quality.
Bringing the Voyage AI lineup to Fireworks means you consolidate retrieval and generation on one platform, keeping proprietary data inside fewer boundaries and under one review, without giving up frontier retrieval quality.
With Voyage on Fireworks you tune retrieval for every workload:
- •voyage-4-large where accuracy matters most
- •voyage-4 balancing accuracy with speed
- •voyage-4-lite optimized for latency and cost
- •voyage-4-nano ideal for local development
- •voyage-multimodal-3.5 when the corpus is interleaved text and visual data
- •rerank-2.5 for refining retrieval results
These sit alongside existing embedding and reranking models already on Fireworks, so you can benchmark against your own data and choose the right model for each use case.
The bar chart below compares the average retrieval quality of the Voyage 4 series of models along with Gemini Embedding 001, Cohere Embed v4, and OpenAI v3 Large. Overall, voyage-4-large is the top-performing model, surpassing voyage-4, voyage-4-lite, Gemini Embedding 001, Cohere Embed v4, and OpenAI v3 Large by an average of 1.87%, 4.80%, 3.87%, 8.20%, and 14.05%, respectively. More details are available from the Voyage 4 series announcement.

Put Voyage to work
Customer-facing support and documentation agent. These answers go straight to customers, so accuracy matters most. Use your most capable retrieval: voyage-4-large for embeddings, rerank-2.5 for precision, and an LLM on Fireworks.
Internal knowledge assistant. The stakes are lower and the volume is higher, so optimize for cost: Voyage 4 Lite with rerank-2.5 is the better balance.
Grounded retrieval inside agentic systems. Long-horizon agents that fetch context mid-loop issue varied, instruction-laden queries, which is exactly what rerank-2.5's instruction following targets. Voyage AI retrieval becomes a tool the agent calls, with the reasoning model served on Fireworks, so the full compound system runs in one place without a per-step hop to another vendor.
Beyond text and beyond RAG. voyage-multimodal-3.5 extends retrieval to context with diagrams, screenshots, images, and video, opening up corpora that text-only search handles poorly. Paired with vision-capable LLMs, multimodal RAG enables use cases that were previously impossible, unlocking new classes of application.
And not every use case needs generation: large-scale semantic search, recommendation, and deduplication run on embeddings and reranking alone, where lower-dimensional vectors keep storage and search cost-effective at millions of items.
Get started
The whole pipeline is one integration away. Wire up a Voyage AI embedding model, add rerank-2.5, and connect a generation model, and you have retrieval and response running behind a single API, ready to post-train and own as your product generates signal.
- •Build it: work through the Voyage and Fireworks cookbook.
- •New to Fireworks? Create an account and you can have the full pipeline running in minutes.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み