Pinecone、独自クラウド向け「Nexus」一般提供へ、モデル性能向上とコスト削減を報告
本文の状態
日本語全文を表示中
詳細モードで約18分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Pinecone
Pinecone はベクトルデータベース「Nexus」の一般提供を開始し、顧客企業のクラウド内に知識層を配置することで、大規模言語モデルの精度向上とコスト削減を実現する技術的進展を発表した。
AI深層分析を開く2026年8月6日 23:02
AI深層分析
キーポイント
Nexus の一般提供開始とデータ所在地の確保
Pinecone は「Nexus」を一般提供(GA)とし、AWS、Google Cloud、Azure 上で顧客クラウド内に実行されるデータプレーンとして運用可能になったことを発表した。
ベンチマークにおける性能とコストの劇的改善
Sierra の τ-Knowledge ベンチマークで GPT-5.5 と GPT-5.2 に Nexus を適用した結果、GPT-5.5 は精度を維持しつつコストを削減し、GPT-5.2 は精度が 12% 向上してコストが 80% 削減された。
顧客サポート業務における自動化率の向上
Pinecone の自社顧客サポートキューへの適用により、エージェントによる自己解決率が 25% から 55% に向上したことが報告されている。
ナレッジベースによるコスト削減効果
Nexus を導入することで、タスクあたりのモデル呼び出し回数が半分になり、1.45ドルだったコストが0.53ドルに低下した。これはモデルの推論能力ではなく、安価な知識へのアクセス不足がボトルネックとなっているためである。
エンタープライズエージェントの生産現場での課題
規制産業における実装では、タスク完了精度の停滞、トークンコストの高騰、高レイテンシ、根拠となる引用の欠如という4つの重大な欠陥が顕在化している。
重要な引用
The durable advantage an enterprise has is its knowledge and how its people do the work.
GPT-5.5 held its accuracy at 77% less cost per task.
Nexus keeps it where it belongs: in a governed layer, in your cloud, shaped by your own experts.
The agent has to find and apply the right policy before it acts, and a plausible answer grounded in the wrong version of a policy scores zero.
編集コメントを表示
編集コメント
Pinecone は「モデル性能の限界」ではなく「知識管理の壁」がエンタープライズ AI のボトルネックであると指摘し、Nexus を通じてその解決を図っている。GPT-5.2 や GPT-5.5 という具体的なバージョン名がベンチマークで言及されている点は、最新のモデル評価における実証データの重要性を物語っている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
5 週間前に「Pinecone Nexus」のパブリックプレビューを開始しました。本日、Nexus は一般提供(GA)となり、お客様のクラウド環境に直接デプロイできるようになりました。
シエラ社のオープンベンチマーク「τ-Knowledge」(エージェント型カスタマーサービス業務向け)において、GPT-5.5 と GPT-5.2 に Nexus の知識層を付与し、ベンチマークが提供するツールを用いて既存モデルと比較検証しました。その結果、GPT-5.5 はタスクあたりのコストを削減しながら精度 77% を維持し、GPT-5.2 は精度を 12% 向上させつつコストを 80% 削減することに成功しました。両モデルとも、ツール呼び出しとモデル呼び出しの回数を約半分に減らすことができました。
また、Nexus を自社のカスタマーサポートチケットキューにも適用しました。その結果、エージェントが単独で処理していたインバウンドチケットの解決率は 25% から 55% に向上しました。
本記事では、ベンチマーク結果が示す意味や、パブリックプレビューを通じて企業が知識をどう管理すべきかを学んだ教訓、そして一般提供によってお客様にどのような変化をもたらすかについて解説します。
今すぐ利用可能な機能
企業が持つ持続的な優位性は、自社の知識と、それを活用する人材の能力にあります。その知識をサードパーティのモデルへ転送するエージェント呼び出しは、優位性の一部を他者に譲渡することと同義です。Nexus は、この知識が本来あるべき場所——管理されたレイヤーとして、お客様のクラウド内に、自社の専門家が構築した形で——保持されます。この原則が、今日リリースされるすべての機能の根幹にあります。Nexus は、標準的な企業導入フロー(評価→パイロット導入→調達→本番環境での運用)に即座に対応可能です。
デプロイメント: Nexus のデータプレーンは、AWS、Google Cloud、Azure 上のお客様のクラウド内で実行されます。ドキュメントやコンパイル済みの知識がインフラストラクチャから外部へ流出することはありません。
モデルの選択。 Nexus は、あなたが選んだモデルで動作します。オープンウェイトモデルも利用可能です。モデルの認証情報はユーザーが提供し、推論呼び出しはあなたのクラウドから指定したプロバイダーへ直接行われます。各ワークフローではアンサンブル構成が可能で、ステップごとに最適なモデルを自動選択できます。
ロックインなし。 Nexus がコンパイルする知識レイヤーはすべてあなたのものです。アーカイブとしてダウンロードも可能です。
統一されたインターフェース。 エージェント、チャットボット、AI 検索、レコメンデーションシステムなど、すべてのアプリケーションが KnowQL を通じて同じ知識レイヤーを照会します。仕様の詳細は spec.knowql.org で確認できます。
プラットフォームとの親和性。 Nexus は Pinecone プラットフォームの一部として設計されており、検索基盤には Pinecone Database が、生産環境向けの知識アプリには Pinecone Marketplace が提供されています。
これらすべてを支えているのは一つの主張です。エンタープライズ向けエージェントは、モデルの性能限界に達するよりも先に、知識の壁にぶつかるのです。この仮説を検証するために、以下のテストを実施しました。
ベンチマーク:同じモデル、異なる知識
τ-Knowledge は、Sierra が公開しているエージェント型カスタマーサービス業務向けのオープンソースベンチマークです。多段階の推論、厳格なポリシー遵守、協調的なツール使用を評価対象とします。このベンチマークは、エージェントがシステムを正しい最終状態へ導けたかどうかで採点されます。最新のドメインは意図的に知識集約型に設計されており、エージェントは行動する前に適切なポリシーを見つけ、適用する必要があります。間違ったバージョンのポリシーに基づいた妥当な回答であっても、得点はゼロとなります。
これが Nexus が構築された目的です。そこで、97 のタスクからなる banking_knowledge ドメインで評価を行いました。このドメインでも、エージェントは行動する前に正しいポリシーを見つけ、適用することが求められます。
| GPT-5.2 | Nexus を使用した GPT-5.2 | GPT-5.5 | Nexus を使用した GPT-5.5 | |
|---|---|---|---|---|
| タスクあたりのツール呼び出し数 | 42.5 | 17.7 | 28.6 | 16.0 |
| タスクあたりのモデル呼び出し数 | 81.7 | 42.6 | 60.9 | 39.4 |
※ 数値は 2026 年 8 月 4 日時点のベンチマーク結果です。
GPT-5.2 は単独でタスクを処理する際、平均して 42.5 回のツール呼び出しと 81.7 回のモデル呼び出しを行い、文書を読み込んで回答を組み立てていました。しかし Nexus を利用すると、その様子は一変します。必要な情報を直接問い合わせるだけで済むため、ツール呼び出しは 17.7 回、モデル呼び出しは 42.6 回に減り、KnowQL クエリもタスクあたり約 6 回で収まります。
モデル呼び出しが半分になり、かつ各呼び出しのコンテキスト量も抑えられることで、1 タスクあたりのコストは 1.45 ドルから 0.53 ドルへと劇的に低下しました。GPT-5.5 も同様の傾向を示し、ツール呼び出しが 28.6 回から 16.0 回へ、モデル呼び出しが 60.9 回から 39.4 回へと削減されました。
このコスト優位性は、GPT-5.2 では 97 タスク中 97 タスクで、GPT-5.5 では 97 タスク中 96 タスクで維持されており、これは一部のタスクに偏った結果ではなく、全体的な傾向として確実です。
なぜこれほど大きな差が生まれるのか。モデル自体には十分な推論能力があります。しかし、安価にアクセスできる「接地された知識(grounded knowledge)」が不足しているのです。エージェントがトークン予算の大半をポリシー文書の検索や再読に費やすと、コストは膨らみます。一方、1 回の呼び出しでクエリ可能な、統合され管理されたレイヤーを提供すれば、推論能力は本来のタスク解決に集中できます。
エンタープライズエージェントが現場で直面する課題
エンタープライズでは、金融、保険、法務、小売、サポートなどにおいて、デモ段階から実運用(プロダクション)への移行が進んでいます。しかしこの移行は、規制の厳しい業界にとって致命的な 4 つの欠陥を同時に露呈させる結果となりました。
- タスク完了と精度: 複雑なコーパスにおいては、現場が求める水準に達しないまま停滞するケースが多いです。
- トークンコスト: エージェントが返す価値に対して、トークン請求額の方が急激に増加しています。
- 高いレイテンシ: タスクあたりの応答遅延が、生産環境のサービスレベルを破綻させています。
回答には出典が明記されていないため、主張の根拠となる文書や条項を遡って確認することができません。
多くの企業はモデル自体に責任を転嫁し、次期バージョンのリリースを待ちます。しかしデータはこのアプローチを支持していません。生産環境でエージェントを実行している 306 チームを対象とした調査では、信頼性がモデル能力よりも上位の開発課題として挙げられ、68% のチームが人間の介入が必要になる前にエージェントの処理ステップ数を 10 段階に制限しています。一方で、コストは安価なモデルを導入しても改善されないことが明確です。統合推論のコストは前年比で約 67% 低下しましたが、平均的な企業の AI 予算は 2024 年の 120 万ドルから 2026 年には 700 万ドルへと急増しました。その理由は、エージェントのタスク実行には多数のモデル呼び出しが必要であり、それぞれでそれまでに収集したコンテキストを再送信しているからです。
ゴールドマン・サックスは、2026 年から 2030 年にかけてトークン消費量が 24 倍に増加すると予測しています。そのため、タスクあたりの無駄が複利のように蓄積していきます。モデルを大きくしても、検索に比例して増大する請求額や、エージェントが到達可能なコンテキストによって制限される完了率の問題は解決しません。
精度、レイテンシ、コスト、信頼性における失敗点には共通の根本原因があります。それは、エージェントが推論を行う前に行う作業です。タスクを読み込み、コンテキストを検索し、結果を確認して再度検索を行い、最初のアクションを実行する前に収集したすべての情報を次のターンで再送信します。トークンとレイテンシの予算の大部分は、エージェントが何らかの判断を下す前に消費されてしまいます。
検索自体には情報の欠落が避けられません。ベクトル検索やハイブリッド検索で返されるのは、関連する文脈を切り離されたテキストの断片のみです。エージェントはこれらの断片を受け取り、リクエストごとに「ポリシーと記録がどう結びついているか」や「ある条項が他方をどのように修飾しているか」を再構築しなければなりません。答えが文章そのものではなく、関係性に依存する場合は、上位 K 件の検索結果では見逃されてしまいます。知識集約型のタスクにおいて、間違ったバージョンのポリシーに基づいて流暢に回答したとしても、得点はゼロです。
アジェンティック RAG は、クエリごとに文脈を再導出し、データやタスクが変更されるたびに再埋め込みを行います。Palantir や Microsoft が提唱する「モデルで事業全体を表現する」アプローチである中央オントロジーは、実際の作業を行わないチームによって事前に作成され、リリースされた瞬間から陳腐化が始まります。知識はクエリ時に組み立てられるか、一度モデル化されて放置されるかのどちらかです。これが現状の限界です。
Nexus の仕組み:一度コンパイルすれば、その後は毎回回答可能に
Nexus は、検索処理をクエリごとのループから外します。レコードとなるシステムデータを、管理されたドメイン固有の知識として事前に一度だけコンパイルし、エージェントはそれ以降の呼び出しでこのコンパイル済みレイヤーを再利用します。これを可能にする三つの要素があります:
マニフェスト。 専門家が自らの言葉で業務内容を記述し、それが「マニフェスト」として完成します。ここで重要なのは、どのエンティティが意味を持ち、それらがどう関係し合い、どのような回答が必要かという形です。ドメインを理解しているのは中央のモデリングチームではなく、現場の専門家自身です。また、マニフェストは企業全体を対象とするのではなく、特定の業務に限定して策定されます。
コンパイルされた知識レイヤー。 マニフェストの指針のもと、Nexus は生データを構造化された知識アーティファクトに変換します。要約、構造化抽出、そしてトップ K 検索では捨てられてしまうエンティティと関係性のグラフなどが含まれます。
KnowQL。 エージェントは、エージェント向けに設計された宣言型言語「KnowQL」を通じて、コンパイルされたレイヤーを照会します。エージェントは必要なもの、質問、出力の形式、スコープ、根拠となるデータ、そして予算を指定するだけです。すると、1 回の呼び出しで、出典が明記され型付けされた回答が返されます。
知識レイヤーで解決した 4 つの失敗
精度: コンパイルされたレイヤーは事実間の関係性を保持し、各フィールドごとの出典と信頼度を維持します。そのため、エージェントは断片的な文章の寄せ集めではなく、相互に関連付けられ根拠のある知識を得られます。また、専門家がすでに矛盾を解消しているため、タスク完了率は向上し、パイロット段階で停滞する要因が取り除かれます。これは監査官も認める回答品質を実現します。
レイテンシ: 事前コンパイルされたレイヤーに対して KnowQL を 1 回呼び出すだけで、「検索→評価→再検索」というループとその往復通信を不要にします。これにより、エージェントはタイムアウトしてユーザーを待たせることなく、生産環境で求められるサービスレベルを満たすことが可能になります。
コスト: 知識を一度コンパイルして再利用することで、エージェントの請求書における最大項目を排除できます。また、再キュレーションプロセスは全体のコーパスではなく変更された部分のみを対象とするため、AI 関連支出が予測可能かつ上限付きとなり、クエリごとにスケールする心配もなくなります。これによりコストが十分に低下し、小規模なモデルやオープンウェイトモデルでも実用が可能になります。
信頼性: ガバナンスはデータレイヤーに組み込まれ、プロンプトで要求されるのではなく構築によって強制されます。アクセス制御は検索時に適用され、各フィールドには出典と信頼スコアが付与されます。個人情報は取り込み時点でタグ付けされ、回答のすべてがソースまで遡れます。コンパイルされたレイヤーは、Pinecone への常時接続なしに自社のクラウド内で実行され、選択したモデル上で動作します。エージェントはセキュリティ審査を通過し、規制産業でもデプロイ可能です。
自社サポートキューで検証
7 月 17 日、Nexus を自社のサポートエージェントのバックエンドに導入しました。
| 指標 | Nexus なし | Nexus あり |
|---|---|---|
| 解決率 | 24.6% | 55.1% |
| 割当率 | 76.5% | 94.2% |
| 支援率 | 60.5% | 87.8% |
現在、サポートチケットの半数以上が、担当者が直接対応することなく解決されています。これは、Nexus が顧客アカウントに関する文脈情報を保持しているおかげです。これにより、エージェントは利用可能なすべての情報に基づいて推論し、質問を「既知の情報」「検索で得られる情報」「顧客または他チームにしかわからない情報」の 3 つのカテゴリに分類できます。このうち「他チームにしかわからない情報」というカテゴリにおいて、Nexus がビジネス指標に対する全体的なパフォーマンスに大きな差をもたらしました。
真実性を保つ知識層
パブリックプレビューに参加した顧客は 300 のコンテキストを作成し、350 万個のソース断片を約 26,000 の構造化され、クエリ可能な知識アーティファクトにまとめました。これらのプロジェクトを通じて取り込まれたコーパスには、サポートナレッジベース、法的契約書、財務報告書、研究論文、議事録、通話記録などが含まれています。私たちは「単一の質問に対して、コーパス全体から複数のファイルを参照する」という制約付きの環境を求めましたが、その期待に応える結果が得られました。
これらの取り組みから得られた最も重要な知見は、顧客が精度を検証するのは非常に早く、通常は評価開始初週に行われるということです。残りの期間では、別の問いに焦点が当てられます。つまり、「この知識層をどのようにして生きているものとして管理していくか」という課題です。ほぼすべての会話で、以下の 3 つの要求が共通して提起されました。
知識の鮮度を保つ仕組みは、手間がかからないものでなければなりません。エンタープライズ内の情報資産は静止したままではありません。新しいチケットが毎日届き、契約書が改訂され、プロセス文書が火曜日に更新されることもあります。プレビュー顧客は、新規ソースデータが完成されたレイヤーに、再構築をゼロから行うことなく、到着次第逐次的に流し込まれることを望んでいます。
Nexus は逐次キュレーションを実現します。新しい情報や変更されたソースは、既存の知識層へ追加されるだけであり、完全な再構築をトリガーすることはありません。
知識レイヤーはビジネスの変化に追従する必要があります。 適切な知識構造は、業務の変化に応じて変わります。収益チームがパイプラインの段階を再編成したり、コンプライアンスチームが新たな規制を引き継いだりします。3 ヶ月目にユーザーが抱く質問と、1 ヶ月目の質問は異なるものです。
現在はこの作業は SME(ドメインエキスパート)が直接担当しています。新しい要件を反映するためにマニフェストを更新し、キュレーションを行い、リリースする。このループは迅速であり、ドメインを理解している人物がコントロールを維持できます。エージェントがこのレイヤーに対して実行するすべてのクエリも、レイヤーに何が含まれるべきかを示すシグナルとなります。Manifest ドライブ型のアーキテクチャはこのシグナルを読み取ることができます。私たちはそこに投資を進めています。
知識レイヤーはソース間の矛盾に対応できる必要があります。 ウィキにはある内容が記され、契約書には別のことが書かれており、そのどちらかが 3 年前の古い情報である場合もあります。検索システムがエージェントに両方の情報を渡すと、自信を持って間違った回答が生成されてしまいます。
キュレーション機能は、SME が判断を下せるよう、矛盾する点を完成された知識層上で浮き彫りにします。知識レイヤーは「自分が知っていること」を把握し、「議論の余地があること」をフラグとして示すべきです。
プレビュー結果から、一つの設計上の確信が裏付けられました。真に持続的な価値を生むのは、専門家が長期間にわたり正確性を保てる「知識レイヤー」です。
静的で宣言的なコンテキストや、中央集権型のオントロジーに基づくアプローチは、一度知識を定義するとその後の変化に対応できず、陳腐化してしまいます。一方、Nexus はドメインの現場を知る担当者の指示に従い、要件の変化に応じて再コンパイルされます。
構築を開始する
もしあなたのエージェントがコーパス上で信頼性に欠ける場合、あるいはトークン数とレイテンシのコストが増加し続けるのに、精度向上の実感が伴わないのであれば、その壁は知識レイヤーに起因しています。これが Nexus が解決するために設計された課題です。
今日から標準的な調達プロセスを通じてこの問題を解消できます。
Pinecone Nexus は現在、一般提供を開始しました。詳細は pinecone.io/nexus をご覧ください。また、今すぐトライアルを始める ことも可能です。
原文を表示
Five weeks ago we opened Pinecone Nexus to Public Preview. Today Nexus is generally available to be deployed in your own cloud.
On τ-Knowledge, Sierra's open benchmark for agentic customer-service work, we gave GPT-5.5 and GPT-5.2 a Nexus knowledge layer and ran them against the same models using the benchmark's own tools. GPT-5.5 held its accuracy at 77% less cost per task. GPT-5.2 achieved 12% more accuracy and saw 80% cost reduction. Both cut their tool calls and model calls roughly in half.
We also pointed Nexus at our own customer support queue. The agent handling inbound tickets went from resolving 25% of them on its own to 55%.
Here’s why the benchmark reads the way it does, what Public Preview taught us about how enterprises manage knowledge, and what GA changes for you.
What's Ready Now
The durable advantage an enterprise has is its knowledge and how its people do the work. Every agent call that ships that knowledge to a third-party model transfers a piece of that advantage to someone else. Nexus keeps it where it belongs: in a governed layer, in your cloud, shaped by your own experts. That principle runs through everything shipping today. Nexus is ready for the standard enterprise motion: evaluate, pilot, procure, run in production. As of today:
Deployment. The Nexus data plane runs in your own cloud, on AWS, Google Cloud, or Azure. Your documents and compiled knowledge never leave your infrastructure.
Model Choice. Nexus runs on the models you choose, including open-weight. You supply the model credentials, and inference calls go from your cloud to the provider you name. Each workflow can use an ensemble, with the right model picked for each step.
No Lock-In. The knowledge layer Nexus compiles is yours. You can download it as an archive.
One Interface. Agents, chatbots, AI search, and recommendation systems all query the same layer through KnowQL. Find the specifications at spec.knowql.org.
Platform Fit. Nexus sits within the broader Pinecone platform, with Pinecone Database as its retrieval foundation and Pinecone Marketplace offering production-ready knowledge apps.
All of that rests on one claim: enterprise agents hit a knowledge ceiling long before they hit a model ceiling. To check it, we ran the following test.
The Benchmark: Same Models, Different Knowledge
τ-Knowledge is Sierra's open-source benchmark for agentic customer-service work: multi-step reasoning, strict policy adherence, coordinated tool use. It grades on whether the agent drives the system to the correct end state. Its newest domains are knowledge-intensive by design. The agent has to find and apply the right policy before it acts, and a plausible answer grounded in the wrong version of a policy scores zero. That is the workload we built Nexus for, so we ran its banking_knowledge\verb|banking_knowledge| domain that has 97 tasks, where the agent has to find and apply the right policy before it acts.
| GPT-5.2 | GPT-5.2 with Nexus | GPT-5.5 | GPT-5.5 with Nexus | |
|---|---|---|---|---|
| Tool calls per task | 42.5 | 17.7 | 28.6 | 16.0 |
| Model calls per task | 81.7 | 42.6 | 60.9 | 39.4 |
Note: Results are from benchmarks ran as of Aug 4, 2026.
GPT-5.2 on its own averaged 42.5 tool calls and 81.7 model calls per task, grinding through documents to assemble an answer. With Nexus it asked for the answer instead: 17.7 tool calls, 42.6 model calls, and about six KnowQL queries per task. Half the model calls, and each one carrying less context, is where a $1.45 task becomes a $0.53 task. GPT-5.5 moved the same way, from 28.6 tool calls and 60.9 model calls down to 16.0 and 39.4. The cost advantage held on 97 of 97 tasks for GPT-5.2 and 96 of 97 for GPT-5.5, so this is not an average hiding a wide spread.
Why the gap is this large: the models by themselves have enough reasoning capability. What they lack is grounded knowledge they can reach cheaply. And it gets expensive for an agent when it spends most of its token budget locating and re-reading policy documents. Give the same model a compiled, governed layer it can query in one call and the reasoning gets spent on the task.
Where Enterprise Agents Break in Production
Enterprises have moved agents out of demo and into production across finance, insurance, legal, retail, and support. The move exposes four failures at once which can be a deal-breaker for enterprises in regulated industries:
Task completion and accuracy on hard corpora stalls short of what production needs.
Token bills climb faster than the value the agents return.
High latency per task breaks production service levels.
And the answers carry no citation, so nobody can trace a claim back to the document and clause it came from.
Most enterprises blame the model and wait for the next frontier release. But the data doesn’t support this approach. In a survey of 306 teams running agents in production, reliability outranked model capability as the top development challenge, and 68% cap their agents at ten steps before a human steps in. Meanwhile it was evident that cost does not respond to cheaper models. Blended inference costs fell about 67% year over year while average enterprise AI budgets rose from $1.2 million in 2024 to $7 million in 2026, because one agent task runs many model calls, each re-sending the context gathered so far. Goldman Sachs projects token consumption to multiply 24x between 2026 and 2030, so the waste per task compounds. A bigger model does not fix a bill that scales with retrieval, or a completion rate capped by the context an agent can reach.
Failure points in accuracy, latency, cost, and trust share one root cause, which is the work an agent does before it reasons. It reads the task, searches for context, reads the result, searches again, and re-sends everything it has gathered on the next turn before it takes a single action. Most of the token and latency budget is spent before the agent decides anything.
The retrieval itself is lossy. A vector or hybrid search returns the top matching chunks of text, stripped of the relationships that connect them. The agent gets fragments and has to reconstruct, on every request, how a policy connects to a record or how one clause qualifies another. When the answer depends on the relationship rather than the passage, top-K retrieval misses it. In a knowledge-intensive task, a fluent answer grounded in the wrong version of a policy scores zero.
Agentic RAG re-derives context on every query and re-embeds whenever the data or the task changes. Central ontologies, the model-the-whole-business approach from Palantir and Microsoft, are authored up front by a team that does not do the work, and they decay from the day they ship. Either the knowledge is assembled at query time, or it is modeled once and left to drift. That is the ceiling.
How Nexus Works: Compile Once, Answer Every Time
Nexus moves the retrieval work out of the per-query loop. It compiles your systems-of-record data into governed, domain-specific knowledge once, ahead of time, and agents reuse that compiled layer on every call. Three parts that make it work:
The Manifest. A subject matter expert describes the work in their own terms, and that description becomes a Manifest: the entities that matter, the relationships between them, and the shape of the answers the work requires. The person who understands the domain defines it, not a central modeling team. A Manifest is scoped to a job rather than to the whole company.
The compiled knowledge layer. Guided by the Manifest, Nexus compiles raw sources into structured knowledge artifacts: summaries, structured extracts, and the entity-and-relationship graph that top-K retrieval throws away.
KnowQL. Agents query the compiled layer through KnowQL, a declarative language built for agents. The agent states what it needs, the question, the output shape, the scope, the grounding, and the budget. It gets back a typed, cited answer in one call.
Four Failures, Solved at the Knowledge Layer
Accuracy: The compiled layer keeps the relationships between facts and carries per-field citations and confidence, so an agent gets connected, grounded knowledge rather than a bag of passages, with conflicts already resolved by the expert. Task completion clears the ceiling that keeps agents stuck in pilot, on answers an auditor will accept.
Latency: One KnowQL call against a precompiled layer replaces the retrieve-evaluate-re-retrieve loop and its round trips. Agents meet production service levels instead of timing out and losing the user.
Cost: Compiling knowledge once and reusing it removes the largest line item in an agent's bill, and re-curation processes only what changed rather than the whole corpus. AI spend becomes predictable and capped instead of scaling with every query, and low enough that smaller and open-weight models become viable.
Trust: Governance lives at the data layer, enforced by construction rather than requested in a prompt. Access control is applied at retrieval. Every field carries a citation and a confidence score. PII is tagged at ingest. Each answer traces back to its source. The compiled layer runs inside your own cloud with no standing Pinecone access, on the models you choose. Agents pass a security review and deploy in regulated industries.
We Ran It On Our Own Support Queue
We put Nexus behind our own support agent on July 17th.
| Metric | Without Nexus | With Nexus |
|---|---|---|
| Resolution rate | 24.6% | 55.1% |
| Assign rate | 76.5% | 94.2% |
| Assist rate | 60.5% | 87.8% |
More than half of our tickets now close without a person touching them. This was made possible because Nexus holds contexts about our customer accounts, so the agent can reason across everything available to it and sort a question into three buckets: what it already knows, what it can look up, and what only the customer or another team can tell it. That third bucket is where Nexus made a significant difference to the overall performance against business metrics.
A Knowledge Layer That Stays True
Public Preview customers created 300 contexts, compiling 3.5 million source chunks into nearly 26,000 structured, queryable knowledge artifacts. The corpora that flowed through those projects: support knowledge bases, legal contracts, financial filings, research papers, meeting minutes, and call transcripts. We asked for bounded corpora where a single question draws on files across the corpus, and that is what we got.
The headline learning from those engagements: customers validate accuracy fast, usually in the first week of evaluation. The rest of the engagement goes to a different question. How do we manage this knowledge layer as a living thing? Three demands came up in almost every conversation.
Keeping Knowledge Current Has To Be Effortless. Enterprise corpora do not hold still. New tickets land daily, contracts get amended, a process doc gets revised on a Tuesday. Preview customers wanted new source data flowing into the compiled layer as it arrives, incrementally, without rebuilding from scratch. Nexus curates incrementally: new and changed sources flow into the existing knowledge layer instead of triggering a full rebuild."
The Knowledge Layer Has To Follow The Business. The right knowledge structure changes as the work changes. A revenue team reorganizes its pipeline stages. A compliance team inherits a new regulation. The questions people ask in month three are not the questions from month one. Today the SME handles this directly: update the Manifest to reflect the new requirements, re-curate, ship. That loop is fast, and it keeps the person who understands the domain in control. Every query an agent runs against the layer is also a signal about what the layer should contain, and a Manifest-driven architecture can read that signal. We are investing there.
The Knowledge Layer Has To Handle Source Conflicts. The wiki says one thing, the contract says another, and one of them is three years stale. A retrieval system hands the agent both, which is how confident wrong answers get made. Curation surfaces conflicts in the compiled knowledge, where the SME can adjudicate them. A knowledge layer should know what it knows and flag what is contested.
The preview confirmed a design conviction. The durable value is a knowledge layer your experts can keep true over time, well past the first curation run. Approaches built on static, declarative context, central ontologies included, define knowledge once and let it decay. Nexus recompiles as the outcome requirements change, guided by the person with direct experience of the domain.
Start Building
If your agents are unreliable on your corpus, or token and latency costs keep climbing without the accuracy to show for it, the ceiling is the knowledge layer. That is the problem Nexus was built for, and as of today you can solve it with a standard procurement conversation.
Pinecone Nexus is generally available now. Learn more at pinecone.io/nexus, or Start Your Trial today.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み