Onton、神経記号モデル「Ontology 1」を公開し検索精度で世界最高を更新
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
サンフランシスコの検索企業 Onton は、従来のキーワードやベクトル検索を凌ぐ精度を持つニューロシンボリックモデル「Ontology 1」を発表し、独立した評価で主要競合を大幅に上回る結果を示した。
AI深層分析を開く2026年8月3日 10:07
AI深層分析
キーポイント
高い検索精度の実証
90件のクエリによるベンチマークで Ontology 1 が平均 Precision@10 0.630 を達成し、Google Shopping (0.543) や Amazon (0.469) を上回った。
ニューロシンボリックアプローチ
販売者のラベルに依存せず、素材や構造などの客観的性質から推論を行い、データ上の矛盾を検出する世界モデルを構築している。
限定的な展開形態
公開 API やモデルのダウンロードは行われず、エンドユーザー向けサイトとパートナーシップ限定での利用に留まっている。
検索精度と勝敗結果
Onton は P@10 で 0.630 を記録し、Google Shopping (0.543) や Amazon (0.469) を上回った。90 クエリ中 52 クエリで他社を圧倒したが、機能仕様に関するクエリでは競合に劣るケースも確認された。
独自インフラによる高速化
Ograph というカスタムグラフデータベースを採用し、GPU 版は CPU 版より約 43 倍高速である。チューニングが進めばさらに性能が向上する見込みだ。
重要な引用
Onton argues this catalog interface has barely changed in nearly 30 years.
It reasons from properties more likely to be objective — fiber, weave, construction — and flags claims the product data contradicts.
Where Ontology 1 loses: Failure cases cluster on functional-spec queries where Amazon's category metadata dominates
The architecture is neurosymbolic: an inspectable knowledge graph that decomposes vague predicates into checkable properties.
編集コメントを表示
編集コメント
検索技術がラベル依存から推論依存へとパラダイムシフトする重要な兆候を示している。ただし、現時点では特定の業界に限定され公開アクセスも制限されているため、即座の一般適用にはまだ時間がかかるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
サンフランシスコに拠点を置く検索・発見企業であるOntonは、複雑な会話型かつマルチモーダルな商品検索を実現するニューロシンボリックモデル「Ontology 1」をリリースしました。3 つの独立した LLM 判定者による 90 クエリのベンチマークでは、Ontology 1 は平均精度@10 で 0.630 を達成。これに対し、Google Shopping は 0.543、Amazon は 0.469 でした。この成果は、両社のカタログの約 1% しかインデックス化されていない状態で達成されたものです。
導入は可能か?
はい、可能です。ただし、ダウンロードして重み(weights)として使うような形ではありません。Ontology 1 は現在、Onton.com でエンドユーザー向けに稼働しており、Onton によると、エージェント型ウェブ上で構築するチームへのパートナーアクセスはケースバイケースで付与されています。モデル自体の公開 API や価格プラン、オープンなチェックポイントはありません。現状での採用形態は、パッケージインストールではなく、パートナーシップによるものです。
企業との適合性:中堅・大規模小売業者、マーケットプレイス、エージェント型コマースプラットフォームが対象です。これらは、要件が長く複雑なクエリにおいて検索精度の低下に悩まされているケースが多いでしょう。一方、カタログ規模が小さい場合は恩恵は限定的です。Ontology 1 が解決しようとする失敗モードは、カタログサイズやリストノイズの影響を強く受けるためです。
対象業界:現在はインテリア・家具分野のみを対象としています。これは Onton がインデックス化している唯一の垂直領域だからです。Onton は同社の手法が e コマースに限定されず、商品データ以外の非製品データもほぼ再設定なしで検索可能になると主張しています。
応用事例:会話型およびマルチモーダルなサイト内検索、ムードボードに基づく発見機能、否定表現を多用するフィルタリング、リストやレビューの信頼性スコアリング、ショッピングエージェントのためのグラウンディング層などが挙げられます。
なぜキーワード検索とベクトル検索では不十分なのか
従来の EC サイトは、検索意図をカテゴリや属性(サイズ、価格、素材、ブランドなど)にマッピングする仕組みが前提となっています。しかし、「ペット可」といったフィルターや「部屋に合う家具」を検索するための機能は存在しません。Onton は、このカタログ型インターフェースがほぼ 30 年間変わっていないと指摘しています。
Ontology 1 は異なるアプローチを採用します。「ペット可のソファセット」のような検索に対して、販売者が付けたラベル(欠落している場合や事実と異なる場合もある)を信頼するのではなく、繊維の種類、織り方、構造など客観的な属性から推論を行います。また、製品データと矛盾する主張にはフラグを立てます。さらに情報の出所も評価します。アルゴリズムを操作するリストがある一方、購入されたレビューが存在するからです。
このモデルは、パターンを重みに吸収するのではなく、明確で検証可能な世界モデルを構築します。「ペット可」という概念について説明がない場合、それを「欠落」として扱い、回答を導き出します。具体的には、清潔さの維持や耐久性を考慮し、ポリエステル製ファブリックがその指標となり得ると判断します。この学習は、「ペット可の椅子」や「掃除可能な青いソファ」などの後続のクエリでも再利用され、ループは継続的に回ります。
ベンチマーク:Subtext-Decor-90
Onton はコードとデータセットを公開した Subtext-Decor-90 を発表しました。Claude Opus 4.8、Gemini 3.1 Pro、GPT-5.5 の 3 つのマルチモーダル評価者が、90 のテキストクエリそれぞれに対して Onton、Amazon、Google Shopping が返した上位 10 件の表示カードを採点しました。P@10 は各評価者間で平均化され、10,000 回のブートストラップ再サンプリングから得られた 95% 信頼区間が算出されました。
評価結果は、Onton が 0.630 [0.571, 0.688]、Google Shopping が 0.543 [0.490, 0.596]、Amazon が 0.469 [0.417, 0.521] でした。Onton は 52 クエリで完全勝利し、Google が 19、Amazon が 16 です。合計が 87 となっているのは、3 つのクエリで Onton の返答件数が 10 件未満だったためです。空欄は非関連として評価されました。もしこれらの空欄を除外して計算し直すと、Onton は 0.665、Google は 0.549、Amazon は 0.459 となります。
3 人の審査員全体でのクリッペンドルフのアルファ係数は 0.465 です。これは絶対的な P@10 の値がノイズを含み、審査員によって変動しやすいことを意味します。ただし、どの審査員も評価順は同じで、エンジン間の順位付けに揺らぎはありません。
画像やマルチモーダルなクエリ 90 件からは除外されました。Amazon Lens はマルチモーダル検索に対応しておらず、Google Lens も製品を専ら返すわけではないためです。Onton では、Google との比較として別途 10 件の画像・マルチモーダルクエリでの評価結果を報告しています。
Onton 1 が劣るケース
失敗事例は、Amazon のカテゴリメタデータが支配的な「機能仕様」系のクエリに集中しています。具体的には、「深夜 3 時に読んでもパートナーを起こさないランプ」(Onton 0.4、Amazon 0.9)や「奥行きのある窓辺に置けるもの」(Onton 0.07、Amazon 0.67)などです。Onton はこれをカタログの網羅性の不足と、単一垂直領域かつ非広告インデックスであることによる限界として説明しています。今後は自己学習ループを通じてこの差を縮めていく方針です。
基盤となるインフラストラクチャ
Ontology 1 の知識グラフは、Ograph という独自開発のグラフデータベース上で動作しています。Onton によると、Ograph のコア 1 つが、14 コアで稼働する SuiteSparse:GraphBLAS を上回り、コアあたりの処理スループットは約 100 倍に達します。また、GPU 版は CPU 版よりも 43 倍高速で、実装の最適化が進むにつれて初期テストでは 1000 倍もの速度差も確認されています。
インタラクティブな解説動画
以下の埋め込み動画では、4 つのパネルに分けて同じ内容を解説しています。実際の Subtext-Decor-90 のクエリとそのスコア、ペット可の条件を段階的に推論するグラフ、自己学習ループ、そして信頼区間や代替スコアリングを含むベンチマークチャートです。
主なポイント
Ontology 1 は、Subtext-Decor-90 ベンチマークで P@10 が 0.630 を記録し、Google Shopping(0.543)や Amazon(0.469)を大きく上回りました。
90 クエリ中 52 で直接勝利しており、競合他社のカタログの約 1% の規模でこの結果を出しています。
アーキテクチャはニューロシンボリック型です。あいまいな述語を検査可能なプロパティに分解する、検証可能な知識グラフを採用しています。
評価者の信頼性は中程度(Krippendorff's alpha は 0.465)ですが、3 人の評価者全員が各エンジンの順位を同じように付けました。
提供形態は製品ファーストです。Onton.com で直接利用可能で、パートナーへのアクセスはケースバイケース。オープンウェイトや公開 API の提供はありません。
技術詳細とベンチマークについてはこちらをご覧ください。Twitter でもフォローしていただければ幸いです。また、15 万人以上の ML 専門家が参加する SubReddit や、ニュースレターにもぜひご登録ください。あ、Telegram も使っていますか?今なら Telegram でも私たちに参加いただけます。
本記事は MarkTechPost にて公開された「Onton が Ontology 1 を発表:世界最高峰の EC 検索エンジンより 2.7 倍高精度な神経記号型検索モデル」の翻訳です。
原文を表示
Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a 90-query benchmark scored by three independent LLM judges, Ontology 1 reached a mean precision@10 of 0.630, against 0.543 for Google Shopping and 0.469 for Amazon. It did this while indexing roughly 1% of their catalogs.
Is it deployable
Yes, but not as weights you download. Ontology 1 is live for end users at Onton.com, and Onton says partner access is granted case by case for teams building on the agentic web. There is no public API, pricing tier, or open checkpoint for the model itself. Adoption today looks like a partnership, not a pip install.
Company fit: Mid-market and enterprise retailers, marketplaces, and agentic-commerce platforms whose relevance stack already loses on long, requirements-heavy queries. Small catalogs see less benefit, because the failure mode Ontology 1 targets scales with catalog size and listing noise.
Industries: Home decor and furniture today, since that is the only vertical Onton indexes. Onton states the methodology generalizes beyond e-commerce, and that Ontology searches non-product data with essentially no reconfiguration.
Applications: Conversational and multimodal site search, moodboard-driven discovery, negation-heavy filtering, listing and review trust scoring, and grounding layers for shopping agents.
Why keyword and vector retrieval break here
Conventional e-commerce assumes intent maps onto categories and attributes: size, price, material, brand. There is no filter for ‘pet-friendly,’ and none for furniture that fits your room. Onton argues this catalog interface has barely changed in nearly 30 years.
Ontology 1 takes a different route. For ‘pet-friendly sectional,’ it does not trust the seller’s label, which may be absent or untrue. It reasons from properties more likely to be objective — fiber, weave, construction — and flags claims the product data contradicts. It also weighs the source, since some listings game the algorithm and some reviews are bought.
The model builds an explicit, inspectable world model rather than absorbing patterns into weights. When it has no account of ‘pet-friendly,’ it treats that as a gap and works the answer out: cleanability and durability, then polyester upholstery as an indicator. The learning is reused on later queries such as ‘pet-friendly chair’ or ‘cleanable blue couch,’ and the loop runs continuously.
The benchmark: Subtext-Decor-90
Onton released Subtext-Decor-90 with code and data. Three multimodal judges: Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5, scored the top 10 visible result cards returned by Onton, Amazon and Google Shopping for each of 90 text queries. P@10 was averaged across judges, with 95% confidence intervals from 10,000 bootstrap resamples.
Results: Onton 0.630 [0.571, 0.688], Google Shopping 0.543 [0.490, 0.596], Amazon 0.469 [0.417, 0.521]. Onton won 52 queries outright, Google 19, Amazon 16. Those sum to 87 because Ontology returned fewer than 10 results on three queries, and empty slots were scored as non-relevant. Excluding those slots instead gives Onton 0.665, Google 0.549, Amazon 0.459.
Krippendorff’s alpha across the three judges is 0.465, so absolute P@10 values are noisy and judge-dependent. All three judges still place the engines in the same order.
Image and multimodal queries were excluded from the 90, because Amazon Lens does not support multimodal queries and Google Lens does not return products exclusively. Onton reports a separate 10-query image and multimodal comparison against Google.
Where Ontology 1 loses
Failure cases cluster on functional-spec queries where Amazon’s category metadata dominates: ‘lamp that won’t wake my partner if I read at 3am’ (Onton 0.4, Amazon 0.9) and ‘something to put on a weirdly deep windowsill’ (Onton 0.07, Amazon 0.67). Onton attributes this to catalog breadth and its single-vertical, non-sponsored index, and expects the self-learning loop to narrow the gap.
The infrastructure underneath
Ontology 1’s knowledge graph runs on Ograph, a custom graph database. Onton reports one Ograph core beating SuiteSparse:GraphBLAS running on 14 cores, roughly 100× the throughput per core, and a GPU build running 43× faster than the CPU variant, with early runs touching 1000× as the implementation is tuned.
Interactive explainer
The embed below walks through the same material in four panels: real Subtext-Decor-90 queries with per-query scores, the pet-friendly reasoning graph drawn step by step, the self-learning loop, and the benchmark chart with confidence intervals and alternate scoring views.
Key Takeaways
Ontology 1 scores P@10 0.630 on Subtext-Decor-90, ahead of Google Shopping (0.543) and Amazon (0.469).
It wins 52 of 90 queries outright while indexing roughly 1% of either competitor’s catalog.
The architecture is neurosymbolic: an inspectable knowledge graph that decomposes vague predicates into checkable properties.
Judge reliability is modest (Krippendorff’s alpha 0.465), but all three judges rank the engines identically.
Availability is product-first — live on Onton.com, partner access case by case, no open weights or public API.
Check out the Technical details and Benchmarks. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the World’s Best E-commerce Search Engines appeared first on MarkTechPost.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み