SAM 3.1: マルプレックスとグローバル推論による高速・高アクセシビリティなリアルタイム動画検出と追跡
MetaはSAM 3の効率向上版「SAM 3.1」を発表した。オブジェクト多重化技術により処理速度を大幅に向上させ、既存モデルとの互換性を維持しつつ、リアルタイムの動画検出と追跡性能を強化した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
*2026 年 3 月 27 日更新:*
過去数ヶ月間で SAM 3 の驚くべき普及ぶりを目の当たりにし、その間も私たちは裏側で動画処理効率を向上させるためのアップデートに取り組んでまいりました。本日、SAM 3.1 をご紹介します。
SAM 3 のドロップイン代替品として、オブジェクトの多重化(multiplexing)を導入することで、モデルが単一のフォワードパスで最大 16 個のオブジェクトを追跡可能となりました。この革新により、中程度の数のオブジェクトを含む動画の処理速度が倍増し、単一の H100 GPU 上で 1 秒あたり 16 フレームから 32 フレームへのスループット向上を実現します。その結果、SAM 3.1 は複雑な動画におけるリアルタイムオブジェクト追跡を可能にしつつ、全体の GPU リソース要件を削減し、より小型でアクセスしやすいハードウェアでも高性能アプリケーションの実行を現実的なものとしています。

この改善は、モデルが複数のオブジェクトを扱う方法の転換によるものです。従来では各オブジェクトに専用のパスが必要でしたが、マルチプレクシングを採用した SAM 3.1 では、追跡対象となるすべてのオブジェクトをまとめて処理するため、冗長な計算やメモリのボトルネックが解消されます。このグローバル推論アプローチにより、パフォーマンスが効率化され、混雑したシーンにおける精度も向上します。
コミュニティの皆様には、SAM 3.1 のモデルチェックポイントのダウンロード、SAM 3 コードベースおよび研究論文へのアップデート確認、そして Segment Anything Playground での更新版モデルの実機テストを推奨いたします。
SAM 3.1 Model CheckpointSAM 3 CodebaseSAM 3 Research PaperExplore the Playground
Meta Segment Anything Model 3 および Segment Anything Playground の紹介
要点:
- Meta Segment Anything Model 3(SAM 3)を発表します。これは、テキスト、例示、視覚的なプロンプトを用いて画像および動画内のオブジェクトを検出・セグメンテーション・追跡するための統一モデルです。
- 今回のリリースの一環として、SAM 3 のモデルチェックポイント、評価用データセット、ファインチューニングコードを公開します。
- また、Segment Anything Playground という新しいプラットフォームも導入しました。これは、誰でも SAM の機能を理解し、クリエイティブなメディア加工のための最先端 AI モデルを実験できる環境を提供するものです。
- Instagram の動画作成アプリ「Edits」では、まもなく SAM 3 を活用した新エフェクトが利用可能になります。これにより、クリエイターは動画内の特定の人物やオブジェクトに効果を適用できるようになります。また、SAM 3 に基づく新しい創作体験は、Meta AI アプリの「Vibes」およびウェブ上の meta.ai にも順次導入されます。
- 別途、単一画像からの 3D オブジェクトおよび人間の再構築を可能にするオープンソースモデル、コード、データセットのスイートである SAM 3D も公開しました。これは、物理世界のシナリオにおける grounded 3D 再構築(grounded 3D reconstruction)の新たな基準を設定するものです。詳細は SAM 3D のブログ記事をご覧ください。
- SAM 3 と SAM 3D は、Facebook Marketplace の新機能「View in Room」を駆動しています。これにより、ユーザーは購入前に、ランプやテーブルなどのインテリアアイテムが自身の空間にどのように映るか、スタイルやフィット感を視覚化できるようになります。
- Conservation X Labs および Osa Conservation とのパートナーシップのもと、SAM 3 を活用した野生生物モニタリング用の初の公開動画データセットも立ち上げました。
私たちは、画像および動画の理解を前進させる、Segment Anything モデルコレクションの次世代版を発表します。Segment Anything Model 3(SAM 3)は、テキストや見本プロンプトなど、これまで最も多く要望されていた機能の一部を導入し、画像や動画全体にわたるあらゆる視覚概念の検出、セグメンテーション、追跡を可能にしました。また、より多くの人々が当社のモデルを利用しやすくすることも目指しています。今回のリリースの一環として、最先端のモデルをメディア加工に応用して実験できる最も簡単な方法である Segment Anything Playground をお披露目します。
本日、SAM 3 のモデル重み(weights)、Segment Anything Playground 上のデモ、および SAM 3 の構築方法を詳細に解説した 研究論文 を公開いたします。さらに、コミュニティのための新たなベンチマークとして機能する Concepts との組み合わせによる Segment Anything(SA-Co)評価データセット も共有します。別途、物体およびシーン再構築用のモデルと、人間の姿勢・形状推定用のモデルを含む SAM 3D も公開いたします。今回のリリースに関する詳細は、SAM 3D のブログ記事でご確認ください。
Meta では、これらの技術的進展を活用して、次世代のクリエイティブメディアツールの構築に取り組んでいます。SAM 3 と SAM 3D は、Facebook Marketplace の新機能「View in Room」を実現するために利用されており、ユーザーが購入前に自宅空間にランプやテーブルなどのインテリアアイテムを配置した際のスタイルやフィット感を視覚的に確認できるよう支援しています。SAM 3 によって可能になる新たな創作体験は、Meta AI アプリの「Vibes」およびウェブ版 meta.ai に順次導入されます。ここでは、AI を活用したビジュアル作成ツールを利用したり、既存の AI 生成動画をリミックスしたりすることが可能です。また、間もなく Edits アプリでも SAM 3 を活用した新しいエフェクトを導入する予定です。クリエイターは、動画内の人物やオブジェクトに動的なエフェクトを適用できるようになり、複雑な編集ワークフローをワンタップで完了させることが可能になります。
Meta Segment Anything Model 3 の紹介
画像や動画内の言語と特定の視覚的要素を結びつけることは、コンピュータビジョンにおける主要な課題の一つです。従来のモデルは、固定されたテキストラベルセットを用いたオブジェクトセグメンテーションに焦点を当てることが多く、その結果、ユーザーの多様な要望に応える能力が制限されています。特に、ユーザーの要求には事前定義されたリストに含まれていない概念のセグメンテーションが含まれることが頻繁にあります。つまり、既存モデルは「人」のような一般的な概念であればセグメント化できますが、「縞模様の赤い傘」といったより微妙なニュアンスを持つ概念については対応に苦慮します。
SAM 3 は、テキストや見本画像によるプロンプトで定義された概念のすべてのインスタンスを検出・セグメントする「プロンプタブルな概念セグメンテーション」機能の導入により、これらの制限を克服します。SAM 3 は、オープンボキャブラリの短い名詞句であるテキストプロンプトと、画像見本プロンプトを受け付けることで、固定されたラベルセットという制約を排除しました。大規模語彙における検出およびセグメンテーション性能を評価するため、私たちは「Segment Anything with Concepts (SA-Co)」ベンチマークを作成しました。これは、画像および動画内でのプロンプタブルな概念セグメンテーションを対象としたものであり、以前のベンチマークと比較してはるかに大規模な語彙の概念認識をモデルに要求するものです。今回のリリースの一環として、再現性とオープンエンドな視覚セグメンテーションにおけるさらなるイノベーションを支援するため、SA-Co を一般公開します。
SAM 3 は、SAM 1 および SAM 2 で導入されたマスク、ボックス、ポイントといったビジュアルプロンプトに加え、単純な名詞句や画像見本などの概念プロンプトなど、多様なプロンプトモダリティをサポートします。これにより、セグメンテーションの柔軟性と使いやすさが向上し、特にテキストだけでは記述が困難または稀な概念に対してその効果が顕著です。
SAM 3 は、短い名詞句で記述されたオブジェクトのセグメンテーションにおいて卓越した性能を発揮し、インタラクティブかつ自然な設定における一般的なユーザー意図を反映しています。また、当モデルは、より複雑なプロンプト(例:「手にお土産箱を持っていないが、座っている人々」)で記述されたオブジェクトのセグメンテーションを行うための、マルチモーダル大規模言語モデル向けの知覚ツールとしても利用可能です。
全体として、SAM 3 は、プロンプト可能概念セグメンテーションベンチマークである SA-Co において、既存システムに対して画像および動画の両方で 2 倍の性能向上をもたらすとともに、インタラクティブな視覚セグメンテーションタスクにおける以前の SAM の能力も改善しています。
AI と人間のアノテーターを用いた新規データエンジンの構築
広範なカテゴリと視覚ドメインにわたって、セグメンテーションマスクおよびテキストラベル付きの高品質アノテーション画像を取得することは大きな課題です。このようなデータは、ウェブ上でスケールして存在していません。オブジェクトカテゴリのすべての出現を網羅的にマスキングすること(特に動画においては)は、人間のアノテーターにとって時間と労力を要する複雑なタスクです。さらに、複数の視覚ドメインにわたる大規模で多様な語彙に対する包括的なカバレッジを構築するには、相当な時間とリソースが必要です。全体として、このプロセスは時間がかかり、かつ高コストとなります。
この課題に対処するため、SAM 3、人間のアノテーター、およびループ内の AI モデルを活用したスケーラブルなデータエンジンを作成しました。これによりアノテーションの速度が劇的に向上し、ネガティブプロンプト(画像や動画に存在しない概念)では人間の約 5 倍、ポジティブプロンプトにおいても困難な微細なドメインにおいて 36% の高速化を実現しています。この人間と AI を組み合わせたシステムにより、400 万を超えるユニークな概念を備えた大規模で多様なトレーニングセットの作成が可能となりました。

SAM 3 や Llama ベースのキャプション生成器などの AI モデルからなるパイプラインが、画像や動画を自動的にマイニングし、キャプションを生成してテキストラベルに解析し、上記の図で「候補」として示される初期セグメンテーションマスクを作成します。
人間と AI の注釈作成者がその後、これらの提案を検証・修正し、データセットのカバレッジを急速に拡大しつつ、データの品質を継続的に向上させるフィードバックループを実現します。AI による注釈作成は、Llama 3.2v モデルを基盤としており、これは注釈作成タスクにおいて人間の精度と同等かそれ以上の性能を発揮するように特別に訓練されたものです。具体的には、マスクの品質が高いかどうかの検証や、画像内の概念に関するすべてのインスタンスが網羅的にマスク処理されているかの確認などが含まれます。
一部の人間による注釈作成タスクを AI 注釈作成者に委譲することで、人間のみで構成される注釈作成パイプラインと比較してスループットを倍以上に引き上げることができます。また、AI 注釈作成者は自動的に簡単な事例をフィルタリングし、現在の SAM 3 のバージョンが失敗する最も困難なケースに人間の注釈作成リソースを集中させることができます。さらに、Wikipedia に基づく概念とその関係性を定義した辞書である「概念オントロジー」を活用して、テキストラベルを共有された概念空間へマッピングし、データ内の頻度の低い概念のカバレッジを増大させます。
このアプローチの有効性は、アブレーションスタディ(要素除去実験)を通じて検証され、AI 注釈と人間注釈の両方を統合することでモデル性能に測定可能な改善がもたらされることを示しています。さらに、完全に自動化されたデータエンジンを用いてデータを生成し、新たな視覚領域やテキスト領域へのカバレッジを自動的に拡大できることも実証されています。
モデルアーキテクチャ
プロンプト可能概念セグメンテーションに優れたモデルを構築するには、個別のタスク固有モデルと比較してすべてのタスクにおいて高いパフォーマンスを維持する必要があります。これは、潜在的なタスク間の競合により、モデル設計およびトレーニングレシピの開発において大きな課題となります。例えば、同一概念の他のインスタンスからそれらを区別する視覚的特徴を必要とするインスタンスの再検出・追跡タスクは、概念のすべてのインスタンスに共通する視覚的特徴を必要とする概念検出タスクと競合します。すべてのタスクを統合されたモデルで解決可能にするためには、適切なアーキテクチャを見つけることが重要なステップです。さらに、新しいタスクやデータが導入される際に壊滅的な忘却(catastrophic forgetting)などの問題を防ぐために、強力なデータレシピの設計も不可欠です。
SAM 3 モデルのアーキテクチャも、Meta からの多くの先行する AI の進歩の上に構築されています。SAM 3 のテキストエンコーダーと画像エンコーダーは、Meta Perception Encoder に由来しており、これは今年 4 月に公開されたオープンソースモデルで、画像認識や物体検出といった日常業務を支援できるより高度なコンピュータビジョンシステムの構築を可能にします。Meta Perception Encoder を採用したことで、以前のエンコーダーの選択肢と比較して性能が大幅に向上しました。検出器コンポーネントは、物体検出にトランスフォーマー(transformer)を初めて使用した DETR モデルに基づいています。SAM 2 で使用されたメモリバンクおよびメモリエンコーダーは、トラッカーコンポーネントの基礎となっています。また、私たちの研究を推進するために、データセット、ベンチマーク、モデル改善など、いくつかのオープンソースコンポーネントも利用しました。
結果
SAM 3 は、画像(SA-Co Gold サブセットで測定)および動画(SA-Co Video で測定)における概念セグメンテーション性能において飛躍的な向上を達成し、既存モデルと比較して cgF1 スコア(モデルが概念を認識し局在化できる度合いを示す指標)を倍増させました。SAM 3 は、Gemini 2.5 Pro に代表される基盤モデルや、GLEE、OWLv2、LLMDet といった強力な専門ベースラインを常に上回っています。ユーザー調査では、最も強力なベースラインである OWLv2 と比較して、約 3 対 1 の割合で SAM 3 の出力が好まれる結果となりました。また、SAM 2 の視覚セグメンテーションタスク(マスクからマスキレットへ、ポイントからマスクへ)においても最先端の結果を達成し、SAM 2 に代表される先行モデルの最先端性能に匹敵するか、あるいはそれを上回る成果を示しました。さらに、ゼロショット LVIS(未掲載)やオブジェクトカウント(CountBench で掲載)といった困難なベンチマークでも顕著な向上が見られました。

この優れた性能は、高速推論を伴っています。SAM 3 は H200 GPU 上で単一画像に対して 100 個を超える検出オブジェクトがある場合でも 30 ミリ秒で動作します。動画においては、推論レイテンシはオブジェクト数に比例してスケーリングし、約 5 つの同時実行オブジェクトに対してほぼリアルタイム性能を維持します。
また、SAM 3 をツールとして活用するマルチモーダル大規模言語モデル(MLLM)である「SAM 3 Agent」を用いることで、「画像の中で馬を制御・誘導するために使用されている物体は何ですか?」といった複雑なテキストクエリのセグメンテーションも可能であることを示しています。この MLLM は、SAM 3 をプロンプトするために名詞句クエリを提案し、返されたマスクを分析して満足いく結果が得られるまで反復処理を行います。参照表現セグメンテーションや推論セグメンテーションのデータに対する学習を行わずとも、ReasonSeg(上記参照)や OmniLabel といった、推論を要する困難な自由テキストセグメンテーションベンチマークにおいて、先行研究を上回る性能を発揮します。
科学分野への応用
SAM 3 はすでに科学分野でのユースケースに応用されています。例えば、Meta は Conservation X Labs と Osa Conservation と協力し、現場の野生生物モニタリングと SAM 3 を組み合わせることで、研究に即した生映像データセットを構築しました。一般公開されている SA-FARI データセット には、100 種以上の動物に関する 10,000 本を超えるカメラトラップ動画が含まれており、各フレーム内のすべての動物に対してバウンディングボックスとセグメンテーションマスクが付与されています。FathomNet は MBARI が主導するユニークな研究協力プロジェクトであり、海洋探査のための AI ツールの発展に取り組んでいます。水中画像に特化したセグメンテーションマスクと新しいインスタンスセグメンテーションベンチマークが、現在 FathomNet データベース を通じて海洋研究コミュニティで利用可能です。SA-FARI と FathomNet は、広範な AI コミュニティによって、陸上および海洋における野生生物の発見・監視・保全に向けた革新的な新手法の開発に活用できます。
オープンソースコミュニティによる今後の探索領域
SAM 3 は、単純なテキストフレーズを用いて画像や短い動画内のオブジェクトをセグメント化する上で強力なパフォーマンスを示していますが、モデルのパフォーマンスはさらに向上させる余地があり、特に困難なシナリオにおいてその可能性が広がります。
SAM 3 は、ゼロショット方式で微細なドメイン外概念(例:「血小板」のように専門知識を要する特定の用語の識別など)への一般化に苦戦します。これは特に医療や科学画像を含むニッチな視覚領域において顕著です。私たちは SAM 3 の能力を拡張するための戦略を実験しましたが、少量のアノテーション済みデータでファインチューニングを行うと、モデルが新しい概念や視覚ドメインに素早く適応できることが分かりました。コードリリースの一環として、コミュニティが自らのユースケースに合わせて SAM 3 を適応させるために活用できるファインチューニング手法を共有します。また、Roboflow と連携し、ユーザーがデータのアノテーション、ファインチューニング、そして特定のニーズに応じた SAM 3 のデプロイを容易に行えるよう支援しています。
さらに、SAM 3 は「ハードカバーの本」のような短いオープンボキャブラリープロンプトに対しては良好に動作しますが、「上段の右から二番目の本」といった長く複雑なフレーズには対応していません。ただし、マルチモーダル大規模言語モデルと組み合わせることで、推論を要するケースを含む、より長く複雑な記述をサポートするようにモデルを訓練することが可能です。
動画に適用すると、SAM 3 は SAM 2 スタイルのマスクレットを用いてすべてのオブジェクトを追跡します。これは、追跡対象となるオブジェクトの数に対して SAM 3 の推論コストが線形にスケールすることを意味します。各オブジェクトは個別に処理され、共有されるフレーム単位の埋め込みのみを利用し、オブジェクト間の通信は行われません。多くの視覚的に類似したオブジェクトが存在する複雑なシーンにおいて、共有されたオブジェクトレベルの文脈情報を組み込むことで、効率性とモデル性能の向上が期待できます。
この分野の研究をさらに推進するためには、まだ多くの取り組みが必要です。私たちは、SAM 3 を活用して構築し、SA-Co ベンチマークを採用し、これらの新しいリソースを活用することで、AI コミュニティの皆様にもご参加いただき、これらの機能をさらに発展させることを願っています。共にオープンサイエンスを加速させ、人々や社会に有益なインパクトのある新たな体験とユースケースを構築していきましょう。
Segment Anything Playground で SAM 3 を探索する
私たちは、これらの研究成果をすべて「Segment Anything Playground」という新しいプラットフォームに統合しました。このプラットフォームを使えば、技術的な専門知識がなくても誰でも最新のモデルを試すことができます。ゼロから始めるフローでは、画像やビデオのアップロードが可能で、利用可能なテンプレートのいずれかを使ってすぐに作業を開始することもできます。これには、顔・ナンバープレート・画面をピクセル化するといった実用的なオプションに加え、スポットライト効果の追加、モーショントレイルの付与、特定のオブジェクトの拡大表示など、楽しいビデオ編集も含まれています。さらに、これらのテンプレートは視覚データの注釈付けを支援し、SAM 3 のストレステストを行う手段としても機能します。私たちは、メディア加工のためのモデル実験を最も簡単に行える場として SAM Playground を設計しました。人々がどのように用它して創造性を高めてくれるのか、今から楽しみです。
SAM 3 は、Meta の Aria Gen 2 研究用グラス などのウェアラブルデバイスによって撮影されたファーストパーソン映像においても高い性能を発揮します。これにより、ファーストパーソン視点からのオブジェクトの堅牢なセグメンテーションとトラッキングが可能となり、ウェアラブルデバイスで撮影されるシーン特有の動的な課題にも対応できます。Aria Gen 2 Pilot Dataset から選ばれた録画の一部が現在、Segment Anything Playground で紹介されています。この統合は、人間視点からの世界理解が不可欠である機械知覚、文脈 AI、ロボティクスなどの分野における研究や応用において SAM 3 が持つ価値を示すものです。
この動画では、SAM 3 Aria Gen 2 の出力を示すために「hands」というプロンプトが使用されています。
原文を表示
*Update March 27, 2026:*
We’ve seen incredible adoption of SAM 3 over the last few months, and during that time, we’ve been working behind the scenes on updates to improve video processing efficiency. Today, we’re pleased to introduce SAM 3.1.
As a drop-in replacement for SAM 3, our updated model delivers a significant boost in video processing efficiency by introducing object multiplexing, which allows the model to track up to 16 objects in a single forward pass. This innovation doubles the processing speed for videos with a medium number of objects, increasing throughput from 16 to 32 frames per second on a single H100 GPU. As a result, SAM 3.1 enables real-time object tracking in complex videos while reducing overall GPU resource requirements, making high-performance applications feasible on smaller, more accessible hardware.

This improvement comes from a shift in how the model handles multiple objects. Previously, each object required its own dedicated pass, but with multiplexing, SAM 3.1 processes all tracked objects together, eliminating redundant computation and memory bottlenecks. This global reasoning approach streamlines performance and enhances accuracy in crowded scenes.
We encourage the community to download the SAM 3.1 model checkpoint, explore the updates to the SAM 3 codebase and research paper, and test drive the updated model on the Segment Anything Playground.
SAM 3.1 Model CheckpointSAM 3 CodebaseSAM 3 Research PaperExplore the Playground
Introducing Meta Segment Anything Model 3 and Segment Anything Playground
Takeaways:
- We’re announcing Meta Segment Anything Model 3 (SAM 3), a unified model for detection, segmentation, and tracking of objects in images and video using text, exemplar, and visual prompts.
- As part of this release, we’re sharing SAM 3 model checkpoints, evaluation datasets, and fine-tuning code.
- We’re also introducing Segment Anything Playground, a new platform that makes it easy for anyone to understand the capabilities of SAM and experiment with cutting-edge AI models for creative media modification.
- In Edits, Instagram’s video creation app, SAM 3 will soon enable new effects that creators can apply to specific people or objects in their videos. New creation experiences enabled by SAM 3 will also be coming to Vibes on the Meta AI app and meta.ai on the web.
- Separately, we’re sharing SAM 3D, a suite of open source models, code, and data for 3D objects and human reconstruction from a single image, setting a new standard for grounded 3D reconstruction in physical world scenarios. Learn more by reading the SAM 3D blog post.
- SAM 3 and SAM 3D are powering Facebook Marketplace’s new View in Room feature, helping people visualize the style and fit of home decor items, like a lamp or a table, in their spaces before purchasing.
- Together with our partners at Conservation X Labs and Osa Conservation, we’re also launching a first-of-its-kind, publicly available video dataset for wildlife monitoring using SAM 3.
We’re unveiling the next generation of the Segment Anything collection of models, advancing image, and video understanding. Segment Anything Model 3 (SAM 3) introduces some of our most highly requested features like text and exemplar prompts — enabling detection, segmentation, and tracking of any visual concept across images and video. We also want to make it easier for more people to use our models. As part of this release, we’re debuting the Segment Anything Playground, the simplest way for anyone to experiment with applying our state-of-the-art models to media modification.
Today, we’re releasing the SAM 3 model weights, a demo on Segment Anything Playground, and a research paper that details how we built SAM 3. Additionally, we’re sharing the Segment Anything with Concepts (SA-Co) evaluation dataset to serve as a new benchmark for the community. Separately, we’re sharing SAM 3D, which includes a model for object and scene reconstruction and another for human pose and shape estimation. More information about this release can be found in our SAM 3D blog post.
At Meta, we’re using these advancements to help build the next generation of creative media tools. SAM 3 and SAM 3D are being used to enable the new View in Room feature on Facebook Marketplace, helping people visualize the style and fit of home decor items, like a lamp or a table, in their spaces before purchasing. New creation experiences enabled by SAM 3 will be coming to Vibes on the Meta AI app and meta.ai on the web, where people can use AI visual creation tools and remix existing AI-generated videos. We’ll also soon be introducing new effects on our Edits app that use SAM 3. Creators can apply dynamic effects to people or objects in their videos — simplifying a complex editing workflow to just one tap.
Introducing Meta Segment Anything Model 3
Linking language to specific visual elements in images or videos is a major challenge in computer vision. Traditional models often focus on object segmentation with a fixed set of text labels, restricting their ability to address the full spectrum of user requests, which frequently involve segmenting concepts not present in predefined lists. This means that existing models can segment frequent concepts like “person,” but struggle with more nuanced concepts like “the striped red umbrella”.
SAM 3 overcomes these limitations by introducing the promptable concept segmentation capability: finding and segmenting all instances of a concept defined by a text or exemplar prompt. SAM 3 accepts text prompts — open-vocabulary short noun phrases — and image exemplar prompts, eliminating the constraints of fixed label sets. To assess large-vocabulary detection and segmentation performance, we created the Segment Anything with Concepts (SA-Co) benchmark for promptable concept segmentation in images and videos that challenges models to recognize a much larger vocabulary of concepts compared to prior benchmarks. As part of this release, we’re making SA-Co publicly available to support reproducibility and further innovation in open-ended visual segmentation.
SAM 3 supports a variety of prompt modalities, including both concept prompts such as simple noun phrases and image exemplars, as well as visual prompts, such as masks, boxes, and points, which were introduced in SAM 1 and SAM 2. This increases the flexibility and usability of segmentation, particularly for concepts that are rare or hard to describe with text alone.
SAM 3 excels at segmenting objects described by short noun phrases, reflecting common user intent in interactive and natural settings. Our model can also be used as a perception tool for multimodal large language models to segment objects described by more complex prompts, such as: “people sitting down, but not holding a gift box in their hands.”
Overall, SAM 3 delivers a 2x gain over existing systems in both image and video on our promptable concept segmentation benchmark, SA-Co, and improves upon previous SAM capabilities in interactive visual segmentation tasks.
Building a Novel Data Engine Using AI and Human Annotators
Obtaining high-quality annotated images with segmentation masks and text labels across a broad range of categories and visual domains is a significant challenge. This type of data doesn’t exist at scale on the web. Exhaustively masking every occurrence of an object category — particularly in video — is a time-intensive and complex task for human annotators. Additionally, building comprehensive coverage for a large and diverse vocabulary across multiple visual domains requires considerable time and resources. Overall, the process is both time-consuming and expensive.
We address this challenge by creating a scalable data engine that leverages SAM 3, human annotators, and AI models in the loop, which allows dramatic speed-ups in annotation — approximately 5x faster than humans on negative prompts (concepts not present in the image/video) and 36% faster for positive prompts even in challenging fine-grained domains. This hybrid human and AI system enabled us to create a large and diverse training set with over 4 million unique concepts.

A pipeline of AI models, including SAM 3 and systems such as a Llama-based captioner, automatically mine images and videos, generate captions, parse the captions into text labels, and create initial segmentation masks, which are shown as “candidates” in the above figure.
Human and AI annotators then verify and correct these proposals, yielding a feedback loop that rapidly scales dataset coverage while continuously improving data quality. AI annotators are based on Llama 3.2v models that were specifically trained to match or surpass human accuracy on annotation tasks, such as verifying if a mask is high quality, or if all instances of a concept are exhaustively masked in an image.
By delegating some human annotation tasks to AI annotators, we more than double the throughput compared to a human-only annotation pipeline. AI annotators also automatically filter out easy examples, focusing valuable human annotation effort on the most challenging cases where the current version of SAM 3 fails. We also leverage a concept ontology — a dictionary of concepts and their relationships based on Wikipedia — to map text labels into a shared concept space and increase the coverage of less frequent concepts in the data.
We validate this approach through ablation studies, demonstrating that integrating AI- and human-annotated labels results in measurable improvements in model performance. We further validate that an entirely automated data engine can be used to generate data to automatically expand coverage to new visual and text domains.
Model Architecture
Building a model that excels at promptable concept segmentation requires us to maintain strong performance on all tasks compared to individual, task-specific models. This presents significant challenges in model design and in the development of a training recipe, due to potential task conflicts. For example, the task of re-detecting and tracking instances requires visual features that distinguish them from other instances of the same concept. This conflicts with the concept detection task, which requires visual features that are similar for all instances of a concept. Finding the right architecture is an important step in being able to solve all tasks in a unified model. Additionally, designing strong data recipes is essential to prevent issues like catastrophic forgetting as new tasks and data are introduced.
The SAM 3 model architecture also builds on many previous AI advancements from Meta. The text and image encoders in SAM 3 are from the Meta Perception Encoder, an open source model we shared in April that enables the building of more advanced computer vision systems that can assist people in everyday tasks, such as image recognition and object detection. Using the Meta Perception Encoder enabled us to achieve a significant leap in performance compared to previous encoder choices. The detector component is based on the DETR model, which was the first to use transformers for object detection. The memory bank and memory encoder used in SAM 2 is the basis for the Tracker component. We also used several open source components, including datasets, benchmarks, and model improvements, to advance our work.
Results
We achieve a step change in concept segmentation performance in images (measured on SA-Co Gold subset) and videos (on SA-Co Video), with SAM 3 doubling cgF1 scores (a measure of how well the model can recognize and localize concepts) relative to existing models. SAM 3 consistently outperforms both foundational models like Gemini 2.5 Pro and strong specialist baselines such as GLEE, OWLv2, and LLMDet. In studies, users prefer SAM 3 outputs over the strongest baseline, OWLv2, approximately three to one. We also achieve state-of-the-art results on the SAM 2 visual segmentation tasks (mask-to-masklet, point-to-mask), matching or exceeding the state-of-the-art performance of previous models like SAM 2. Furthermore, we see notable gains on challenging benchmarks like zero-shot LVIS (not shown) and object counting (shown on CountBench).

This excellent performance comes with fast inference — SAM 3 runs in 30 milliseconds for a single image with more than 100 detected objects on an H200 GPU. In video, the inference latency scales with the number of objects, sustaining near real-time performance for approximately five concurrent objects.
We also show that a multimodal large language model (MLLM) that uses SAM 3 as a tool, called SAM 3 Agent, can segment more complex text queries such as, “What object in the picture is used for controlling and guiding a horse?” The MLLM proposes noun phrase queries to prompt SAM 3 and analyzes the returned masks, iterating until the masks are satisfactory. Without training on any referring expression segmentation or reasoning segmentation data, SAM 3 Agent surpasses prior work on challenging free-text segmentation benchmarks that require reasoning, such as ReasonSeg (shown above) and OmniLabel.
Applications to Science
SAM 3 is already being applied for use cases in scientific fields. For example, Meta collaborated with Conservation X Labs and Osa Conservation to combine on-the-ground wildlife monitoring with SAM 3 to build an open dataset of research-ready, raw video footage. The publicly available SA-FARI dataset includes over 10,000 camera trap videos of more than 100 species, annotated with bounding boxes and segmentation masks for every animal in each frame. FathomNet is a unique research collaboration led by MBARI that is working to advance AI tools for ocean exploration. Segmentation masks and a new instance segmentation benchmark tailored for underwater imagery are now available to the marine research community via the FathomNet Database. SA-FARI and FathomNet can be used by the broader AI community to develop innovative new ways to discover, monitor, and conserve wildlife on land and in the ocean.
Future Areas of Exploration for the Open Source Community
While SAM 3 demonstrates strong performance for segmenting objects in images and short videos with simple text phrases, the model performance can be further improved, especially in challenging scenarios.
SAM 3 struggles to generalize to fine-grained out-of-domain concepts in a zero-shot manner, such as identifying specific terms that require domain knowledge like “platelet,” especially in niche visual domains involving medical or scientific imagery. We experimented with strategies to extend the capability of SAM 3 and found that the model quickly adapts to new concepts and visual domains when fine-tuned on small quantities of annotated data. As part of our code release, we’re sharing fine-tuning approaches that the community can leverage to adapt SAM 3 for their use cases. We’re also partnering with Roboflow to enable people to annotate data, fine-tune, and deploy SAM 3 for their particular needs.
Additionally, while SAM 3 performs well with short open-vocabulary prompts, such as “a hardcover book,” the model doesn’t support longer, complex phrases like, “the second to last book from the right on the top shelf.” However, when paired with multimodal large language models, the model can be trained to support longer, more complex descriptions including cases that require reasoning.

When applied to video, SAM 3 tracks every object with a SAM 2-style masklet, which means the cost of SAM 3 inference scales linearly with the number of objects being tracked. Each object is processed separately, utilizing only shared per-frame embeddings, without inter-object communication. Incorporating shared object-level contextual information could aid in improving efficiency and model performance in complex scenes with many visually similar objects.
There’s plenty more work to be done to propel research in this field even further. We hope the AI community will join us by building with SAM 3, adopting the SA-Co benchmark, and leveraging these new resources to help push these capabilities further. Together, we can accelerate open science to build impactful new experiences and use cases that benefit people and society.
Explore SAM 3 on the Segment Anything Playground
We’re bringing all of this work together in the Segment Anything Playground, our new platform that enables anyone to try our latest models — no technical expertise needed. The start-from-scratch flow enables uploading an image or video, or it’s possible to jump right in using one of the available templates. These include practical options like pixelating faces, license plates, and screens, as well as fun video edits such as adding a spotlight effect, motion trails, or magnifying specific objects. Additionally, the templates assist in annotating visual data and provide a way to stress test SAM 3. We’ve designed SAM Playground to be the simplest way to experiment with our models for media modification, and we can’t wait to see how people use it to enhance their creativity.
SAM 3 also performs well on first-person footage captured by wearable devices like Meta’s Aria Gen 2 research glasses. This enables robust segmentation and tracking of objects from a first-person perspective, handling the dynamic challenges of wearable-captured scenes. Select recordings from the Aria Gen 2 Pilot Dataset are now featured on the Segment Anything Playground. This integration demonstrates SAM 3’s value for research and applications in areas like machine perception, contextual AI, and robotics, where understanding the world from the human perspective is crucial.
The prompt “hands” is used in this video showing SAM 3 Aria Gen 2 output.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み