テキスト条件付き JEPA:意味豊かな視覚表現を学習する手法
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Apple Machine Learning
研究者らは、マスクされた位置の視覚的不確実性を軽減するため、画像キャプションを活用した「Text-Conditional JEPA(TC-JEPA)」を提案し、より意味豊かな視覚表現の学習を実現しました。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
画像ベースのJoint-Embedding Predictive Architecture (I-JEPA) は、マスクされた特徴量の予測を通じて視覚的自己教師あり学習への有望なアプローチを提供します。しかし、マスクされた位置における本質的な視覚的不確実性により、特徴量予測は依然として困難であり、意味表現を学習できない可能性があります。本研究では、画像キャプションを用いて予測の不確実性を低減するText-Conditional JEPA (TC-JEPA) を提案します。具体的には、入力テキストトークンに対してスパースなクロスアテンション(sparse cross-attention)を計算する微細なテキストコンディショナーを用いて、予測されたパッチ特徴量を調整します。このような…
原文を表示
Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature prediction. However with the inherent visual uncertainty at masked positions, feature prediction remains challenging and may fail to learn semantic representations. In this work, we propose Text-Conditional JEPA (TC-JEPA) that uses image captions to reduce the prediction uncertainty. Specifically, we modulate the predicted patch features using a fine-grained text conditioner that computes sparse cross-attention over input text tokens. With such…
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み