VideoFlexTok:柔軟な長さの粗から細への動画トークン化手法
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Apple Machine Learning
Apple Machine Learning は、動画を時空間グリッドで表現する既存手法を改善し、柔軟な長さに適応できる粗から細への動画トークン化技術「VideoFlexTok」を発表した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ビジュアルトクナイザーは、高次元の生ピクセルを圧縮表現に変換して下流モデルに供給します。圧縮機能を超えて、トクナイザーはどのような情報が保持され、どのように整理されるかを決定します。ビデオトクナイズにおける事実上の標準的なアプローチは、ビデオをトクンの時空間 3 次元グリッドとして表現することです。各トクンは、元の信号に対応する局所的な情報を捉えます。これにより、テキストからビデオを生成するモデルなどの下流モデルは、ビデオの固有の複雑さに関わらず、「ピクセル単位」ですべての低レベルの詳細を予測することを学習する必要が生じ、その結果…
原文を表示
Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and how it is organized. A de facto standard approach to video tokenization is to represent a video as a spatiotemporal 3D grid of tokens, each capturing the corresponding local information in the original signal. This requires the downstream model that consumes the tokens, e.g., a text-to-video model, to learn to predict all low-level details “pixel-by-pixel” irrespective of the video’s inherent complexity, leading to…
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み