動画生成器が汎用ビジョンモデルへ
TLDR AI は、動画生成技術が単なるコンテンツ作成ツールを超え、世界を理解・推論する汎用ビジョンモデルとしての役割を果たす可能性を分析している。
キーポイント
動画生成の二重機能性
動画を生成するプロセス自体が、物理法則や因果関係を学習し、世界モデルを構築する手段となっている。
推論能力への転換
生成された動画の質は、背後にあるモデルが現実世界の動きや物体の挙動をどれだけ正確に理解しているかの指標となる。
汎用ビジョンモデルへの進化
従来の静的画像認識から脱却し、時間軸を含む動的な理解を通じて、より高度な視覚推論能力を獲得する道筋を示す。
重要な引用
Video generators are becoming general-purpose vision models.
The ability to generate realistic video implies a deep understanding of the world.
影響分析・編集コメントを表示
影響分析
この記事は、動画生成技術の進歩が単なるエンターテインメントやコンテンツ制作の効率化を超え、AI が現実世界を理解するための基盤技術として再定義されるべきであることを示唆しています。今後の AI 開発において、生成能力と推論能力を切り離して考えるのではなく、統合された「世界モデル」として捉える視点の重要性が高まります。
編集コメント
動画生成技術が「世界を理解する能力」の指標として再評価されるという視点は、AI の次の進化段階を予見する重要な示唆です。生成と推論の境界が曖昧になる中で、モデルの真価は出力の美しさではなく、その背後にある物理理解の深さで測られるようになるでしょう。
Google DeepMind、トロント大学、ロンドン大学カレッジ、オックスフォード大学、MIT、ルンド大学の共同研究チームが発表した論文です。
※Google DeepMind 在籍時の研究成果
ECCV 2026
TL;DR
コンピュータビジョンを「特定のタスクに特化した時代」から、「汎用的な視覚知能」の時代へと進化させる試みです。これは、自然言語処理(NLP)が特定のタスクモデルから汎用的な言語知能へ進化した過程と似ています。
GenCeption は、事前学習済みの動画生成モデルを流用し、単一の統合された汎用フィードフォワード型ビジョンモデルへと再構築しました。このモデルはテキスト指示によって制御されながら、多様な視覚タスクにおいて最先端(SOTA)の性能を発揮します。さらに、驚異的な学習効率と興味深い創発的挙動も確認されています。
視覚生成事前学習
GenCeption は動画生成モデルを表現学習のための事前学習基盤として活用し、単一のアーキテクチャ内で多様なタスクに対する事後学習(ポストトレーニング)を実行します。
統合型かつフィードフォワード型
GenCeption は密な(dense)タスクと疎な(sparse)タスクの両方を処理可能です。従来の多段階の生成バックボーンを、単一ステップで動作するフィードフォワードモデルへと変換しました。
最先端性能と創発的挙動
GenCeption は多様な視覚タスクにおいて統合型モデルとして最先端の性能を誇り、驚異的な学習効率と興味深い創発的挙動を示します。
ブラウザは video タグに対応していません。

手法。 GenCeption は、動画生成拡散モデルを事前学習の基盤として活用し、大規模な空間時間的な世界の前提知識や、ネイティブなビジョン・言語アライメントを獲得します。その後、多様な知覚タスクに対応するために、主に合成データで微調整された順方向(フィードフォワード)モデルへと適応させるポストトレーニングを行います。GenCeption は複数の視覚タスクで高い性能を発揮し、興味深い「新興行動」を示すことで、シミュレーションから実世界へのシームレスな転送や、分布外オブジェクトカテゴリへの一般化を可能にします。
パラダイムシフト。 これは、特定のタスクに特化したコンピュータビジョンモデルから、自然言語処理(NLP)がタスク特化型モデルから汎用的な言語知能へと進化したのと同様に、完全に統合された一般主義的なビジョンモデルへの転換を意味しています。

テキスト指示に従って操作される GenCeption は、単一のタスク非依存アーキテクチャ内で、1 回の順方向パスで深さ推定、表面法線・カメラ姿勢推定、前景および表情参照によるオープンワールドセグメンテーション、2D/3D キーポイント予測など、幅広い知覚タスクをこなすことができます。
Results


SOTA 性能: GenCeption は、深さ推定(DepthAnything3)、4D リカバリー(D4RT)、視覚セマンティック理解(VGGT-Ω)、セグメンテーション(SAM3)、人間動作分析(Sapiens)、動画生成・編集(DAVID, Genmo, Lotus-2)といった個別タスクに特化した最先端モデルと同等か、それ以上の性能を発揮します。ここでいう「スペシャリスト」は各タスクごとに個別に訓練されたモデルを指し、「ジェネラリスト」は単一のモデルが複数のタスクを同時に学習したものを意味します。
データ効率: 微調整用のデータ量を同一条件で比較した場合、動画生成バックボーンは V-JEPA や VideoMAE V2 といった他の事前学習手法を上回ります。また、モデルサイズやデータ量の増加に伴って性能が向上するスケーリング特性も確認されており、D4RT や VGGT-Ω と同等の性能を達成するには、必要な訓練データを 7 倍から 500 分の 1 に減らすだけで十分です。
-->
ファイルは assets/ フォルダに配置するか、動画ファイルが大きい場合はホストされた URL を使用してください。
-->
Overview Video
概要動画 — 近日公開予定
-->
Architecture
Your browser does not support the video tag.
GenCeption のアーキテクチャ概要は、テキストから動画生成する拡散モデルを流用した、シンプルかつ強力な構成です。入力された動画と、出力したい内容を指定するテキストプロンプトを受け取ると、合成データを中心に学習させたこの統一モデルが、単一の順方向パスで多様な密・疎な知覚タスクを実行できます。
密なビジョンタスクは RGB 環境空間上で統合され、潜在空間内で効率的に教師あり学習が可能です。一方、疎なビジョンタスクは、拡散トランスフォーマー(DiT)への追加入力として学習可能なトークンを付加することで実現されます。
Video Any-Task Capabilities
GenCeption は、人間が現実世界で行うように、異なるビジョンタスクをシームレスに切り替えることができます。
Your browser does not support the video tag.
Native Vision-Language Alignment
自然言語の記述を与えると、GenCeption は色や空間関係、動きについて推論し、参照される対象を正確にセグメント化します。また、テキストから動画生成する事前学習で得た知識を活用することで、ロケットなど未見の対象にも一般化して対応できます。
Your browser does not support the video tag.
Understanding the Complex Visual World
複雑な視覚世界を理解する境界を押し広げる例として、GenCeption は困難な条件下(複雑な動作、遮蔽、自己視点、マルチビューなど)において 4D の人体キーポイントを予測します。これはロボット工学や AR/VR における実世界のアプリケーションを支える、堅牢な 4D 人体理解の基盤となっています。
Your browser does not support the video tag.
Grounded 4D Reconstruction( grounded な 4D 再構築)
単一の動画から、GenCeption はピクセルごとの幾何学情報とカメラ姿勢を予測し、シーンを 4D の点群へと昇華させます。これにより、自由視点でのフライスルーや、テキスト指示に基づいた物体の言語的 grounding が可能になります。
Your browser does not support the video tag.
Emergent Behaviors(創発する振る舞い)
GenCeption は主に合成された人間の動画で微調整されていますが、シミュレーションから実写映像へ、さらに分布外のカテゴリへとシームレスに転移します。これは生成モデルのバックボーン内部に、普遍的な「世界モデル」が存在することを示す証拠です。
- 純粋に合成動画のみで訓練されたこのモデルは、ゼロショットで実世界の映像にも転移可能であり、シミュレーションから実世界への完璧な転送を実現します。
- 単一物体の合成動画で訓練されたモデルは、複数のインスタンスが含まれる実写動画に対してもゼロショットで一般化できます。
人間データのみで訓練されたこのモデルは、未知のカテゴリである動物やロボットにも汎化します。
謝辞
ここに記載する謝辞は網羅的なものではありません。個別に名を挙げることはできませんが、本プロジェクトの支援、フィードバック、議論を通じて貢献してくださった多くの方々にも心から感謝いたします。
このプロジェクト全体を通じて貴重なサポートを提供してくれた Jeremiah Harmsen, Joëlle Barral, Chris Dyer, Rahul Sukthankar, Junlin Zhang, Miki Rubinstein, Thabo Beeler, Di Qiu, Jesús Pérez, Alberto García García, Sergio Orts Escolano, Erroll Wood, Iker J. de los Mozos, Emily Conn の各氏に深く感謝いたします。
また、研究に関する洞察深い議論や技術的なフィードバック、実装の詳細についての有益な対話に貢献してくれた以下の皆様(姓の五十音順)にもお礼申し上げます。Thiemo Alldieck, Yanrui Bin, Ang Cao, Minghao Chen, Carlton Chu, Dima Damen, Shlomi Fruchter, Valentin Gabeur, Robert Geirhos, Tengda Han, Zeren Jiang, Linyi Jin, Ruofan Liang, Ziwei Liao, Haotong Lin, Xingchao Liu, Yibo Liu, Shangbang Long, Viorica Patraucean, Cordelia Schmid, Hao Shao, Jianyuan Wang, Luyu Wang, Thaddäus Wiedemer, Ziyi Wu, Yuxi Xiao, Junyu Xie, Wei Yu, Shangzhan Zhang, Shuhong Zheng。
本プロジェクト完了後の著者たちの個人的な考えについては、以下のブログもご覧ください。
引用
@inproceedings{wang2026genception,
title = {Video Generation Models are General-Purpose Vision Learners},
author = {Wang, Letian and Zhang, Chuhan and Kabra, Rishabh and Uijlings, Jasper and
Waslander, Steven and Zisserman, Andrew and Carreira, Joao and He, Kaiming and
Andriluka, Misha and Bazavan, Eduard Gabriel and Zanfir, Andrei and Sminchisescu, Cristian},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}
「動画生成モデルは汎用的な視覚学習器である」という論文のプロジェクトページです。
原文を表示
1Google DeepMind 2University of Toronto 3University College London 4University of Oxford 5MIT 6Lund University
**
*Work done while at Google DeepMind.
ECCV 2026
TL;DR
Pushing computer vision from the task-specific era toward general-purpose visual intelligence — just as NLP evolved from task-specific models into general-purpose language intelligence.
GenCeption repurposes a pre-trained video generative model into a single unified, general-purpose, feed-forward vision model that solves a wide range of vision tasks with SOTA performance — all steered by text instructions, with exceptional learning efficiency and intriguing emergent behaviors.
Visual Generative Pretraining
GenCeption leverages video generation models as representation pretraining, and conducts multi-task post-training in a unified architecture.
Unified & Feed-Forward
GenCeption handles both dense and sparse vision tasks, transforming the multi-step generative backbone into a single-step feed-forward model.
SOTA & Emergent
GenCeption is a SOTA unified model on various vision tasks, with exceptional learning efficiency and intriguing emerging behaviors.
Methodology
Your browser does not support the video tag.

Methodology.** GenCeption leverages a video generative diffusion model as a *pre-training* base to capture rich spatio-temporal world priors and native vision-language alignment at scale. During multi-task *post-training*, the model is adapted to a feed-forward model fine-tuned on predominantly synthetic data to handle diverse perception tasks. GenCeption shows strong performance on multiple vision tasks with intriguing *Emerging Behaviors*, enabling seamless sim-to-real transfer and generalization to out-of-distribution object categories.
Paradigm Shift. This highlights a paradigm shift from specialized, task-specific computer vision models toward fully unified, generalist vision models — just as NLP evolved from task-specific models into general-purpose language intelligence.
Results


SOTA performance: GenCeption is competitive with or outperforms state-of-the-art models dedicated to individual tasks (DepthAnything3, D4RT, VGGT-Ω, SAM3, Sapiens, DAVID, Genmo, Lotus-2). Our *specialist* denotes a model trained on each task individually, whereas the *generalist* represents a single model trained jointly across multiple tasks.
Data efficiency: under matched fine-tuning data, the video-generative backbone beats alternative pre-training paradigms (V-JEPA, VideoMAE V2), shows preliminary scaling properties, where it improves with larger models and more data, and reaches performance comparable to D4RT and VGGT-Ω with 7× to 500× less training data.
Architecture
Your browser does not support the video tag.
Architecture overview of GenCeption, a simple yet powerful architecture adapted from text-to-video diffusion models. Given an input video and a text prompt specifying the desired output, our unified model, trained majorly on synthetic data, is capable of performing a wide range of dense and sparse perception tasks, with a single forward-pass of the model. The dense vision tasks are unified in the RGB ambient space where supervision can be applied in latent space efficiently, and the sparse vision tasks are realized by adding learnable tokens as additional inputs to the diffusion transformer (DiT).
Video Any-Task Capabilities
GenCeption is able to seamlessly switch between different vision tasks, just like how humans do in real life.
Your browser does not support the video tag.
Native Vision-Language Alignment
Given a natural-language expression, GenCeption accurately segments the referred object — reasoning about colors, spatial relationships and motion — and generalizes to unseen objects (e.g., a rocket) by leveraging knowledge from text-to-video pre-training.
Your browser does not support the video tag.
Understanding the Complex Visual World
As an example of pushing the boundaries of understanding the complex visual world, GenCeption predicts 4D human keypoints under challenging conditions (complex motion, occlusion, ego-centric, multi-view, etc.) — robust 4D human understanding that underpins real-world applications across robotics and AR/VR.
Your browser does not support the video tag.
Grounded 4D Reconstruction
From a single video, GenCeption predicts per-pixel geometry and camera pose that lift the scene into a 4D point cloud — enabling free-viewpoint fly-throughs and language grounding of objects from a text instruction.
Your browser does not support the video tag.
Emergent Behaviors
Although fine-tuned predominantly on synthetic human videos, GenCeption transfers seamlessly from simulation to real footage and to out-of-distribution categories — evidence of a universal "world model" inside generative video backbones.
Trained purely on synthetic videos, the model transfers zero-shot to real-world footage — seamless sim-to-real transfer.
Trained on synthetic single-object videos, the model generalizes zero-shot to real videos with multiple instances.
Trained only on humans, it generalizes to unseen categories — animals and robots.
Acknowledgements
The following acknowledgements are not exhaustive. We are also grateful to many others whose support, feedback, and discussions contributed to this work but who are not individually acknowledged here.
We are deeply grateful to Jeremiah Harmsen, Joëlle Barral, Chris Dyer, Rahul Sukthankar, Junlin Zhang, Miki Rubinstein, Thabo Beeler, Di Qiu, Jesús Pérez, Alberto García García, Sergio Orts Escolano, Erroll Wood, Iker J. de los Mozos, and Emily Conn for their invaluable support throughout this project.
We also thank the following individuals (listed alphabetically by last name) for insightful research discussions, technical feedback, and valuable conversations on implementation details: Thiemo Alldieck, Yanrui Bin, Ang Cao, Minghao Chen, Carlton Chu, Dima Damen, Shlomi Fruchter, Valentin Gabeur, Robert Geirhos, Tengda Han, Zeren Jiang, Linyi Jin, Ruofan Liang, Ziwei Liao, Haotong Lin, Xingchao Liu, Yibo Liu, Shangbang Long, Viorica Patraucean, Cordelia Schmid, Hao Shao, Jianyuan Wang, Luyu Wang, Thaddäus Wiedemer, Ziyi Wu, Yuxi Xiao, Junyu Xie, Wei Yu, Shangzhan Zhang, and Shuhong Zheng.
Also see the authors' personal thoughts after the project.
Citation
@inproceedings{wang2026genception,
title = {Video Generation Models are General-Purpose Vision Learners},
author = {Wang, Letian and Zhang, Chuhan and Kabra, Rishabh and Uijlings, Jasper and
Waslander, Steven and Zisserman, Andrew and Carreira, Joao and He, Kaiming and
Andriluka, Misha and Bazavan, Eduard Gabriel and Zanfir, Andrei and Sminchisescu, Cristian},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}Project page for “Video Generation Models are General-Purpose Vision Learners.”
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み