動画拡散潜在変数からの三角形スプラット生成(5 分読了)
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
研究チームが、動画拡散モデルの潜在表現から三角形スプラットを直接生成する手法を発表し、3D 再構築の効率化を実現した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
幾何学的に正確なシーン生成のための順伝播潜在三角形スプッティング。
動画拡散潜在から明示的な表面アライメントされた三角形スプッティングを、単一の順伝播パスで復号する。
Orest Kupyn1,2, Goutam Bhat1, Philipp Henzler1, Fabian Manhardt1, Christian Rupprecht1,2, Federico Tombari1,3
1 Google Research
2 University of Oxford, Visual Geometry Group
3 Technical University of Munich
FLAT は、圧縮された動画拡散潜在を明示的な非体積的シーンパラメータに直接マッピングできることを示している。3D ガウス(Gaussian)を復号するのではなく、三角形スプッティングをワンパスで予測することで、幾何学的精度を向上させつつ競争力のある視覚品質を維持し、軽量なリファインメント後に単純な三角形レンダラーによるラスタライゼーションや物理ベースのインタラクションを可能にする。
直接三角形復号
FLAT は、多くの順伝播シーンパイプラインで一般的に用いられる「生成後最適化」という経路を避け、圧縮された動画拡散潜在を明示的な三角形スプッティングへ直接変換する。
幾何学特化型トレーニング
レイ中心の三角形パラメータ化とプロダクトウィンドウレンダリング関数が、三角形回帰を安定化させる。これにより、小さな方向誤差が勾配フローを破綻させることを防ぐ。
不透明アセットへのリファインメント
軽量なテスト時リファインメントステップにより、予測された三角形の集合体が、標準的なレンダリングやゲームエンジン風のインタラクションに適した完全な不透明表現に変換されます。
生成されたシーンを明示的な三角形ジオメトリとして検査する。
FLAT は、シンプルな三角形レンダラーですぐに探索できるシーンを出力します。これにより、重厚なレンダリングエンジンへの依存なく、ビューアーは高速かつあらゆるデバイスでポータブルになります。タッチ対応デバイスでは、シーン内でドラッグして周囲を見渡したり、画面上の移動ボタンを使用してナビゲートしたりできます。
White Room
Loading White Room...
ナビゲーション
W A S D で移動、ドラッグで視点移動、R でリセット。
ヒント
ビューポート内のどこでもダブルクリックすると、デフォルトの視点にスナップします。
タッチによる移動
外観と表面構造は整合性を保ちます。
私たちが目指すのは、画像のリアリズムだけでなく幾何学的な精度です。これらの対になったレンダリングは、FLAT の新規視点と表面法線が視点間を通じて一貫性を保ち、幾何学信号を外観だけで隠すのではなく明瞭にしていることを示しています。
新規視点
表面法線
image
image
01 / 07
01 / 07
BibTeX
@misc{kupyn2026flat,
title = {FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation},
author = {Orest Kupyn and Goutam Bhat and Philipp Henzler and Fabian Manhardt and Christian Rupprecht and Federico Tombari},
year = {2026},
note = {Preprint}
}
原文を表示
Feedforward Latent Triangle Splatting for geometrically accurate scene generation.
Decode explicit surface-aligned triangle splats from video diffusion latents in a single forward pass.
Orest Kupyn1,2, Goutam Bhat1, Philipp Henzler1, Fabian Manhardt1, Christian Rupprecht1,2, Federico Tombari1,3
1 Google Research
2 University of Oxford, Visual Geometry Group
3 Technical University of Munich
FLAT shows that compressed video diffusion latents can be mapped directly to explicit non-volumetric scene parameters. Instead of decoding 3D Gaussians, it predicts triangle splats in one pass, improving geometric accuracy while preserving competitive visual quality and enabling rasterization with simple triangle renderers and physics-based interaction after lightweight refinement.
Direct Triangle Decoding
FLAT turns compressed video diffusion latents into explicit triangle splats directly, avoiding the usual generate-then-optimize path used by many feedforward scene pipelines.
Geometry-Specific Training
Ray-centered triangle parameterization and a product window rendering function stabilize triangle regression, where small orientation errors would otherwise break gradient flow.
Refinement to Opaque Assets
A lightweight test-time refinement step converts the predicted triangle soup into a fully opaque representation that fits standard rendering and game-engine-style interaction.
Inspect generated scenes as explicit triangle geometry.
FLAT outputs scenes that can be explored immediately with a simple triangle renderer. This makes the viewer fast and portable across devices, without depending on a heavy rendering engine. On touch devices, drag inside the scene to look around and use the on-screen movement buttons to navigate.
White Room
Loading White Room...
Navigation
W A S D move, drag to look, R to reset.
Tip
Double-click anywhere in the viewport to snap back to the default view.
Touch Movement
Appearance and surface structure stay aligned.
We target geometric accuracy, not only image realism. These paired renders show that FLAT's novel views and surface normals stay consistent across viewpoints, making the geometry signal legible instead of hiding it behind appearance alone.
Novel View
Surface Normals


01 / 07
01 / 07
BibTeX
@misc{kupyn2026flat,
title = {FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation},
author = {Orest Kupyn and Goutam Bhat and Philipp Henzler and Fabian Manhardt and Christian Rupprecht and Federico Tombari},
year = {2026},
note = {Preprint}
}関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み