Hunyuan3D-Buffalo 1.0: 統合型マルチモーダルモデルで 3D 生成・編集を支援
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Hugging Face Daily Papers
腾讯は画像生成の進展を踏まえ、3D モデリングのデータ不足という課題に対処するため、理解・生成・編集を統合した「Hunyuan3D-Buffalo 1.0」を発表し、87M スケールの独自コーパスを用いて最先端性能を実現した。
AI深層分析を開く2026年8月5日 14:31
AI深層分析
キーポイント
統合フレームワークの提案
同社は 3D の理解、テキストから 3D への生成、指示に基づく編集、およびテキストに根ざした部分生成を単一アーキテクチャで実現する「Hunyuan3D-Buffalo 1.0」を発表した。
大規模コーパスの構築
スケーラブルな学習を可能にするため、Nano3D-v2 を用いて生成された 87M スケールの 3D マルチモーダルコーパス(理解サンプル 2500 万組、テキスト -3D ペア 5000 万組、編集ペア 1200 万組)を構築した。
ハイブリッドアーキテクチャ
同フレームワークは、意味的・構造的・空間的理解を行う「Hunyuan3D-VLM」と、高忠実度 3D 合成を行う「Hunyuan3D DiT」を組み合わせることで機能している。
構造保存による編集
編集と部分生成においては、拡散プロセスにソースオブジェクト表現を条件付けすることで、全体の構造や未編集領域を保持する仕組みを採用した。
重要な引用
unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture
we construct an 87M-scale 3D multimodal corpus
Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks
編集コメントを表示
編集コメント
3D データの不足という根本的なボトルネックに、8700 万規模のデータセットで取り組んだ点は技術的に極めて意義深い。同社のプロジェクトページへのリンクも提供されており、実装や詳細な実験結果を確認する道が開かれている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
画像生成の近年の進展は、理解・生成・編集を統合した統一型マルチモーダルモデルの可能性を示しました。しかし、統一型の 3D モデリングはまだ、大規模で幾何学的に整合性のある編集データの不足など、限られたマルチモーダルデータによって制約されています。
この課題に対処するため、私たちは「Hunyuan3D-Buffalo 1.0」を提案します。これは、単一のアーキテクチャ内で 3D の理解、テキストから 3D への生成、指示に基づく 3D 編集、そしてテキストに紐づくパーツ生成をサポートする統一フレームワークです。
スケーラブルな学習を実現するために、私たちは 8700 万件規模の 3D マルチモーダルコーパスを構築しました。これには、Nano3D-v2 を用いて生成された 2500 万件の理解サンプル、5000 万件のテキストと 3D のペア、そして 1200 万件の編集ペアが含まれています。
アーキテクチャ的には、このフレームワークは意味的・構造的・空間的な理解を行う「Hunyuan3D-VLM」と、高忠実度の 3D 合成を実現する「Hunyuan3D DiT」を組み合わせています。VLM は生成のためのマルチモーダルな意味条件を提供し、編集やパーツ生成では、拡散プロセスにソースオブジェクトの表現を条件として付与することで、全体の構造と未編集領域を維持します。
広範な実験により、Hunyuan3D-Buffalo 1.0 がテキストから 3D への生成や 3D 編集ベンチマークにおいて最先端またはトップクラスの性能を発揮し、優れた理解能力とパーツ生成能力を示すことが確認されました。さらに分析では、生成と理解の両方が編集性能を向上させることが明らかになり、統一型 3D マルチモーダル学習の有効性が実証されています。
プロジェクトページ:https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
原文を表示
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み