ネイティブ多モーダルモデル(GitHub リポジトリ)
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
GitHub で公開されたリポジトリが、画像や音声など複数のデータタイプを同時に処理するネイティブ型多モーダルモデルを紹介している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
**
このリポジトリは、構造的な移行を体系的に追跡しています。それは、生きた感覚信号に対する根本的な盲目性を抱える後期融合/接ぎ木構成である「モジュラーアセンブリ」から、複数のモダリティが *統一されたトランスフォーマー空間* または *共通バックボーン* に本質的に統合される「ネイティブ多モーダルモデリング (NMM)」への移行です。
⭐ このリポジトリをスターして、最新の画期的な作品を追跡してください。見落としがある可能性のあるモデルについては、PR を歓迎します。
🗺️ NMM 建築分類法
私たちは、統合深度(中期融合 vs. 早期融合)と機能的入出力の二重性に基づく二軸レンズを通じて、NMM エコシステムを形式化しています。
#
パラダイム
入力 → 出力
核となるアイデア
🟦
M2T — マルチからテキストへ
多モーダル → テキスト
推論のために、クロスモーダル入力を純粋に言語的な応答に grounded します。
🟩
M2G — マルチからターゲットへ
多モーダル → モダリティ固有
ネイティブ表現を通じてモダリティ固有の出力を直接合成し、時間的・音響的一貫性を達成します。
🟪
M2M — マルチからマルチへ
多モーダル → 多モーダル
理解と生成が単一のネットワーク内で相互的な射影として自然に共存する統一されたパラダイムです。
🟦 1. マルチからテキストへ (M2T) ユニモーダル生成
**
*論理的推論のためにクロスモーダル入力を言語ストリームに grounded するネイティブスケーリングフレームワーク。*
🧱 後期融合ベースライン参照
*浅いプロジェクタを介してモジュラー的に組み立てられ、生きた感覚信号に対して盲目です。
**
- LLaVA [Liu et al., 2023] — 💻 GitHub · 📄 Paper
- DeepSeek-VL [Lu et al., 2024] — 💻 GitHub · 📄 Paper
- Qwen-Image [Wu et al., 2025] — 💻 GitHub · 🌐 Blog
🔗 Mid-Fusion (Naturally Interacted Regime)
*Foundational pioneers maintaining explicit, modality-aware boundaries.*
- CogVLM [Wang et al., 2023] — 💻 GitHub
- Qwen-Audio [Chu et al., 2023] — 💻 GitHub · 🌐 Project Page
*Massive state-of-the-art evolved mid-fusion architectures:*
- Qwen2.5-VL [Qwen Team, 2025] — 💻 GitHub · 🌐 Blog
- Qwen3-VL [Qwen Team, 2025] — 💻 GitHub · 📄 Paper
- InternVL-3.5 [Chen et al., 2025] — 💻 GitHub · 🤗 HF Collection
*Scale-driven industrial mid-fusion implementations:*
- GLM-4.5V / GLM-V [ZhipuAI, 2025–2026] — 💻 GitHub · 🤗 HF Model
- Kimi K2 / K2.5 [Moonshot AI, 2025–2026] — 🌐 Project Page · 💻 GitHub Org
🧬 Dense / Native M2T Scaling
- MiniCPM-V 4.x [Yu et al., 2025] — 💻 GitHub
- Nemotron 3 Nano Omni [NVIDIA, 2026] — 💻 GitHub · 📄 Paper
- MiMo-V2.5 [Xiaomi MiMo Team, 2026] — 💻 GitHub · 🌐 Project Page
- Gemma-4 / Qwen3.6 — Timeline benchmarks driving advanced contextual reasoning (forthcoming).
🟩 2. Multi-to-Target (M2G) Scenario-based Generation
*Bypassing traditional post-hoc decoders to synthesize photorealistic spatiotemporal physics or continuous speech directly.*
🎬 Advanced Video / World Simulators
- Wan 2.2-T2V-A14B [Wan Team, 2025] — 🤗 HF Model — Unifies video patches into native generation spaces with continuous physics.
- HunyuanVideo & HunyuanVideo-1.5 [Tencent, 2024–2025] — 💻 GitHub · 🤗 HF Model (1.5)
- Kling-Omni [Kuaishou, 2025] — 🌐 Project Page
🎙️ Speech-Centric Native Frameworks
- OmniVoice [Zhu et al., 2026] — 💻 GitHub · 🌐 Project Page
- MiniCPM-o 2.6 / 4.5 [OpenBMB, 2025–2026] — 🤗 HF Model · 💻 GitHub
- Seedream 3.0 [Gao et al., 2025] — 📄 Tech Report · 🌐 Project Page
- HiDream-I1 — 💻 GitHub
📅 Timeline Milestone Generators
- LTX-2 / LTX-Video [Lightricks, 2024–2026] — 💻 GitHub
- Ming-Flash-Omni [Ant Group / inclusionAI, 2025] — 💻 GitHub · 📄 Paper
🟪 3. Multi-to-Multi (M2M) Symmetric Modeling
*Omni-directional unified spaces establishing a symmetric paradigm where comprehension and generation natively coexist.*
🔥 Early-Fusion (Native Convergent Regime)
*Born-native designs treating all modalities equivalently via one unified backbone & embedding space.*
- Transfusion [Zhou et al., 2024] — 📄 Paper
- Chameleon ★ [Meta AI, 2024] — 💻 GitHub · 📄 Paper
- AnyGPT ★ [Zhan et al., 2024] — 💻 GitHub · 📄 Paper
🔮 Early Unified Predictors
- Moshi ★ [Défossez et al., 2024] — 💻 GitHub · 📄 Paper — Real-time conversational audio-text dual-stream processing.
- Emu3 / Emu3.5 ★ [BAAI, 2024–2025] — 🌐 Project Page · 📄 Paper — Next-token sequence prediction unifying understanding and synthesis.
🧩 Interleaved Sequence Modeling
- BAGEL-7B [ByteDance Seed Team, 2025] — 🤗 HF Model · 🌐 Project Page · 📄 Paper
- OneCAT-3B [Meituan & SJTU, 2025] — 💻 GitHub · 🤗 HF Model
- Show-o2-7B [Xie et al., 2025] — 💻 GitHub · 📄 Paper
🌌 Bidirectional Unification Frontiers
*Collapsing representation boundaries.*
- Janus-Pro ★ [DeepSeek-AI, 2025] — 🤗 HF Model · 📄 Paper
- Llama-4 Scout / Maverick [Meta AI, 2025] — 🌐 Llama Site — Advanced interleaved-scale exploration.
- LLaDA-V [Ml-GSAI, 2025] — 💻 GitHub · 🌐 Project Page · 📄 Paper
- Lance [ByteDance, 2026] — 📄 Paper — Leading edge of complete native convergence.
- TUNA-2 [Liu et al., 2026] / Mamoda 2.5 [Shi et al., 2026] / LongCat-Next — Forthcoming.
★ *Denotes early exploratory or foundational dual-regime architectures.*
🛠️ The Technical Roadmap Dimensions
Following the systemic structure detailed across Sections §3–§7** of the roadmap paper, the core components of the NMM lifecycle are curated below.
🧩 1. Architecture · §3
- Integration Depth Mapping — Structural mechanics of joint multimodal backbones vs. single unified transformer spaces.
- Input–Output Decoupling — Eliminating modality-aware boundaries & shallow projectors.
📊 2. Data Curriculum · §4
- Interleaved Data Curricula — Pre-training token mixtures combining web-scale text, audio waves, & video streams.
- Post-Training Engineering — Multi-modal instruction tuning & alignment token datasets.
🎯 3. Training Strategies · §5
- マルチ目的損失レシピ — 連続・離散目標の統合(自己回帰次期トークン予測+拡散ステップ)。
- スケーリングダイナミクス — 1T+ モデルの最前線における勾配安定性を維持するためのトークン割り当て戦略の計算。
⚡ 4. 推論・展開 · §6
- フルデュプレックスオーケストレーション — リアルタイムインタラクション(<100 ms)のための動的 KV キャッシュ退去とマルチスケールアテンションパターン。
- ハードウェアネイティブコンパイル — 統一されたクロスモーダルトークンルーティングのための分散 CUDA 計算カーネル。
🧪 5. 評価ベンチマーク · §7
- 対称的評価行列 — ターゲットモダリティの崩壊に陥ることなく、インターリーブされたマルチモーダルシーケンスを検査できるシステムをベンチマークするもの。
🤝 コントリビューション
コントリビューションは大歓迎です!注目すべきネイティブマルチモーダルモデルが欠落している場合や、古いリンクを見つけた場合は、Issue を開くかプルリクエストを送ってください。
推奨されるエントリー形式は以下の通りです:
- <モデル名> [<著者 / チーム>, <年>] — `💻 GitHub` · `📄 Paper`
✍️ 引用
当社の形式化、分類体系、またはロードマップフレームワークがあなたの研究に役立った場合は、以下の決定版論文を引用してください:
@article{TencentYoutuLab2026toward,
title = {Toward Native Multimodal Modeling: A Roadmap},
author = {Siyu An and Junru Lu and Junnan Dong and others},
journal = {arXiv preprint},
year = {2026}
}
NMM-Roadmap コミュニティによって維持されています · オープンなマルチモーダル研究のために ❤️ を込めて作成されました。
原文を表示
This repository systematically tracks the structural transition from Modular Assembly — late-fusion / grafted compositions that suffer from a fundamental blindness to raw sensory signals — to Native Multimodal Modeling (NMM), where multiple modalities are intrinsically integrated into a unified transformer space or joint backbone.
⭐ Star this repo to track the latest landmark works. PRs are warmly welcomed for any model we may have missed.
🗺️ The NMM Architectural Taxonomy
We formalize the NMM ecosystem through a dual-dimensional lens based on Integration Depth (mid-fusion vs. early-fusion) and Functional Input–Output Duality:
| # | Paradigm | Input → Output | Core Idea |
|---|---|---|---|
| 🟦 | M2T — Multi-to-Text | multimodal → text | Ground cross-modal inputs into purely linguistic responses for reasoning. |
| 🟩 | M2G — Multi-to-Target | multimodal → modality-specific | Direct synthesis of modality-specific outputs through native representations to achieve temporal & acoustic coherence. |
| 🟪 | M2M — Multi-to-Multi | multimodal → multimodal | A unified paradigm where understanding and generation naturally coexist as reciprocal projections within a single network. |
🟦 1. Multi-to-Text (M2T) Unimodal Generation
Native scaling frameworks that ground cross-modal inputs into linguistic streams for logical reasoning.
🧱 Late-Fusion Baseline References
*Modularly assembled via shallow projectors; blind to raw sensory signals.*
- LLaVA [Liu et al., 2023] — 💻 GitHub · 📄 Paper
- DeepSeek-VL [Lu et al., 2024] — 💻 GitHub · 📄 Paper
- Qwen-Image [Wu et al., 2025] — 💻 GitHub · 🌐 Blog
🔗 Mid-Fusion (Naturally Interacted Regime)
*Foundational pioneers maintaining explicit, modality-aware boundaries.*
- CogVLM [Wang et al., 2023] — 💻 GitHub
- Qwen-Audio [Chu et al., 2023] — 💻 GitHub · 🌐 Project Page
*Massive state-of-the-art evolved mid-fusion architectures:*
- Qwen2.5-VL [Qwen Team, 2025] — 💻 GitHub · 🌐 Blog
- Qwen3-VL [Qwen Team, 2025] — 💻 GitHub · 📄 Paper
- InternVL-3.5 [Chen et al., 2025] — 💻 GitHub · 🤗 HF Collection
*Scale-driven industrial mid-fusion implementations:*
- GLM-4.5V / GLM-V [ZhipuAI, 2025–2026] — 💻 GitHub · 🤗 HF Model
- Kimi K2 / K2.5 [Moonshot AI, 2025–2026] — 🌐 Project Page · 💻 GitHub Org
🧬 Dense / Native M2T Scaling
- MiniCPM-V 4.x [Yu et al., 2025] — 💻 GitHub
- Nemotron 3 Nano Omni [NVIDIA, 2026] — 💻 GitHub · 📄 Paper
- MiMo-V2.5 [Xiaomi MiMo Team, 2026] — 💻 GitHub · 🌐 Project Page
- Gemma-4 / Qwen3.6 — Timeline benchmarks driving advanced contextual reasoning (forthcoming).
🟩 2. Multi-to-Target (M2G) Scenario-based Generation
Bypassing traditional post-hoc decoders to synthesize photorealistic spatiotemporal physics or continuous speech directly.
🎬 Advanced Video / World Simulators
- Wan 2.2-T2V-A14B [Wan Team, 2025] — 🤗 HF Model — Unifies video patches into native generation spaces with continuous physics.
- HunyuanVideo & HunyuanVideo-1.5 [Tencent, 2024–2025] — 💻 GitHub · 🤗 HF Model (1.5)
- Kling-Omni [Kuaishou, 2025] — 🌐 Project Page
🎙️ Speech-Centric Native Frameworks
- OmniVoice [Zhu et al., 2026] — 💻 GitHub · 🌐 Project Page
- MiniCPM-o 2.6 / 4.5 [OpenBMB, 2025–2026] — 🤗 HF Model · 💻 GitHub
- Seedream 3.0 [Gao et al., 2025] — 📄 Tech Report · 🌐 Project Page
- HiDream-I1 — 💻 GitHub
📅 Timeline Milestone Generators
- LTX-2 / LTX-Video [Lightricks, 2024–2026] — 💻 GitHub
- Ming-Flash-Omni [Ant Group / inclusionAI, 2025] — 💻 GitHub · 📄 Paper
🟪 3. Multi-to-Multi (M2M) Symmetric Modeling
Omni-directional unified spaces establishing a symmetric paradigm where comprehension and generation natively coexist.
🔥 Early-Fusion (Native Convergent Regime)
*Born-native designs treating all modalities equivalently via one unified backbone & embedding space.*
- Transfusion [Zhou et al., 2024] — 📄 Paper
- Chameleon ★ [Meta AI, 2024] — 💻 GitHub · 📄 Paper
- AnyGPT ★ [Zhan et al., 2024] — 💻 GitHub · 📄 Paper
🔮 Early Unified Predictors
- Moshi ★ [Défossez et al., 2024] — 💻 GitHub · 📄 Paper — Real-time conversational audio-text dual-stream processing.
- Emu3 / Emu3.5 ★ [BAAI, 2024–2025] — 🌐 Project Page · 📄 Paper — Next-token sequence prediction unifying understanding and synthesis.
🧩 Interleaved Sequence Modeling
- BAGEL-7B [ByteDance Seed Team, 2025] — 🤗 HF Model · 🌐 Project Page · 📄 Paper
- OneCAT-3B [Meituan & SJTU, 2025] — 💻 GitHub · 🤗 HF Model
- Show-o2-7B [Xie et al., 2025] — 💻 GitHub · 📄 Paper
🌌 Bidirectional Unification Frontiers
*Collapsing representation boundaries.*
- Janus-Pro ★ [DeepSeek-AI, 2025] — 🤗 HF Model · 📄 Paper
- Llama-4 Scout / Maverick [Meta AI, 2025] — 🌐 Llama Site — Advanced interleaved-scale exploration.
- LLaDA-V [Ml-GSAI, 2025] — 💻 GitHub · 🌐 Project Page · 📄 Paper
- Lance [ByteDance, 2026] — 📄 Paper — Leading edge of complete native convergence.
- TUNA-2 [Liu et al., 2026] / Mamoda 2.5 [Shi et al., 2026] / LongCat-Next — Forthcoming.
★ Denotes early exploratory or foundational dual-regime architectures.
🛠️ The Technical Roadmap Dimensions
Following the systemic structure detailed across Sections §3–§7 of the roadmap paper, the core components of the NMM lifecycle are curated below.
| 🧩 1. Architecture · §3 Integration Depth Mapping — Structural mechanics of joint multimodal backbones vs. single unified transformer spaces. Input–Output Decoupling — Eliminating modality-aware boundaries & shallow projectors. | 📊 2. Data Curriculum · §4 Interleaved Data Curricula — Pre-training token mixtures combining web-scale text, audio waves, & video streams. Post-Training Engineering — Multi-modal instruction tuning & alignment token datasets. |
|---|
| 🎯 3. Training Strategies · §5 Multi-Objective Loss Recipes — Unifying continuous-discrete objectives (autoregressive next-token prediction + diffusion steps). Scaling Dynamics — Computed token-allocation strategies to maintain gradient stability at the 1T+ MoE frontier. | ⚡ 4. Inference & Deployment · §6 Full-Duplex Orchestration — Dynamic KV-cache eviction & multi-scale attention patterns for real-time interaction (
🤝 Contributing
Contributions are very welcome! If a notable native multimodal model is missing or you find an outdated link, please open an Issue or send a Pull Request.
The preferred entry format is:
- **<Model Name>** [<Authors / Team>, <Year>] — [`💻 GitHub`](https://...) · [`📄 Paper`](https://...)✍️ Citation
If our formalization, taxonomy, or roadmap framework assists your research, please cite our definitive paper:
@article{TencentYoutuLab2026toward,
title = {Toward Native Multimodal Modeling: A Roadmap},
author = {Siyu An and Junru Lu and Junnan Dong and others},
journal = {arXiv preprint},
year = {2026}
}Maintained by the NMM-Roadmap community · Made with ❤️ for open multimodal research.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み