ComfyUI、Black Forest Labs の FLUX 3 Video をパートナーノード経由で利用可能に
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ComfyUI Blog
Black Forest Labs は FLUX 3 Video を ComfyUI のパートナーノード経由で利用可能にし、単一モデルによる動画生成とネイティブ音声合成が実現された。
AI深層分析を開く2026年8月6日 04:45
AI深層分析
キーポイント
統合型マルチモーダルアーキテクチャの採用
Black Forest Labs は動画、音声、画像、行動予測を別々のシステムでつなぐのではなく、単一アーキテクチャ内で同時に訓練されたモデルを公開した。
ネイティブ音声合成と多様なスタイル対応
FLUX 3 Video は最大 20 秒の生成が可能で、吹き替えではなくフレームと同時に音声を生成し、映画調からアニメまであらゆるスタイルに対応する。
ComfyUI での実装とワークフロー提供
同社は ComfyUI の最新バージョンまたは Comfy Cloud で利用可能とし、T2V や I2V などの具体的なワークフローをダウンロード用テンプレートとして公開した。
重要な引用
Black Forest Labs' first unified multimodal model — one model for video, audio, image, and action-prediction
Rather than separate systems stitched together at inference time, all four modalities are trained jointly in one architecture
編集コメントを表示
編集コメント
単一モデルによるマルチモーダル統合は、生成 AI の品質と効率性を飛躍的に高める重要なステップである。ComfyUI を利用する開発者は、この新機能を即座に試すことで、従来の手法との違いを実感できるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
FLUX 3 Video が ComfyUI の Partner Nodes を通じて利用可能になりました。本日リリースされるのは動画生成機能です。オープンウェイト版は近日公開予定です。
Black Forest Labs が開発した初の統一型マルチモーダルモデル「FLUX 3」は、映像・音声・画像・行動予測のすべてを一つのモデルで処理します。このモデルは、世界がどのように見え、動き、音を立て、そしてどう相互作用するかという深い理解に基づいて構築されています。
その背景には、独自のトレーニング手法があります。推論時に別々のシステムをつなぎ合わせるのではなく、4 つのモダリティ(映像・音声・画像・行動予測)を単一のアーキテクチャで同時に学習させています。これにより各モダリティが互いに強化し合い、真の意味での汎用モデルが実現しました。その結果、より現実に忠実な創作が可能になり、映画館のような特定のスタイルに限定されない多様な表現や、深いカスタマイズ性が得られます。
モデルの主な特徴
FLUX 3 Video は、1 回の生成で最大 20 秒の動画を制作できます。音声もフレーム生成と同時にネイティブに作成されるため、後から吹き替えする必要はありません。セリフ、効果音、環境音を一度に生成でき、必要に応じてサイレントプレート(無音版)を出力して後から音楽や効果音を付け足すことも可能です。
スタイルの幅広さがこのモデルの大きな特徴です。多くの動画生成モデルは、どんなプロンプトを与えても映画のような映像美に収束しがちですが、FLUX 3 ではハイリアリスティックな描写からアニメーション、そしてシネマティックな表現まで、あらゆるスタイルを自由に実現できます。
具体的な活用シーン:
- テキストから動画へ:シンプルなプロンプトから複雑な指示まで、現実世界の知識に基づいた詳細な生成が可能。
- 画像から動画へ:開始フレーム、終了フレーム、あるいは時系列のキーフレーム群を指定して動画を生成。
- 動画の継続:既存のクリップに続き、動きの勢い、構図、シーンの論理を一貫性を持って拡張。
複数のシーン — 1 つの生成内で複数のシーンやカメラアングルを扱えます。
多言語ダイアログ — 言語やアクセントに正確に対応し、リップシンクも維持します。
テキストとタイポグラフィ — シーンに自然に溶け込む文字表現で、単なる上書きではありません。
エージェント連携 — ComfyUI でクリップを生成・連鎖させ、より長く一貫性のあるストーリーを作成できます。
同じアーキテクチャは動画にとどまりません。画像生成、高速バリアント、オープンウェイトも近日公開予定です。
始め方
ComfyUI を最新バージョンに更新するか、Comfy Cloud を開いてください。
以下のワークフローをダウンロードするか、テンプレートライブラリから探してください。
FLUX 3 I2V ワークフローのダウンロード
FLUX 3 T2V ワークフローのダウンロード
FLUX 3 ストーリーボード ワークフローのダウンロード
プロンプトを入力し、画像をアップロードして再生時間を設定したら実行します。
FLUX 3 は、ComfyUI パートナーノードを通じて利用可能なモデルリストに新たに加わりました。いつものように、創作を楽しんでください!
原文を表示
FLUX 3 Video is now available in ComfyUI through Partner Nodes. Video generation is what ships today. Open weights are coming soon.
Black Forest Labs’ first unified multimodal model — one model for video, audio, image, and action-prediction, built on a richer understanding of how the world looks, moves, sounds, and how to interact with it.
The training is the reason. Rather than separate systems stitched together at inference time, all four modalities are trained jointly in one architecture, and each modality sharpens the others. What comes out is a true general model: creations that are truer to life, stylistically diverse well beyond cinematic, and deeply customizable.
Model Highlights
FLUX 3 Video generates up to 20 seconds in a single generation, with audio produced natively alongside the frames rather than dubbed on afterward — dialogue, sound effects, and ambience in the same pass, and optional when you want silent plates to score yourself. Stylistic range is the differentiator. Most video models converge on a cinematic aesthetic regardless of what you ask for; FLUX 3 makes any style possible, from hyper-realistic to animation to cinematic.
What that covers in practice:
Text to video — Simple prompts or complex ones, with real world knowledge behind the details.
Image to video — Drive from a start frame, an end frame, or a sequence of keyframes in order.
Video continuation — Extend an existing clip while carrying momentum, framing, and scene logic forward.
Multiple scenes — Several scenes and camera angles inside a single generation.
Multilingual dialogue — Accurate speech across languages and accents, with lipsync that holds.
Text and typography — Type that sits naturally in the scene rather than pasted on top.
Agentic chaining — Generate and Chain clips together in ComfyUI for a longer, coherent story.
The same architecture reaches past video. Image generation, fast variants, and open weights are on the way.
Getting Started
Update ComfyUI to the latest version, or open Comfy Cloud.
Download the workflows below, or find them in the template library.
Download FLUX 3 I2V Workflow
Download FLUX 3 T2V Workflow
Download FLUX 3 Storyboard Workflow
Write your prompt, upload an image, set duration, and run.
FLUX 3 joins the growing roster of models available through ComfyUI Partner Nodes.
As always, enjoy creating!
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み