Black Forest Labs、動画生成に挑戦するマルチモーダル基盤モデル「FLUX 3」を発表
本文の状態
日本語全文を表示中
詳細モードで約13分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Replicate
Black Forest Labs は画像、音声、動画を統合学習する新マルチモーダル基盤モデル「FLUX 3」を発表し、物理法則や因果関係を内包した高品質な生成を実現すると発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月10日 16:07
AI深層分析
キーポイント
統合アーキテクチャによる物理法則の学習
画像、音声、動画を同時に学習する新アプローチにより、音と動作の整合性や質量への従順さなど、物理法則をモデル内部に組み込んでいる。
テキストから動画への変換能力
単純な文章から詳細な指示まで柔軟に対応し、エンジン音や砂の質感など具体的な動作原理を理解したリアルな動画生成が可能である。
マルチモーダル間の相互制約
異なるモダリティが互いに制約となり合うことで、単一の視点では見えない現実の因果関係や時間的ダイナミクスをより正確に捉える。
プロンプトの簡潔化が可能
FLUX 3 は詳細なキーワードや過剰な記述を必要とせず、モデルに動きや美観の詳細を任せることで満足できる結果が得られる。
複数のカットデフォルト動作
特に指定がない限り、FLUX 3 は複数のシーンカットを生成する傾向がある。
重要な引用
Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws.
The modalities stop being separate and start being evidence about one underlying reality.
I like that FLUX 3 doesn't necessarily need your prompts to be dense or overwrought with details or keywords. Sometimes, you can offload the details of motion and aesthetics to the model itself, and you'll probably end up with a satisfying result.
You'll notice that FLUX 3 will default to multiple cuts of scenes unless otherwise specified.
編集コメントを表示
編集コメント
物理法則を学習に組み込むというアプローチは、生成 AI の信頼性を高める上で極めて重要な方向性である。特に動画生成において時間的整合性と物理挙動の再現が課題となる中、この技術的転換は業界全体に影響を与える可能性がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

Black Forest Labs は、画像生成モデル「FLUX.1」をリリースしたことで歴史に名を残しました。同社は、高速かつ高品質なオープンソースの画像生成を実現する先駆的な独立系研究機関の一つです。そして今、その挑戦は動画へと向けられ、私たちは大きな期待を抱いています。
FLUX 3 は、同ラボが新たに開発したマルチモーダル基盤モデルです。他の研究機関とは異なる独自のアプローチを採用しており、画像・音声・動画を学習対象として統合されたアーキテクチャで構築されています。これにより、多様な出力を生成することが可能になりました。リリースブログからの抜粋をご紹介します。
「画像は、特定の瞬間における空間構造と関係性を捉えます。動画には時間の次元が加わり、時間的な動態や物理法則が明らかになります。音声からは、視覚だけでは検出できない機械現象と音響の間の因果関係が読み取れます。言語はこれらの知覚を目標、抽象概念、指示へと結びつけます。一つのモダリティから学べば、その投影に関する良いモデルを得られます。しかし、それらを同時に学習すれば、相互の制約がさらに多くのことを教えてくれます。音は衝撃と一致し、運動は質量に従い、未来は過去から導かれるべきです。こうして各モダリティは独立したものではなく、一つの根本的な現実を裏付ける証拠へと変わっていくのです。」
このトレーニングアプローチから推測すると、このモデルには物理法則が深く組み込まれているようです。音は動きを、動きは音を補完します。動画は空間関係をすでに符号化している画像に時間軸の要素を加えます。FLUX 3 は、現実をよりよく捉える重みのネットワークを作るための試みです。
この投稿では、FLUX 3 が特に得意とする分野を実例とともに解説します。
テキストから動画へ
FLUX 3 を動かすために必要なものは多くありません。単純な一文でも、詳細に記述された文章でも、どちらでも問題なく処理できます。重要なのは、単に見た目を真似るだけでなく、仕組みそのものを理解しているような描写にも対応できる点です。
砂丘を全速力で走るオフロードトロフィートラックが、背後に巨大な砂煙を上げながら疾走する。山頂を越える瞬間にサスペンションが跳ね、一瞬浮き上がる様子が捉えられている。カメラは地面すれすれの低アングルで、高速で追従している。超リアルな 8K 画質。真昼の厳しい砂漠の日差し、エンジンの轟音、タイヤが砂を踏み砕く音が聞こえるようだ。
Cinematic underwater wildlife video of a large octopus slowly crawling across a dark rocky seafloor. Low, close tracking angle at eye level, showing its textured mantle, expressive eyes, and eight arms moving independently across wet stone. Suction cups attach and release naturally as the arms pull the body forward; subtle skin ripples and realistic muscular motion. Deep teal water, drifting particles, soft shafts of filtered light, muted rust-red and brown octopus coloration, highly detailed photorealistic skin texture, shallow depth of field, natural documentary cinematography.
I like that FLUX 3 doesn't necessarily need your prompts to be dense or overwrought with details or keywords. Sometimes, you can offload the details of motion and aesthetics to the model itself, and you'll probably end up with a satisfying result.
A young man staring out the window of a train rumbling along in rural Switzerland, beautiful mountains in the background whooshing by. Handheld camera feel, film shot.
A diver descends through crystal-clear turquoise water into a jungle cenote, sunbeams piercing down from the opening far above. Beneath the surface, the walls of a submerged Mayan temple come into view, ancient glyphs carved into the stone, fallen pillars scattered across the floor. Fish dart through the ruins as the diver's flashlight beam sweeps across the carvings. Hyper-realistic, 8k, National Geographic underwater cinematography, muffled bubbles and the diver's steady breathing.
FLUX 3 では、特に指定がない限り、シーンのカット割りを複数生成するデフォルト動作になります。
雨に濡れた地下駐車場で夜行される高速度のバイク追跡シーン。ライダーはタイトなカーブで大きく体を傾け、ヘッドライトがコンクリートの柱を横切り点滅します。タイヤは濡れた舗装面で唸りを上げ、車体が支持梁に接触して火花を散らしながらバランスを取り直します。カメラアングルは低めの追跡ショットとハンドルバーの視点(POV)の間で切り替わります。超リアルな 8K 画質、ネオンに照らされた水たまりがヘッドライトを反射し、エンジン音とタイヤの音が響き渡ります。
コスタリカの熱帯雨林の茂みを歩くハイカー。緑色の葉の層を突き抜ける金色の光の筋。上空でケツァールが飛び立ち、枝の上には動きのないナマケモノがぶら下がっています。森の地面からは霧が立ち上ります。カメラは背後から一定の歩行ペースで追跡し、やがて広大な樹冠を捉えるワイドショットへと上昇します。超リアルな 8K 画質、豊かな生物多様性、遠くで鳴るオオカミザルの声と足元の葉擦れの音が聞こえます。
Image to video
FLUX 3 は、変換をガイドするために開始フレームと終了フレームを使用できます。例えば、同じ車とシーンを使いながら、状態が「廃車」から「走行可能」へと変化する場合などが有用なテストケースです。
開始フレームと終了フレーム。 2 つの画像を与え、その間に起こしてほしい物理的な変化を記述してください。
Start

End

同じ錆びたジャンクヤードの車が、修復された赤い車へと自然に変化する単一の連続ショット。車両のアイデンティティ、カメラアングル、ホイールベース、ボディ形状、背景は終始完全に統一してください。錆の欠片がきれいな赤いペイントに溶け込み、凹みがゆっくりと平滑化し、割れたフロントガラスが透明になり、外されたトリム部品が元に戻り、パンクしたタイヤが空気を含んで膨らみ、タイヤ周辺の雑草が後退します。金属パネルは物理的な動きを伴って元の位置へ再形成されます。その後、車はエンジンがかかり、高速道路へと走り出します。カットもテレポートもないこと、急なジャンプや車両のアイデンティティの変化もないこと、現実的な機械的変形、ドキュメンタリー風のカメラワーク、自然光を維持してください。
ビデオの連続生成
start_video 入力パラメータが新たに追加されました。FLUX 3 に既存のクリップを与えると、その最後のフレームから続きを生成します。同じ運動量やフレーミングの論理、そして音声を保持したまま、驚くほど滑らかな映像を出力します。
ベースとなるクリップ:
スケートボーダーが日差しに照らされた通りを転がり、スケートパークの奥にあるクォーターパイプのランプへと近づき、スピードを上げながら舗装のひび割れの上を車輪がカチカチと音を立てて進みます。
連続するクリップ:
スケートボーダーはランプから飛び出し、空中でボードを掴み、着地は完璧に決まって滑り去ります。車輪はランプのトランジション部分に強く接地します。
同じアイデアでも異なるスポーツの場合:
ベースとなるクリップ:
サーファーが日の出時に波をかき分けて沖へ漕ぎ出し、海飛沫が早朝の光を捉えながら、地平線に向かって一定のリズムで漕ぎ続けます。
連続するクリップ:
サーファーがターンを決め、漕ぎ込んでくる波を捉え、立ち上がって波の面を滑り、背後で巻く波を駆け抜ける。
複数のシーンとカメラアングル
タイムスタンプ付きのショットを一つのプロンプトに直接記述できます。FLUX 3 は生成内でそれらを切り替えるため、後からつなぎ合わせる作業は不要です。
[0-4 秒]: 満員のスタジアムで歓声を上げる群衆のワイドショット、夜空に浮かぶ照明が眩しく光る。[4-8 秒]: スプリンターの足がスタートブロックから爆発的に飛び出すクローズアップへカット、塵が舞い上がる。[8-12 秒]: スプリンターがゴールラインを突破し、胸を張って群衆が沸き立つ様を追うトラッキングショットへカット。ハイパーリアル、8K、オリンピック放送の映画撮影スタイル、耳をつんざくような歓声と足音。
どんなスタイルにも対応
FLUX 3 は、多くの動画モデルがデフォルトとする「超写実的なシネマティック」なスタイルに縛られていません。必要に応じて、完全にスタイライズされた表現も可能です。
ストップモーション・クレイメーション風。指紋の跡や親指の形をした目を持つ小さな粘土キャラクターが、曲がりくねった家々が並ぶミニチュアの村を歩く。煙突からは綿のような煙が立ち上り、手作りの一歩ごとに揺らぐ。愛らしいストップモーションの美学、継ぎ目の痕や道具の跡がはっきりと見える、温かみのあるミニチュアセットの照明、柔らかい足音。
ディテールも確かに保たれています。キャラクターに指紋の模様が確認できますね?
コミック風のインク画で描かれた街並みを、手書きと 3D CGI を融合したスタイルで表現。建物にはハロトーン(半色調)のドットシェーディングを施し、キャラクターは太いインクの輪郭線で強調。色彩チャンネルがわずかな間隔で分裂するクロマティック・アベレーションのグリッチフレームが随所に現れ、すぐに元に戻る。ダイナミックなダッチアングル(傾斜した構図)のカメラワークを採用。
『スパイダーバース』のような美学に、コミックパネルのエネルギー感、そしてベースを強調したヒップホップビートを加えています。
おもしろい使い方
FLUX 3 を使えば VHS フィルムの映像が作れることに気づいた人が少なくありません。そのためには、撮影機材自体の説明が必要です。例えば、ビデオカメラの種類、露出不足の描写、オートフォーカスの不具合、不安定な構図、そして「何を撮ろうとしているのか全く分かっていない」撮影者の様子などです。
1997 年の一般向け家庭用ビデオカメラで撮影された、夜の郊外通りの生々しい映像。カメラを握る人物は、若者たちのグループの後ろを追いかけています。最初は地面を向いていましたが、皆が止まって叫ぶ瞬間にカメラを空へ向けます。屋根の上を横切る巨大な円盤型の UFO が、静かに地平線の一方から他方へと移動します。その光は駐車中の車や家の窓を照らし出します。カメラはその飛行全体を、一度もカットすることなく追跡し続けます。最後には UFO が木々の陰に消えていきます。
安価なハンドヘルド型ビデオカメラで撮影されたような質感です。焦点が合っていないソフトフォーカス、オートフォーカスが迷う様子、露出オーバーしたポーチの照明、滲んだハイライト、アナログ特有のトラッキングノイズ、インターレース方式の VHS 転送による映像。声や足音はかすれて聞こえ、編集も音楽も映画のような仕上げは一切ありません。
1995 年の家庭用ビデオ映像を再現したサンプルです。湖畔でのキャンプ旅行の夜、カメラは焚き火から始まり、ゆっくりと黒く静かな水面へとパンします。すると湖から二つ目の月が昇り、遠くの岸のすぐ上に浮かび、その反射がカメラの方へ伸びていきます。人々が水辺に向かって走り出し、カメラが揺れます。懐中電灯の光が手前を白く飛ばし、やがて月は湖へと折りたたまれるように沈み、姿を消します。これは安価な Hi8 カムコーダーで撮影された 1 枚の連続カットです。ソフトフォーカス、オートフォーカスの迷い、インターレース化された VHS 転送によるノイズ、かすれた音声、露出の不具合があり、編集や音楽、映画のような仕上げは一切施されていません。
ホームビデオは楽しいものですが、fofr が提供するこれらの例も非常に魅力的です。これらはモデルの想像力を本当に厳しく試すものです。
FLUX 3 の実行方法
FLUX 3 で動画を生成する方法をご紹介します。
Cloudflare AI Gateway を利用する場合
最も高速な AI ゲートウェイで FLUX 3 を実行したい場合は、以下の呼び出しを使用してください。
AI ゲートウェイを経由するリクエストでは、ログ記録、キャッシュ、レート制限など、ゲートウェイが提供する多彩な機能を利用できます。FLUX 3 はサードパーティ製モデルのため、Cloudflare がプロバイダーの認証情報を管理し、Unified Billing(統一課金)を通じて利用料を請求します。Black Forest Labs の API キーは不要です。
プロンプト作成のコツ
- 映像だけでなく音も描写する。FLUX 3 は動画と同時に音声も生成します。「金属が裂けるような悲鳴」や「ネオン変圧器の低い buzzing」といった、物の音を具体的に伝えることで、出力される映像そのものにも変化が生じます。
- 複数の演出要素がある場合はタイムスタンプを活用しましょう。 [0-4 秒]:... のようにプロンプトにカットポイントを明示的に指定することで、モデルがタイミングを推測するのではなく、意図したペースで映像を生成できます。 (原文の技術表記:
[0-4s]: ...)
- モルフィングの形態に合わせて再生時間を調整してください。
imageとend_imageを同時に使用する場合は、durationは「自動」ではなく整数値を指定する必要があります。 (原文の技術表記:auto)
被写体だけでなく、カメラワークを重視しましょう。「バンパーの高さでカメラが追従する」「低アングルへのハードカット」といった具体的な指示を与えることで、FLUX 3 は生成の根拠を得られます。
画像から動画への変換では、過度な計画は不要です。開始フレーム1枚と「次に何が起こるか」を記述した文章1文があれば十分です。物理演算部分はモデルが自動的に補完します。
原文を表示

Black Forest Labs made history when they released their FLUX.1 image model. They were one of the first independent research labs to pioneer fast, high-quality, open source image generation. Now, they are taking a swing at video, and we are intrigued.
FLUX 3 is the lab’s new multimodal foundation model. Compared to other labs, they are taking a new approach. BFL has developed a model with a consolidated architecture, learning from images, audio, and video to generate multimodal outputs. Taken straight from their release blog:
“Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.
Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.”
Seemingly, this training approach means that this model has more laws of physics baked in. Audio informs motion and vice versa. Video introduces temporality to images which already encode spatial relationships. FLUX 3 is an attempt at creating a net of weights that better encapsulates reality.
This post walks through some things FLUX 3 is best at, with real examples for each.
Text to video
FLUX 3 doesn’t need much to work with. Feed it a plain sentence or a dense, specific one, and it holds up either way. That includes things that require actually knowing how something works, not just what it looks like.
A desert off-road trophy truck racing across open dunes at full speed, kicking up a huge rooster tail of sand and dust behind it, suspension soaking up jumps as it crests a ridge and goes briefly airborne. Camera tracks alongside at high speed, low to the ground. Hyper-realistic, 8k, harsh midday desert light, the roar of the engine and the crunch of sand under tires.
Cinematic underwater wildlife video of a large octopus slowly crawling across a dark rocky seafloor. Low, close tracking angle at eye level, showing its textured mantle, expressive eyes, and eight arms moving independently across wet stone. Suction cups attach and release naturally as the arms pull the body forward; subtle skin ripples and realistic muscular motion. Deep teal water, drifting particles, soft shafts of filtered light, muted rust-red and brown octopus coloration, highly detailed photorealistic skin texture, shallow depth of field, natural documentary cinematography.
I like that FLUX 3 doesn’t necessarily need your prompts to be dense or overwrought with details or keywords. Sometimes, you can offload the details of motion and aesthetics to the model itself, and you’ll probably end up with a satisfying result.
A young man staring out the window of a train rumbling along in rural Switzerland, beautiful mountains in the background whooshing by. Handheld camera feel, film shot.
A diver descends through crystal-clear turquoise water into a jungle cenote, sunbeams piercing down from the opening far above. Beneath the surface, the walls of a submerged Mayan temple come into view, ancient glyphs carved into the stone, fallen pillars scattered across the floor. Fish dart through the ruins as the diver’s flashlight beam sweeps across the carvings. Hyper-realistic, 8k, National Geographic underwater cinematography, muffled bubbles and the diver’s steady breathing.
You’ll notice that FLUX 3 will default to multiple cuts of scenes unless otherwise specified.
A high-speed motorcycle chase through a rain-soaked underground parking garage at night. The rider leans hard through tight turns, headlights strobing past concrete pillars, tires screeching on wet pavement. Sparks fly as the bike clips a support beam and rights itself. Camera cuts between a low chase angle and a handlebar POV. Hyper-realistic, 8k, neon-lit puddles reflecting the headlights, screaming engine and echoing tires.
A hiker moves through a dense Costa Rican rainforest canopy, shafts of golden light breaking through layers of green leaves. A toucan takes flight overhead, a sloth hangs motionless in the branches above, and mist rises off the forest floor. The camera follows from behind at a steady walking pace, then rises to a wide canopy shot. Hyper-realistic, 8k, lush biodiversity, howler monkey calls echoing in the distance and leaves rustling underfoot.
Image to video
FLUX 3 can use a start frame and an end frame to guide a transformation. A useful test is to keep the same car and scene while changing its condition from abandoned to roadworthy.
Start and end frame. Give it two images and describe the physical change you want to see between them.
Start

End

Single continuous shot of the same rusted junkyard car transforming naturally into the restored red car. Keep the exact same vehicle identity, camera angle, wheelbase, body shape, and background throughout. Rust flakes dissolve into clean red paint, dents slowly pull themselves smooth, the cracked windshield becomes clear, missing trim reappears, flat tires inflate, and weeds pull back from around the tires. Metal panels reform in place with continuous physical motion. The car then starts and drives forward onto the highway. No cuts, no teleporting, no sudden jumps, no change of vehicle identity, realistic mechanical transformation, documentary camera, natural light.
Video continuation
The start_video input is a new parameter. You give FLUX 3 an existing clip and it keeps going from the final frame. It preserves the same momentum, framing logic, and audio. It is surprisingly fluid.
Base clip:
A skateboarder rolls down a sunlit street and approaches a quarter-pipe ramp at the end of a skatepark, picking up speed, wheels clicking over pavement cracks.
Continuation:
The skateboarder launches off the ramp, catching air, grabbing the board mid-flight, then lands cleanly and rolls away, wheels landing hard on the ramp’s transition.
Same idea, a different sport:
Base clip:
A surfer paddles out through rolling waves at sunrise, ocean spray catching the early light, paddling steadily toward the horizon.
Continuation:
The surfer spins around, paddles hard to catch a rising wave, pops up to their feet, and rides the wave’s face as it curls behind them.
Multiple scenes and camera angles
You can write timestamped shots directly into a single prompt. FLUX 3 will cut between them in one generation, so no stitching is required.
[0-4s]: Wide shot of a packed stadium crowd roaring, floodlights blazing against the night sky. [4-8s]: Cut to a close-up on a sprinter’s feet exploding out of the starting blocks, dust kicking up. [8-12s]: Cut to a tracking shot alongside the sprinter crossing the finish line, chest forward, crowd erupting. Hyper-realistic, 8k, Olympic broadcast cinematography, deafening crowd noise and pounding footsteps.
A style for everything
FLUX 3 isn’t locked into the same hyper-realistic cinematic default a lot of video models default to. It’ll go fully stylized if you ask.
Stop-motion claymation style. A small clay character with visible fingerprint textures and thumbprint eyes walks through a miniature clay village, past crooked clay houses with chimney smoke made of cotton wisps, wobbling slightly with each handmade step. Charming stop-motion aesthetic, visible seams and tool marks, warm miniature set lighting, soft foley footsteps.
Details really hold up. Do you see the fingerprint markings on the figure?
A character sprints through a comic-inked cityscape rendered in a hybrid hand-drawn and 3D CGI style, halftone dot shading across the buildings, bold ink outlines on the character, occasional chromatic-aberration glitch frames splitting the color channels for a beat before snapping back, dynamic Dutch-angle camera work. Into the Spider-Verse aesthetic, comic panel energy, bass-driven hip-hop beat.
Fun stuff
Quite a few people discovered FLUX 3 can make VHS footage. To get there, describe the recording itself: the camcorder, the bad exposure, the autofocus, the shaky framing, and the sense that the person filming had no idea what they were about to capture.
Raw 1997 consumer camcorder home video on a suburban street at night. The camera operator runs behind a group of teenagers, pointed at the pavement at first, then whips up toward the sky as everyone stops and shouts. A huge disk-shaped UFO silently crosses from one horizon to the other above the rooftops, its lights washing over parked cars and house windows. The camera follows the entire flight in one continuous unbroken take until the UFO disappears behind the trees. Cheap handheld camcorder, soft focus, autofocus hunting, blown porch lights, smeared highlights, analog tracking noise, interlaced VHS transfer, muffled voices and footsteps, no edits, no music, no cinematic polish.
Raw 1995 home video from a family camping trip beside a lake at night. The camera starts on a campfire and slowly pans to the black water; a second moon rises out of the lake and hangs just above the far shore, its reflection stretching toward the camera. Everyone runs to the water and the camera shakes, the flashlight blows out the foreground, then the moon folds back into the lake and vanishes. One continuous take on a cheap Hi8 camcorder, soft focus, autofocus hunting, interlaced VHS transfer, analog noise, muffled voices, imperfect exposure, no edits, no music, no cinematic polish.
Home videos are fun, but these examples from fofr are also compelling. They really stress-test the imagination of the model.
Running FLUX 3
Here’s how to generate a video with FLUX 3:
On Cloudflare AI Gateway
If you want to run FLUX 3 on the fastest AI gateway, use this call:
Requests through AI Gateway can use logging, caching, rate limiting, and tons of other gateway features. FLUX 3 is a third-party model, so Cloudflare handles the provider credentials and bills usage through Unified Billing. You do not need a Black Forest Labs API key.
Prompting tips
- Describe the audio, not just the picture. FLUX 3 generates sound in the same pass as the video. Telling it what things sound like, such as “the shriek of tearing metal” or “the low buzz of the neon transformer,” actually changes the output, not just the visual.
- Use timestamps for anything with more than one beat. [0-4s]: ... style prompting gives the model explicit cut points instead of leaving pacing up to chance.
- Match your duration to your morph. If you’re using image and end_image together, duration needs to be a whole number, not auto.
- Lead with the camera, not just the subject. “The camera tracks alongside at bumper height” or “hard cut to a low angle” gives FLUX 3 something concrete to hold onto.
- Don’t over-plan image-to-video. A single start frame and one sentence about what happens next is usually enough. The model fills in the physics.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み