FLUX 3 が早期アクセス開始、画像・動画・音声を統合学習するマルチモーダル基盤モデル
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
FLUX 3 は画像、動画、音声を統合的に学習する新しいマルチモーダル基盤モデルとして早期アクセスを開始し、物理世界における知覚・予測・行動の実現を目指す。
AI深層分析を開く2026年7月31日 22:44
AI深層分析
キーポイント
統一アーキテクチャによる多様なモダリティの同時学習
FLUX 3 は画像、動画、音声を個別に扱うのではなく、これらを統合した単一のアーキテクチャで同時に学習する設計を採用している。
物理世界モデルへのアプローチと Self-Flow の活用
同社は計算資源とデータリソースを大幅に拡張し、既存の「Self-Flow」手法に基づいて動画、画像、音声の生成と理解を効率的に統合している。
相互制約による現実表現の精度向上
複数のモダリティが互いの制約条件となることで、音は衝撃に一致し、運動は質量に従うなど、一つの基盤となる現実をより正確に捉えることが可能になる。
コンテンツ作成と物理 AI における初期成果
コンテンツ制作や物理 AI 分野での初期結果が示唆する通り、このアプローチは物理的・デジタル環境全体で知覚し、予測し、行動するモデルへの正しい道筋である。
マルチモーダル同時学習と能力
FLUX 3 は計算リソースとデータ規模を大幅に拡大し、動画・画像・音声を同時に学習させることで、テキストプロンプトや入力参照からこれらを組み合わせて生成する能力を獲得した。
重要な引用
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture
The modalities stop being separate and start being evidence about one underlying reality.
models that perceive, predict, and act across physical and digital environments
FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video.
編集コメントを表示
編集コメント
画像、動画、音声を統合的に学習するアプローチは、AI が物理世界をより深く理解するための重要なステップである。同社が示す「相互制約」による現実表現の強化は、今後の生成 AI の発展方向性を示唆している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
*FLUX 3 は現在、早期アクセス版として利用可能です。*
FLUX 3 は、画像・動画・音声を統合されたアーキテクチャ内で同時に学習する新しいマルチモーダル基盤モデルです。なぜなら、学ぶべきはこれらの要素を個別に切り離したものではなく、世界そのものの表現だからです。具体的には、物体がどのように結合し、物がどう動き、出来事がどのような音を伴うのかを理解する必要があります。
単一のモダリティ(情報源)だけで世界の完全な記述を得ることはできません。それぞれは同じ現実の異なる側面を捉えた投影に過ぎず、異なるセンサーによって記録される過程で情報が失われます。画像は特定の瞬間における空間構造と関係性を捉えます。動画には時間の次元が加わり、時間的なダイナミクスや物理法則が浮かび上がります。音声は、視覚だけでは検出できない機械現象と音響の間の因果関係を明らかにします。言語はこれらの知覚を目標、抽象概念、指示へと結びつけます。
一つの情報源から学ぶだけでは、その投影に関する良いモデルしか得られません。しかし、これらすべてを同時に学習すれば、各情報源が互いに制約し合うことでさらに多くのことがわかります。「音は衝撃と一致するはず」「運動は質量に従うべきだ」「未来は過去から導かれるはず」といった整合性です。モダリティは別々のものではなく、一つの現実に対する証拠として機能し始めます。
FLUX 3 は、その原則を完全に実装した最初のモデルであり、物理環境とデジタル環境の両方で知覚・予測・行動する「現実世界の視覚知能」の実現に向けた重要なマイルストーンです。コンテンツ制作や物理 AI における初期の結果は、この方向性が正しいことを示しています。
FLUX 3:1 つのモデルが担う多様な能力

FLUX 3 は、私たちの「Self-Flow」というアプローチを基盤としています。これは、同じアーキテクチャ内でマルチモーダルな生成と理解を効率的に統合する手法です。このアプローチに基づき、計算資源とデータ規模を大幅に拡張し、動画・画像・音声を同時に学習させることで FLUX 3 を訓練しました。

*Self-Flow と Flow Matching (FM) の比較。左:各モダリティごとの生成誤差(Fréchet 距離)。FM を 100 に正規化しており、数値が低いほど優れています。右:微調整後の操作タスクにおける成功率。4 つのタスクグループの平均値です。数値が高いほど優れています。*
Capabilities & Early Evaluations
FLUX 3 は、テキストプロンプトのみから、あるいは画像や動画といった入力リファレンスを提示することで、複数のモダリティを混合し、映像と音声を同時に生成する能力を持っています。ここでは、同モデルの主要な機能をいくつか紹介します。
Video
FLUX 3 は、一度の生成で最大 20 秒に及ぶ多様な動画を作成できます。すべての出力にはネイティブ音声も含まれます。
その中核的な機能は以下の通りです。
- テキストから動画への変換(Text-to-video)
- 画像から動画への変換(Image-to-video)。開始フレームからの連続アニメーションや、視覚リファレンスとしての画像利用が可能です。
- リファレンスクリップからの動画生成(Video-to-video)。ソース動画の主要要素(例えば同一キャラクターなど)を維持しつつ、新しいシーンや文脈へと展開します。
- 入力された映像と音声に基づく、生成型ビデオ・オーディオの継続
- キーフレームから動画への変換。定義された瞬間間の制御されたトランジションを実現します。
- 多言語での対話対応
- 従来の映画作品に限定されない、幅広いビジュアルスタイルとアスペクト比への対応
- 個別のクリップを結合し、より長く複雑なマルチショットシーケンスへと拡張するエージェント機能
- 高いスタイルの多様性。FLUX 3 Video は、カジュアルな家庭用カメラ映像からアニメーション、そしてシネマティックな表現まで、あらゆるスタイル範囲を容易に処理します。
- 強力なタイポグラフィ生成とアニメーションデザインの実装
以下の予備分析では、音声付きの 720p 解像度で 10 秒間のテキストから動画への生成サンプルを作成しました。

*評価は初期段階であり、さらなる改善を期待しています*
モデルおよびその周辺ハッチネス(検証環境)はまだ開発中のため、本稿の結果は暫定的なものです。早期アクセス期間中にさらに向上する見込みです。
初期評価では、FLUX 3 は他社製品との比較で以下の割合で優位性を示しました:Grok Imagine Video では最大 69%、Kling v3 Pro で 60%、Happy Horse v1 で 59%、Happy Horse 1.1 で 57%、Seedance 2.0 および Gemini Omni Flash でそれぞれ 52% です。また、Runway Gen-4.5 との比較では 77%、Luma Ray 3.2 との比較では 93% のケースで FLUX 3 が選ばれました。
開発中ではありますが、FLUX 3 Video はすでに人間の表情の描写、物理現象と音の対応付け、多言語対応において特に高い性能を発揮しています。さらにこれらの機能を組み合わせることで、数分間にわたる連続したシーンの生成も可能になります。視覚的なリファレンス(参照画像)を活用することで、登場人物がシーン全体を通じて一貫して描かれるよう保証されます。
FLUX 3 Video は現在、早期アクセスページで利用可能です。
Image
FLUX 3 は、多様なスタイルやアスペクト比、解像度に対応した画像の合成と編集が可能です。トレーニング途中での初期評価では、既存の FLUX バージョンと比較して大幅な進化が確認されました。特に複雑なプロンプトへの対応力やテキスト生成能力が向上しており、多言語で高精度なテキストをレンダリングできるほか、以下に示すような幅広い出力スタイルも生み出せます。

動画の評価と同様に、これらはリリース前の暫定的な結果です。公開までにさらに改善が見込まれます。FLUX 3 Image の早期アクセスプログラムは、今後数週間で開始する予定です。
Action
FLUX 3 は世界理解の範囲を「行動予測」へと広げました。その実現には2 つのアプローチを採用しています。1 つ目は、FLUX 3 にネイティブな行動予測機能を直接統合し、Self-Flow で取り組んできた初期成果を拡張する手法です。2 つ目は、事前学習済みの動画バックボーンを「動的変化を意識した基盤」として活用し、限られたタスク固有データで専門的な行動モデルを微調整できるようにするアプローチです。
2 つ目のパートナーとして、Mimic Robotics は FLUX 3 の早期アクセス権を最初に獲得した企業のひとつです。私たちは Mimic と共同で「FLUX-mimic」を開発しました。これは、FLUX 3 のバックボーンに、巧みな操作や実環境での展開における Mimic のロボット学習の専門知識を組み合わせた動画アクションモデルです。
物理 AI とコンテンツ制作が同じ基盤上で成り立っている理由と、それが Audi での実際の生産タスクでどのように検証されているかについて詳しくは、当社の論文をご覧ください。https://bfl.ai/blog/flux-3-mimic
ローンチ計画
今後数週間から数ヶ月にかけて、以下の機能を順次提供していきます。各機能は、円滑な展開とフィードバックの収集、そして厳格な安全性テストを確保するため、まず早期アクセス期間を経て公開されます。これらすべての機能は、同じ基盤となるマルチモーダル・フローマッチングモデルから構築されています。
- API およびプライベートウェイトへのアクセスを通じて実現する動画・音声の生成と編集(「FLUX 3 Video」)
- 特定の研究機関および商用パートナーを通じたアクション予測。まずは Mimic Robotics から開始します(「FLUX-mimic」と「FLUX 3 Action」)
- API およびプライベートウェイトへのアクセスを通じて実現する画像合成と編集(「FLUX 3 Image」)
- コンテンツ制作(動画、音声、画像)およびアクション予測のためのオープンウェイトによるマルチモーダル・バックボーンへのアクセス(「FLUX 3 Dev」)
また、基盤となるアプローチに関する技術的な詳細も今後公開していきます。
次のステップ
多機能で能力に富んだ統合型マルチモーダルモデル、そしてそれがもたらす未来について、私たちはまだ表面に触れたばかりです。インタラクティブな画像・動画編集からシミュレーション、コンピュータ操作、さらには物理的な AI まで、可能性の地平は広がり続けています。
これらの新機能を順次展開していく一方で、すでに次世代モデルの開発に取り組んでいます。私たちの目標は、知覚、行動、言語予測を一つの統合モデルに集約することです。
FLUX 3 を活用して探索や開発を行いたい方は、こちらからご連絡ください。また、当社のミッション実現にご協力いただける方も大歓迎です。ドイツおよび米国で採用活動を行っていますので、ぜひご応募ください。
原文を表示
*FLUX 3 is now available in Early Access.*
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.
No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.
Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.
FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path.
FLUX 3: One model, multiple capabilities.

FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.

*Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).*
Capabilities & Early Evaluations
As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model’s key capabilities below.
Video
FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation.
Its core capabilities include the following (all outputs come with native audio generation):
- Text-to-video generation.
- Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references.
- Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context.
- Generative video-audio continuation from input video and audio.
- Keyframe-to-video generation for controlled transitions between defined moments.
- Multilingual dialogue.
- A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output.
- Agentic chaining of individual clips into longer, multi-shot sequences.
- High style diversity -- FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.
- Strong typography generation and animated designs.
For the preliminary analysis below, we generated 10-second text-to-video clips in 720p with audio.

*Evaluations are early and we expect further improvements*
As the model and the harness around it are still in development, these results are preliminary, and we expect further improvements during the early access phase. Across early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons.
While still in development, FLUX 3 Video is already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities. Furthermore, these capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that the characters remain consistent across all scenes.
FLUX 3 Video is now available in Early Access here
Image
FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations conducted during midtraining, FLUX 3 already shows a significant improvement over earlier versions of FLUX: its ability to handle complex prompts and text generation has improved significantly. The model produces a wide range of output styles (see the following samples), and is able to render high-accuracy text in multiple languages.

As with video evaluations, these are preliminary results, and we expect further improvements before release. We will open up an early access phase for FLUX 3 Image in the following weeks.
Action
FLUX 3's world understanding extends to action prediction. We have taken two routes to it: integrating native action prediction into FLUX 3 directly, scaling up our initial work in Self-Flow; and using the pretrained video backbone as a dynamics-aware foundation that specialized action models can be finetuned from with limited task-specific data.
For the second, mimic robotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation and production deployment. Read our thesis on why physical AI and content creation run on the same foundation, and how it's being tested on real production tasks at Audi.
Launch Plan
Over the next few weeks and months, we will make the following capabilities available, each after an early access phase for ensuring smooth rollout, collecting feedback and rigorous safety-testing. All capabilities are built from the same underlying multimodal flow matching model. These capabilities and models include:
- Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
- Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
- Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
We will also release more technical details on the underlying approach.
What’s next?
We are only beginning to scratch the surface of versatile, capable, unified multimodal models, and what they will enable. From interactive image & video editing, simulation to computer use and physical AI, the frontier is wide open. While we gradually roll out these new capabilities, we are already working on the next generation models. Our goal is to unify perceptual, action and language prediction in the same unified model.
If you are interested in exploring and building with FLUX 3, get in touch here. If you are interested in contributing to our mission, join us! We are hiring in Germany and the US.
AI算出
主要ニュースainew評価高い
記事は FLUX 3 の早期アクセス開始を報じており、画像・動画・音声を統合学習する画期的なマルチモーダル基盤モデルとしての新事実と技術的特徴(Self-Flow アプローチ)を具体的に記述しているため、新規性と関連性は高い。ただし、日本企業や日本固有の情報が含まれていないため、日本の文脈での関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 50
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み