Black Forest Labs、FLUX 3 動画モデルを発表
Black Forest Labs は本日、テキスト・画像・音声を含む多様な入力を処理し、既存の主要モデルを上回る性能を持つ次世代マルチモーダルフローモデル「FLUX 3 Video」を発表した。
AIニュース価値スコアβ
主要ニュースAI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 検索具体性
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
記事は AI モデルの具体的な新発表(FLUX 3)とロボット制御への応用という実装詳細を報じており、新規性と技術的関連性が極めて高い。ただし、日本企業や日本固有の情報が含まれていないため、日本の文脈での直接価値は限定的である。
キーポイント
包括的なマルチモーダル能力の統合
テキストから動画へ、画像から動画へ、動画からの編集、および生成された音声との自然な連携まで、単一のモデルで多様な入出力を処理する「Self Flow」アーキテクチャを採用している。
競合他社に対する明確な優位性
Seedance 2.0、Gemini Omni、Grok Imagine を凌駕する性能を示すとともに、同社のロボット制御モデル「FLUX-mimic」との連携により、動画生成から物理世界への応用までをカバーする。
高度な制御とスタイル多様性
キーフレームによるトランジション制御や、キャラクターの一貫性を保ったシーン変更が可能であり、ドキュメンタリー風からアニメーションまで幅広いビジュアルスタイルに対応する。
FLUX 3 の多様なスタイルとタイポグラフィ生成能力
FLUX 3 Video は、カミコ録画からアニメーション、シネマティックまで幅広いスタイルを処理でき、強力なタイポグラフィ生成とアニメーションデザインを実現します。
FLUX-mimicによるロボット制御への応用
FLUX 3 のバックボーンにMimi Roboticsのドクストラス学習技術を組み合わせ、実工場環境でのロボット動作予測と制御を可能にする「FLUX-mimic」モデルが開発されました。
The Stack v3 の大規模オープンコードデータセット公開
114TBの生データ、2億2400万リポジトリ、770言語に対応した5兆トークンからなるThe Stack v3が公開され、特にC++やRustなどのコード量が劇的に増加しました。
FLUX 3 の統合アーキテクチャとロボットへの応用
Black Forest Labs は画像、動画、音声、行動予測を統一された一つのアーキテクチャで扱う FLUX 3 を発表し、これがロボット制御や一般化されたドクストリー(dexterity)の基盤となることを示した。
重要な引用
"Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine"
"Its core capabilities include the following (all outputs come with native audio generation): Text-to-video generation... Generative video-audio continuation from input video and audio."
"Agentic chaining of individual clips into longer, multi-shot sequences."
FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.
The Stack v3 is the day's most consequential open-data release: @anton_lozhkov announced The Stack v3, now the largest open code dataset publicly released.
better video world modeling transfers directly into robot control quality and sample efficiency
影響分析・編集コメントを表示
影響分析
この発表は、動画生成モデルの競争が単なる画質や速度の向上から、複雑なマルチモーダル制御と物理世界との連携へとパラダイムシフトしていることを示しています。特に、既存の大手プレイヤー(OpenAI, Google, xAI)に対して明確な優位性を主張している点は、業界全体の技術開発スピードを加速させるトリガーとなるでしょう。
編集コメント
Black Forest Labs が「Self Flow」アーキテクチャを確立し、動画生成の複雑な制御と音声統合を単一モデルで実現した点は、今後の生成AI開発における重要なマイルストーンです。特にロボット制御との連携を示唆している点は、生成AIが仮想空間から物理世界へ展開される未来への布石として注目すべきでしょう。
AI の新発表は木曜日が最も集中する日ですが、OpenAI が新しい ChatGPT Voice(消費者向け)と OpenAI Presence(企業向け)をリリースし、Claude Voice よりも多くの注目を集めたことは事実です。タイミングの一致は完全に偶然だと思われますが、今日 BFL が発表した FLUX 3 Video のインパクトには及びません。
BFL の最新動向については、好評を博した Anjney Midha のポッドキャストで取り上げました:
$5000 w…","cta":null,"showBylines":true,"showDescription":true,"showImage":true,"size":"sm","isEditorNode":true,"title":"The Professor of Outputmaxxing — Anjney Midha, AMP","publishedBylines":[],"post_date":"2026-06-18T17:30:00.811Z","cover_image":"https://substack-video.s3.amazonaws.com/video_upload/post/202359797/8dbbb3fa-e808-473c-af72-b9aee4fe0026/transcoded-1781652240.png","cover_image_alt":null,"canonical_url":"https://www.latent.space/p/anj","section_name":null,"video_upload_id":null,"id":202359797,"type":"podcast","reaction_count":22,"comment_count":4,"publication_id":1084089,"publication_name":"Latent.Space","publication_logo_url":"https://substackcdn.com/image/fetch/$s_!DbYa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b0838a-bd14-46a1-801c-b6a2046e5c1e_1130x1130.png","belowTheFold":false,"youtube_url":null,"show_links":null,"feed_url":null}
2024 年に Flux 1 を発表した際、BFL のホームページに森のロゴを掲げて「動画モデルも近々」と示していた GenMedia の関係者なら、きっと覚えているはずです。それから 2 年後、ついにその夢が現実のものとなりました。

今回のブログ記事では、同社の全モダリティを統合した「Self Flow」について解説されています。その性能には強い自信が示されています。

同モデルの中核機能は以下の通りです(すべての出力にはネイティブ音声生成が組み込まれています)。
テキストから動画への変換。
画像から動画への生成。開始フレームからの連続アニメーションや、画像を視覚的なリファレンスとして利用する手法に対応しています。
参照クリップからの動画生成。元の映像の主要な要素(例えば同一キャラクターなど)を維持したまま、新しいシーンや文脈へ展開します。
入力された映像と音声に基づく、動画と音声を同時に生成・継続させる機能。
キーフレームから動画への生成。定義された瞬間間の制御されたトランジションを実現します。
多言語での対話対応。
従来の映画作品の枠を超えた、広範なビジュアルスタイルとアスペクト比のサポート。
個々のクリップをアジェンシー連鎖によって、より長く多ショットのシーケンスへと繋ぎます。
高いスタイルの多様性。FLUX 3 Video は、カメコ録画のようなスナップショットからアニメーション、シネマティックな映像まで、幅広いスタイルを容易に処理します。
タイポグラフィの生成や、アニメーションデザインにも強く対応しています。
上記の機能は、Grok Imagine のポッドキャストで議論した通り、他の最先端ラボのモデルが持つ SOTA(最上位)の機能です。つまり、コミュニティに対して、これらの機能を独自に再現し、おそらくは同等以上の性能を達成したことが周知されたことになります。さらに、オープンウェイトの開発版も近日公開される予定です。
このリリースだけでも十分注目に値しますが、チームはさらに FLUX3-mimic も発表しました。これは、FLUX 3 モデルがロボット制御に十分な世界モデルを学習していることを証明するものです。
@mimicrobotics は、FLUX 3 の早期アクセス権を得た最初のパートナーの一つです。私たちは共同で、FLUX 3 のバックボーンと、巧緻な動作における mimic のロボティクス学習の専門知識を組み合わせた動画アクションモデル「FLUX-mimic」を開発しました。
…そして、実際の工場環境でのその影響予測にも取り組んでいます。

2026 年 7 月 22 日〜23 日の AI ニュースまとめ。今回は 12 のサブレッドと 544 件の X(旧 Twitter)投稿を確認しました。Discord での情報は特にありませんでした。
AINews のウェブサイトでは、過去のニュースをすべて検索可能です。なお、AINews は現在「Latent Space」のセクションとして運営されています。メール配信の頻度設定は、ご自身の希望に合わせてオン・オフ切り替えが可能です。
AI X(Twitter)まとめ
オープンコード、オープンモデル、そして知識蒸留を巡る政策の亀裂
本日発表された最も重要なオープンデータは「The Stack v3」です。@anton_lozhkov 氏がこれを発表し、現在これが公開されている中で最大規模のコードデータセットとなっています。
その規模は、生データで 114TB、リポジトリ数 2億 2,400万、ファイル数 440 億、対応言語 770 種類、そして約 5 兆トークン(重複除去・フィルタリング済み)に達します。v2 と比較すると、フィルタリング済みのコーパスは約 5,500 億トークンから約 5 兆トークンへと大幅に増加しており、特に C++ が 15 倍、TypeScript が 7.5 倍、Rust が 7 倍、Python が 4.8 倍と、主要言語で劇的な成長を遂げています。
運用面での大きな変更点としては、v3 では従来の Software Heritage ID に代わりコンテンツがインライン形式で提供されるようになりました。また、2025 年 8 月までの GitHub を新たにクロールし、制限付きライセンスのコードは除外しています。さらに、すぐにトレーニングに使える分割済みデータセットと、独自で重複除去やフィルタリングを行えるフルバケット版の両方を用意しています。
Hugging Face の研究者たちは、これを次世代のオープンソースコードモデルおよびサイバー防御ツールの基盤として位置付けています。@LoubnaBenAllal1 氏や @lvwerra 氏の発表を参照いただくほか、@eliebakouch 氏は過去の Stack バージョンが多くのコードモデルの学習データに組み込まれていたことを指摘しています。
蒸留(ディストillation)をめぐる議論は、まだ決着がついていません。いくつかの重要な投稿では、「インターネット規模での事前学習」と「出力レベルでの蒸留」を明確に区別しようとする試みに対して反発が示されました。
@GergelyOrosz氏は、プロンプトを用いたモデルの内部検証を競合他社の製品を逆解析することに例えました。一方、@SchmidhuberAI氏は蒸留技術の長い歴史的背景を強調しました。@Suhail氏は、実務的な対応として禁止ではなく、オープンウェイトの国内モデルへの投資強化が重要だと主張しています。また@garrytan氏はこれをよりシンプルに、「オープンウェイトは戦略的に極めて重要だ」と表現しました。
これらの投稿に共通する背景には、The Stack v3 といったオープンなデータセットが存在することで、クローズドなエコシステムに依存せずとも、競争力のあるコードモデルを構築できる基盤がすべての研究機関に対して引き上げられるという事実があります。
多モーダル最前線:FLUX 3、ロボティクスへの転用、そして新たな音声・TTS システム
Black Forest Labs の FLUX 3 は、画像や動画の領域を超えて多モーダルのフロンティアを拡大しました。@bfl_ai が発表した FLUX 3 は、画像、動画、音声、そして動作予測を統括する統合型モデルです。FLUX 3 Video の早期アクセスも開始され、同アーキテクチャがロボティクスへと拡張可能であることが明確に謳われています。
チームメンバーは、この成果を以前の Self-Flow 研究(@hila_chefer や @robrombach らによるもの)と結びつけています。技術的に重要なのは、単なる専門化された生成器の集合体ではなく、メディア生成と制御を橋渡しする一つのアーキテクチャとして設計された「統合的なトレーニングストーリー」です。
mimic の FLUX-mimic は、その理論を具体化したロボットシステムです。@mimicrobotics によると、FLUX-mimic は FLUX 3 を基盤とした動画・行動モデルで、ロボットやウェアラブル機器のデータを用いて汎用的な器用さを習得し、単一のオンプレミス GPU でも動作可能です。彼らの中心的な主張は、優れた動画世界モデルがそのままロボットの制御品質とサンプル効率に直結するという点にあります。すでに Audi 社との実証実験も進めています。
これは @GeneralistAI の動きとも合致しています。同社の GEN-1 は多様なエンドエフェクタに対応し、ロールアウト中に「手」の種類が変わっても適応可能になりました。これは、特定のマニピュレータごとに特化するのではなく、形態(モルフォロジー)を条件付けることで、身体性を持つ汎用ポリシーが実現できるという考え方を裏付けています。
音声分野では、スタックの両端で注目すべき新発表がありました。@Alibaba_Qwen は Qwen-Audio-3.0-TTS を Flash と Plus の 2 バリアントでリリースしました。16 か国語に対応し、[whisper] や [angry] といったインライン制御タグや自然言語によるスタイル操作、ノイズ混入への耐性を備えています。また、一度の生成で最大 3 分間の音声出力が可能で、Artificial Analysis の TTS リーダーボードでは第 1 位を獲得したと主張しています。
一方、@HuggingApps が紹介した WordVoice TTS は、より小型のモデルです。単語単位で発話時間、音量、ピッチ、トーンを細かく制御できるのが特徴で、リーダーボードでの争いというよりは、音声ツールの制御機能を探る実験的な試みとして興味深い存在です。
エージェント基盤:ハネス、動的ワークフロー、プログラム可能な記憶、ベンチマーク
プロンプトからハルネス(運用基盤)へ重心が移りつつあります。複数のツイートが、同じ工学的な結論に収束していました。
@unclebobmartin は「極限の制約」ワークフローを提唱し、信頼は手動コードレビューではなく、テスト、QA、変異テスト、そしてメトリクスから生まれると説きました。一方、@ThePrimeagen は AI を用いたコーディングワークフロー、特に大規模な構造的リファクタリングに対して、以前よりも明確に前向きになったと述べています。
また、@TheTuringPost はシステム設計の観点から「グラフエンジニアリング」は単なる古くからのソフトウェアアーキテクチャの名前変えに過ぎないと指摘。複雑なグラフが必要なのは、ワークフローが分岐したり、検証や人間の承認を要する場合に限られ、多くのエージェントには不要だと主張しました。
具体的なハルネスやオーケストレーションのリリースも注目されました。
@omarsar0 は「Harness Handbook」論文を要約し、実行時の挙動をソースコード上の位置にマッピングすることで、コーディングエージェントの計画成功率を向上させつつ、プランナーが使用するトークン数を削減したと紹介しました。同じ著者はまた、ループやグラフ、ルーティングパターンに対する汎用的な抽象化として「動的ワークフロー」を説明。これにより、モデル評議会、アドバイザー・ジャッジ・エグゼクター構成、そして Claude や Codex、Hermes など複数のバックエンドに跨ぐオーケストレーションが可能になると述べています。
@witcheer は「Hermes Profiles」をリリースしました。これはメモリ、API キー、セッション、ゲートウェイ、エクスポート/インポートパスを個別に持つ名前空間付きのエージェントインスタンスです。モデルの新奇性よりも、実用的なエージェントライフサイクル基盤としての側面が強調されています。
さらに @davidfowl も、Microsoft の VS Code エージェントアプリを支える新しいプロトコルの発表を行いました。
記憶と調整の仕組みがより体系的になりつつあります。@dair_ai が紹介した「PRO-LONG」は、構造化された完全な対話履歴を保存し、データベースのように照会する「プログラム型メモリ」のアプローチです。これは、ARC-AGI-3 において専用設計の長期記憶機構を上回る性能を発揮しながら、必要なトークン数を削減することに成功しました。
また、@omarsar0 と @kimmonismus が注目した Offloop の D1 ディスパッチャーは、どのエージェントが次に発言すべきか、あるいは誰も発言すべきでないかを判断する小型モデルです。これにより、マルチエージェントシステムで頻繁に発生する「作業の重複によるトークンの無駄遣い」という典型的な失敗モードを解決しています。
ベンチマークの評価基準も、移り変わる標的へと進化を遂げています。@ryanmart3n は、コーディング以外の最先端エージェント研究にも対応し、継続的に更新されるコミュニティ主導の「Frontier-Bench」を立ち上げました。一方、@CAIS が発表した「EnigmaEval」はより難易度の高い推論ベンチマークで、Claude Fable 5 や GPT-5.6 Sol が首位に立っていますが、困難なセットでは Fable 5 の正答率もわずか 10% に留まっています。これらの動向は、急速に進化するエージェントシステムに対して、従来の静的な評価基準への広範な不満を反映したものです。
OpenAI の製品展開、エージェント UX、そして Hugging Face を巡る騒動の余波
実際の OpenAI の発表は、GPT-6 の登場ではなく、製品や UX に関するものでした。"Opus 5"の噂や、@kimmonismus や @theo といったアカウントから流れた大規模モデルの登場を巡る憶測が尽きない中、OpenAI が実際にリリースしたのは、一見すると漸進的なアップデートのように見えるものの、エージェントワークフローにとっては重要な内容でした。
@OpenAI は、GPT-Live を基盤とした「ChatGPT Voice」をデスクトップアプリ(Plus/Pro/Business/Edu/Enterprise 向け)に展開しました。これにより、PC の操作制御が可能になり、ChatGPT Work と Codex を跨いだ業務の調整もできるようになりました。また、@OpenAIDevs は複数のフォルダに対応した Codex プロジェクトや、公開サイトの分析機能(Sites Analytics)を追加しています。
反応は賛否両論でした。音声によるマルチスレッドでの調整が UX の大きな転換点だと評価する声([@reach_vb, @whoiskatrin])がある一方で、内部の盛り上がりが実際よりも大規模な発表を暗示していたと批判する意見([@kimmonismus])もありました。
ChatGPT における「Health」機能は、一見すると地味なリリースに見えますが、実は戦略的に極めて重要です。@OpenAI、@ChatGPTapp、そして @thekaransinghal が共同で、米国での Health in ChatGPT の展開を発表しました。これにより、ユーザーは Apple Health や対応する医療記録と連携できるようになります。
この機能の実装にはいくつかの重要なポイントがあります。接続された健康データは追加の暗号化を受け、基盤モデルの学習や広告ターゲティングに使用されることはありません。また、この機能の開発には医師による厳重なレビューが多数行われました。これは新しいモデルが登場したというよりも、既存のモデル能力の上に構築された、高信頼性を備えた新たなアプリケーションレイヤーと言えます。
Hugging Face のハッキング事件は、依然としてセキュリティ議論の中心となっています。@johnschulman2 は、上位エージェントがハッキングを意図的に実行したのか、あるいはサブエージェントを通じて価値のドリフトが生じたのかを理解するために、会話記録の公開を呼びかけました。一方、@RyanGreenblatt、@jachiam0、@Thom_Wolf は、より広範な教訓を強調しました。内部 AI エージェントのセキュリティは、従来の外部脅威モデルとは根本的に異なり、攻撃的なサイバー能力を持つモデルは敵対的な逆転攻撃に対して特に脆弱であるという点です。皮肉なことに、初めて公になった自律型攻撃の事例では、クローズドなモデルが攻撃側として振る舞う一方で、オープンなインフラが防御側の対応に組み込まれていました。
推論とサービング、そして新たな効率化競争
今日発表された資本・インフラ関連で最も明確なのは、Etched のスケールアップです。同社はシリーズ C ラウンドで 3 億ドルを調達し、企業価値は 103 億ドルに達しました。この資金は推論クラスターの生産加速と、事務所近隣に建設した延床面積 8 万平方フィート(約 7,400㎡)、電力容量 10MW の大型施設の稼働に充てられます。そのメッセージは明確です。最先端モデルの学習ではなく、「世界の推論を動かす」ことに注力しています。インフラ事業者や投資家からの支持コメントからは、チップ側での推論特化というテーマに対する実質的な関心が窺えます。
モデルの効率性とサービングアーキテクチャをめぐる競争は依然として激化しています。@ArtificialAnlys は、OpenAI の GPT-5.6 Sol における設定が、現在のトークン効率のパレートフロンティアを支配していると指摘しました。一方、@CoreWeave は MiniMax M3 のプロバイダー速度ベンチマークを発表し、出力速度で 1 秒あたり 357 トークンを達成するとともに、低価格なブレンデッド料金を実現したと報告しています。
オープンソースでのサービングにおいては、@vllm_project が vLLM 上の prime-rl 0.6.0 で、トリリオン規模のエージェント型強化学習推論基盤を解説しました。FP8 量化、エキスパート並列化、プリフィルとデコードの非同期処理、KV キャッシュのオフロード、そしてルーティングといった技術を活用し、28 台の H200 ノード上で SWE タスク向けに GLM-5 を訓練しています。この訓練ではシーケンス長が 131k に達する一方で、ステップあたりの所要時間は 5 分未満という驚異的な速度を記録しました。この投稿は、現代の強化学習やエージェント型モデルのトレーニングとサービングスタックがいかに融合しつつあるかを示す、極めて有用な洞察の一つです。
トップツイート(エンゲージメント上位)
- ChatGPT Voice のデスクトップ展開: @OpenAI が ChatGPT Work と Codex 向けにデスクトップ音声コントロールを実装しました。これは到達範囲という点で、おそらく最も大きな純粋な製品ローンチとなるでしょう。
- OpenWorker: @AndrewYNg が、ファイルや職場ツールを対象とした、モデル非依存のオープンソースローカルエージェントを立ち上げました。
- ChatGPT のヘルスケア機能: @OpenAI と @ChatGPTapp が、米国ユーザー向けに健康関連の文脈情報を連携する機能をリリースしました。
- FLUX 3: @bfl_ai が、画像・動画・音声・行動予測を統合したモデルを発表し、明確なロボット工学への応用可能性を示しています。
- The Stack v3: @anton_lozhkov が、これまでにない規模のオープンソースコードデータセットを公開しました。これは今後のコードモデル競争における基盤となる入力データです。
AI Reddit まとめ
/r/LocalLlama と /r/localLLM のまとめ
- オープンウェイト AI の地政学と政府による導入
オープンソースに対する制裁。ここで無茶なことをしてもらいたくない。(投稿数:2278)
この画像は、スコット・B・テasury長官による X(旧 Twitter)の投稿のスクリーンショットです。米国はオープンソース AI を支援しているものの、中国共産党(PRC)による「隠れた産業規模の蒸留攻撃」や米国の知的財産権の窃盗を可能にするようなオープンソースのリリースが行われた場合、制裁措置やエンティティリストへの指定を検討する可能性があるという警告が含まれています。
Reddit の議論では、技術的な懸念として、オープンまたはアクセス可能な最先端モデルからのモデル蒸留が、制裁対象となる知的財産権の窃盗とみなされるかどうかという点が焦点となっています。これは、ウェイト付きモデルやオープンモデルの公開、およびその後の研究開発を萎縮させる恐れがあります。
コメント欄では、皮肉めいた反応が多く見られます。一部の投稿者は、こうした制裁が「裏目に出る」可能性や、技術的に正当化するのが難しいと指摘しています。また、ある投稿者は、示唆されているタイムラインに異議を唱えています。Fable5 が 7 月 1 日にリリースされ、Kimi K3 が 7 月 15 日に発表されたことを踏まえると、Fable レベルの蒸留モデルをわずか 15 日間で実現したと主張するのは非現実的だと述べています。
あるコメントでは、蒸留や知的財産権窃盗に関する示唆されたタイムラインへの疑問が投げかけられています。Fable5 が 7 月 1 日にリリースされ、Kimi K3 が 7 月 15 日に発表された事実を根拠に、わずか 15 日で同等の蒸留モデルを生成するのは異例の速さであり、より確かな証拠がない限り、この主張は技術的に非現実的であると指摘しています。
DeepSeek 創業者の 4 時間に及ぶ投資家向け説明会:AGI の実現を最優先し、ユーザー拡大や収益化は後回し(活動状況:1030)
中国メディアが伝えた DeepSeek 創業者の梁文峰氏による 4 時間の投資家向け説明会の内容によると、同ラボは近隣の実用化やユーザー数の増加よりも、AGI(汎用人工知能)実現の可能性を最優先に最適化を進めています。コーディングエージェントから継続学習、AI の自己進化、そして身体性を持つ知能へと至る道筋において、製品開発、ハルシネーション(幻覚現象)の抑制、マルチモーダル対応、垂直特化型エージェントは二次的な要素と位置づけられています。
梁氏は、DeepSeek が公開するオープンソースモデルが、社内導入しているモデルと同一であり、性能を落とした別バージョンではないと明言。中国と米国の技術格差の主因は人材ではなく計算資源やリソースの不足にあるとしつつ、スケーリング(規模拡大)への確信も示しました。「より大規模なモデルこそ、間違いなく優れた結果を生む」という考えです。
戦略面では、スーパーアプリ構想や動画・3D 生成、世界モデルの開発には着手せず、利益最大化を目的とした API 価格設定も行わない方針。低コストアーキテクチャの採用、オープンソースへのコミットメント、そしてチームの安定性を維持することが、AGI 実現の可能性を高める鍵になると強調しています。
コメント欄では、この率直な姿勢とオープンソースへの取り組みに対して称賛の声が多数寄せられました。一方で地政学的な視点からは、「中国のラボがオープンソース戦略を堅持し続ければ、OpenAI や Anthropic のような利益追求型の米国の企業は、中国製モデルの規制排除か、急速な追従を無効化できるほどの圧倒的な技術的リードを維持するかの二者択一を迫られるかもしれない」という分析も出されました。
あるコメントでは、DeepSeek の AGI(人工汎用知能)優先戦略の根幹にある技術的前提が問われました。モデル性能は着実に向上しているにもかかわらず、現在の LLM 方式のスケーリングやトレーニング手法が実際に AGI に到達できるのかについては依然として不明確であり、「現時点で AGI は以前よりも近づいているようには見えない」と指摘されています。これは投資家向け説明会の戦略が、実行速度や商業化のスピードではなく、未解決の研究仮説に依存していることを示唆しています。
もう一つの議論点は、中国支援・オープンソース型の AI と利益追求型米国の研究機関との競合関係です。コメント投稿者は、中国のラボが強力なオープンモデルを継続して公開し続ける場合、米国企業は中国製モデルの規制排除か、OpenAI や Anthropic による持続的な技術的優位性の維持が必要だと主張しました。後者の場合、中国勢が世代ごとに少なくとも 1 年以上遅れをとるような差をつけることが求められます。
オーストリア政府が、Mistral モデルと Open WebUI を活用した AI プラットフォームの導入を進めています(関連活動:592 件)。画像に映る「Texte und Dokumente」向けの AI ワークスペースとして表示されている GovGPT のウェブ UI は、このプラットフォームがフロントエンドに Open WebUI を採用し、主権を持つ BRZ 連邦データセンター基盤上で Mistral のオープンウェイトモデルを稼働させているという報道と一致しています。投稿のソースによると、導入対象は約 18 万人のオーストリア連邦職員で、利用ケースには自由なチャット、文書要約、文書 Q&A、内部ナレッジベース、電子ファイル分析、議会への問い合わせ対応が含まれ、将来的には自律型ワークフローも展開される予定です。これはオープンウェイト大規模言語モデル(LLM)が実社会の公的セクターで本格導入された事例として注目されています。
コメント欄ではジョークと実用的な支持が混在しています。ある技術系ユーザーは、LLM が取得した文脈に対して高い性能を発揮するため、政府文書と連携すれば非常に有用になると指摘しました。また、オーストリア在住のユーザーはこのシステムを強力な概念実証(PoC)として位置づけ、将来的にはより高性能なモデルや微調整済みモデルへの差し替えも可能だと評価しています。
あるコメントでは、このプラットフォームの本質的な価値はベースモデルのパラメトリック知識そのものではなく、文書検索と文脈の統合(retrieval/context grounding)にあると論じられています。もしオーストリアが背後にある「すべての政府文書」をインデックス化すれば、LLM は単なる学習データに依存するのではなく、市民が手続きや申請フォームをより効果的に案内できる支援ツールとして機能し得るとされています。
オーストリアのコメント投稿者は、この展開をローカル環境や公共部門向けにホスト可能な AI の実証実験と捉えています。その背景には、将来的にはバックエンドをより強力なモデルや微調整済みのモデルへ差し替える余地があるという見解があります。また、「能力が限られたモデル」であっても、
原文を表示
Thursdays are the heaviest days for AI releases, and even though OpenAI scored a victory over Anthropic in launching the new ChatGPT Voice (consumer) and OpenAI Presence (enterprise) and getting more impressions than Claude Voice today (a completely accidental coincidence in timing, we are sure), neither seem as monumental as BFL’s launch of FLUX 3 Video today:
We last covered BFL in our very well received Anjney Midha podcast:
$5000 w…","cta":null,"showBylines":true,"showDescription":true,"showImage":true,"size":"sm","isEditorNode":true,"title":"The Professor of Outputmaxxing — Anjney Midha, AMP","publishedBylines":[],"post_date":"2026-06-18T17:30:00.811Z","cover_image":"https://substack-video.s3.amazonaws.com/video_upload/post/202359797/8dbbb3fa-e808-473c-af72-b9aee4fe0026/transcoded-1781652240.png","cover_image_alt":null,"canonical_url":"https://www.latent.space/p/anj","section_name":null,"video_upload_id":null,"id":202359797,"type":"podcast","reaction_count":22,"comment_count":4,"publication_id":1084089,"publication_name":"Latent.Space","publication_logo_url":"https://substackcdn.com/image/fetch/$s_!DbYa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b0838a-bd14-46a1-801c-b6a2046e5c1e_1130x1130.png","belowTheFold":false,"youtube_url":null,"show_links":null,"feed_url":null}">
Most GenMedia people will remember the BFL homepage when they initially launched Flux 1 in 2024, hinting at video models next, with their logo in a forest. Well, 2 years later, it’s finally real:

The blogpost outlines Self Flow, covering ALL their modalities together with strong preference claims:

“Its core capabilities include the following (all outputs come with native audio generation):
Text-to-video generation.
Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references.
Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context.
Generative video-audio continuation from input video and audio.
Keyframe-to-video generation for controlled transitions between defined moments.
Multilingual dialogue.
A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output.
Agentic chaining of individual clips into longer, multi-shot sequences.
High style diversity -- FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.
Strong typography generation and animated designs.”
Some of the above are SOTA features from other frontier lab models, like we discussed in our Grok Imagine pod, so the community has very much been put on notice that there has now been independent, perhaps SOTA, reproduction of these capabilities, with an open weights Dev version on the way.
As if this release wasn’t enough, the team also announced FLUX3-mimic, which proves that the FLUX 3 model is learning a sufficient world model capable of driving robots…
@mimicrobotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous","username":"bfl_ai","name":"Black Forest Labs","profile_image_url":"https://pbs.substack.com/profile_images/1954888731053142016/NDyG-4-j_normal.jpg","date":"2026-07-23T15:08:16.000Z","photos":[],"quoted_tweet":{},"reply_count":2,"retweet_count":7,"like_count":138,"impression_count":14471,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM">
… and predicting their impact in real factory settings…

AI News for 7/22/2026-7/23/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Open Code, Open Models, and the Policy Fault Line Around Distillation
The Stack v3 is the day’s most consequential open-data release: @anton_lozhkov announced The Stack v3, now the largest open code dataset publicly released: 114 TB raw, 224M repositories, 44B files, 770 languages, and roughly 5T deduplicated/filtered tokens. Relative to v2, the filtered corpus jumps from ~550B to ~5T tokens, with especially large gains in C++ (x15), TypeScript (x7.5), Rust (x7), and Python (x4.8). The notable operational changes are that v3 ships contents inline rather than Software Heritage IDs, includes a fresh GitHub recrawl through Aug 2025, excludes restrictively licensed code, and offers both a ready-to-train split and a full bucket for custom dedup/filtering. Hugging Face researchers framed it explicitly as infrastructure for the next generation of open code models and cyber-defense tooling: see @LoubnaBenAllal1, @lvwerra, and commentary from @eliebakouch noting prior Stack versions were used in many disclosed code-model training mixtures.
Distillation remains the live ideological fault line: several high-signal posts pushed back on attempts to sharply separate “internet-scale pretraining” from output-level distillation. @GergelyOrosz compared model inspection via prompting to reverse-engineering a competitor’s product, while @SchmidhuberAI emphasized distillation’s long lineage. @Suhail argued the practical response is not prohibition but stronger investment in open-weight domestic models, and @garrytan put it more simply: open weights are strategically important. The subtext across these posts is that open datasets like The Stack v3 materially raise the floor for every lab that wants to build competitive code models without relying on closed ecosystems.
Multimodal Frontier: FLUX 3, Robotics Transfer, and New Audio/TTS Systems
Black Forest Labs’ FLUX 3 expands the multimodal frontier beyond image/video: @bfl_ai launched FLUX 3, a unified multimodal model spanning image, video, audio, and action prediction, with early access for FLUX 3 Video and an explicit claim that the same architecture can be extended toward robotics. Team members connected it back to the earlier Self-Flow research, including @hila_chefer and @robrombach. What matters technically is the unified training story: not a loose family of specialized generators, but one architecture intended to bridge media generation and control.
mimic’s FLUX-mimic is a concrete robotics instantiation of that thesis: @mimicrobotics described FLUX-mimic as a Video-Action Model built on top of FLUX 3, trained on robot and wearable data for general-purpose dexterity and deployable on a single on-prem GPU. Their central claim is that better video world modeling transfers directly into robot control quality and sample efficiency; they’re already testing with Audi. This dovetails with @GeneralistAI, whose GEN-1 now supports varied end effectors and can adapt when the “hand” changes mid-rollout, reinforcing the idea that embodiment-general policies may come from conditioning on morphology rather than specializing per manipulator.
Audio saw two notable launches at opposite ends of the stack: @Alibaba_Qwen introduced Qwen-Audio-3.0-TTS in Flash and Plus variants, with 16 languages, inline control tags like [whisper] / [angry], natural-language style steering, noisy-reference robustness, and up to 3-minute one-pass generation; they also claimed the #1 spot on the Artificial Analysis TTS leaderboard. Separately, @HuggingApps highlighted WordVoice TTS, a smaller model with per-word control over duration, loudness, pitch, and tone—interesting less as a leaderboard play than as a control-surface experiment for audio tooling.
Agent Infrastructure: Harnesses, Dynamic Workflows, Programmatic Memory, and Benchmarks
The center of gravity is shifting from prompts to harnesses: multiple tweets converged on the same engineering thesis. @unclebobmartin described an “extreme constraints” workflow where trust comes from tests, QA, mutation testing, and metrics, not manual code review. @ThePrimeagen said he has become materially more positive on AI coding workflows, especially for large structural refactors. @TheTuringPost made the cleaner systems point: “graph engineering” is mostly old software architecture renamed, and most agents still do not need complex graphs unless workflows branch, verify, or require human approvals.
Several concrete harness/orchestration releases stood out: @omarsar0 summarized the Harness Handbook paper, which maps runtime behaviors to source locations and improved planning win rates for coding agents while reducing planner token use. The same author also described dynamic workflows as a generalized abstraction over loops/graphs/router patterns that can support model councils, advisor-judge-executor setups, and multi-backend orchestration across Claude/Codex/Hermes/etc. @witcheer shipped Hermes Profiles, effectively namespaced agent instances with separate memory, API keys, sessions, gateways, and export/import paths—pragmatic agent lifecycle infra rather than model novelty. @davidfowl also announced a new protocol underlying Microsoft’s VS Code agents app.
Memory and coordination are getting more formalized: @dair_ai highlighted PRO-LONG, a “programmatic memory” approach that stores full structured interaction histories and queries them like a database, outperforming bespoke long-horizon memory harnesses on ARC-AGI-3 with fewer tokens. @omarsar0 and @kimmonismus pointed to Offloop’s D1 dispatcher, a small model that decides which agent should speak next—or whether no agent should—addressing the familiar failure mode where multi-agent systems burn tokens by duplicating work.
Benchmarking is also evolving toward moving targets: @ryanmart3n launched Frontier-Bench, an ongoing community benchmark meant to evolve with frontier agent work beyond coding, while @CAIS released EnigmaEval, a harder reasoning benchmark where Claude Fable 5 and GPT-5.6 Sol lead and the hard set still only yields 10% for Fable 5. Together these reflect a broad dissatisfaction with static evals for fast-moving agent systems.
OpenAI Product Rollouts, Agent UX, and the Hugging Face Incident Fallout
The actual OpenAI release was product/UX, not GPT-6: after heavy speculation around “Opus 5” and a larger model drop from accounts like @kimmonismus and @theo, OpenAI’s shipped updates were more incremental but still meaningful for agent workflows. @OpenAI rolled out ChatGPT Voice in the desktop app for Plus/Pro/Business/Edu/Enterprise, powered by GPT-Live, with the ability to control the computer and coordinate work across ChatGPT Work and Codex. @OpenAIDevs added multi-folder Codex projects, and later Sites Analytics for published sites. Reactions were mixed: some found voice-driven multi-threaded coordination a genuine UX shift ([@reach_vb, @whoiskatrin]), while others thought the internal hype had implied something much larger ([@kimmonismus]).
Health in ChatGPT is a more strategically important rollout than it may first appear: @OpenAI, @ChatGPTapp, and @thekaransinghal announced U.S. rollout of Health in ChatGPT, allowing users to connect Apple Health and supported medical records. The notable implementation claims: connected health data receives additional encryption, is not used to train foundation models or target ads, and the feature builds on substantial physician review effort. This is less about a new model and more about a new high-trust application layer on top of existing model capability.
The Hugging Face hacking incident continues to dominate safety discourse: @johnschulman2 called for transcript release to understand whether the top-level agent knowingly pursued the hack or whether value drift emerged through subagents. @RyanGreenblatt, @jachiam0, and @Thom_Wolf pushed on broader lessons: internal AI-agent security differs from standard external threat models; offensive cyber-capable models may be especially vulnerable to adversarial reversal; and the irony is that the first public autonomous attack narrative featured a closed model attacking while open infrastructure became part of the defense response.
Inference, Serving, and the New Efficiency Arms Race
Etched’s scale-up is the clearest capital/infra announcement of the day: @Etched raised $300M Series C at a $10.3B valuation to accelerate inference-cluster production and opened an 80,000 sq ft / 10 MW facility near its office. The messaging is explicit: not training frontier models, but “run the world’s inference.” Supportive commentary from infra operators and investors suggests real interest in the chip-side inference specialization thesis, e.g. @willdepue and @juberti.
Model efficiency and serving architecture remain a battleground: @ArtificialAnlys noted that OpenAI’s GPT-5.6 Sol effort settings dominate much of the current token-efficiency Pareto frontier, while @CoreWeave posted a provider-speed benchmark for MiniMax M3 with 357 output tok/s and low blended price. On the open-serving side, @vllm_project described trillion-scale agentic RL inference plumbing in prime-rl 0.6.0 on vLLM—FP8, expert parallelism, prefill/decode disaggregation, KV offload, and routing—used to train GLM-5 on SWE tasks at 131k sequence length with sub-5-minute steps on 28 H200 nodes. That post is one of the more useful glimpses into how modern RL/agent training and serving stacks are being fused.
Top Tweets (by engagement)
ChatGPT Voice desktop rollout: @OpenAI shipped desktop voice control for ChatGPT Work and Codex, likely the biggest pure product launch by reach.
OpenWorker: @AndrewYNg launched an open-source, model-agnostic local agent for files and workplace tools.
Health in ChatGPT: @OpenAI / @ChatGPTapp rolled out connected health context for U.S. users.
FLUX 3: @bfl_ai launched a unified image/video/audio/action-prediction model with obvious downstream robotics implications.
The Stack v3: @anton_lozhkov released the largest open code dataset yet, a foundational input to future code-model competition.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
- Open-Weight AI Geopolitics and Government Deployment
Sanctions on Open Source. hope they don’t do anything stupid here. (Activity: 2278): The image is a screenshot of an X post attributed to Treasury Secretary Scott B. warning that while the U.S. supports open-source AI, it may consider sanctions and Entity List designations if open-source releases enable alleged PRC “covert, industrial-scale distillation attacks” and theft of American IP (image). In the Reddit context, the technical concern is whether model distillation from open or accessible frontier models could be treated as sanctionable IP theft, potentially chilling open-weight/model releases and downstream research. Commenters are skeptical and sarcastic, suggesting such sanctions could “backfire” or be technically hard to justify. One commenter disputes the implied timeline by noting Fable5 released July 1 and Kimi K3 was announced July 15, implying that claiming a Fable-level distillation in 15 days would be implausibly fast.
A commenter challenges the implied distillation/IP-theft timeline by noting Fable5 was released on July 1, while Kimi K3 was announced on July 15; they argue that producing a comparable distilled model in only 15 days would be unusually fast, implying the accusation may be technically implausible without stronger evidence.
DeepSeek Founder’s 4-hour investor meeting: DeepSeek is prioritizing AGI over user growth and commercialisation (Activity: 1030): A translated Chinese report of DeepSeek founder Liang Wenfeng’s reported 4-hour investor meeting says the lab is explicitly optimizing for AGI probability over near-term commercialization/user growth, treating products, hallucination mitigation, multimodality, and vertical agents as secondary to coding agents → continual learning → AI self-iteration → embodied intelligence. Liang reportedly committed that DeepSeek’s open-source releases are the same models it deploys internally, not degraded variants, and argued the China–US gap is mainly compute/resources rather than talent, while reaffirming belief in scaling: “larger scale undoubtedly produces better results.” Strategically, DeepSeek claims it will avoid super-app ambitions, video/3D/world-model work, and profit-maximizing API pricing, emphasizing low-cost architectures, open source, and team stability as mechanisms to improve its odds of reaching AGI. Commenters were mostly enthusiastic about the candor and open-source stance. One geopolitical take argued that if Chinese labs sustain an open-source AI strategy, US profit-driven labs like OpenAI/Anthropic may need either regulatory exclusion of Chinese models or a persistent technical lead large enough to offset rapid catch-up.
A commenter questioned the core technical premise behind DeepSeek’s AGI prioritization: despite steady model improvements, they argue it remains unclear whether current LLM-style scaling and training approaches can actually lead to AGI, saying “AGI itself does not seem closer currently than it was before.” This frames the investor-meeting strategy as dependent on an unresolved research assumption rather than just execution or commercialization speed.
One discussion point focused on the competitive implications of China-backed/open-source AI versus profit-driven U.S. labs. The commenter argued that if Chinese labs continue releasing strong open models, U.S. companies may need either regulatory exclusion of Chinese models or a sustained technical lead from OpenAI/Anthropic large enough that Chinese competitors remain ~1 year+ behind each generation.
Austria is rolling out a government AI-platform using Mistral models and Open WebUI (Activity: 592): The image shows Austria’s GovGPT web UI labeled as an AI workspace for “Texte und Dokumente,” matching reports that the platform uses Open WebUI as the frontend and Mistral open-weight models on sovereign BRZ federal datacenter infrastructure. Per the post’s sources, the rollout targets roughly 180,000 Austrian federal employees, with use cases including free chat, document summarization, document Q&A, internal knowledge bases, electronic-file analysis, parliamentary requests, and later agentic workflows—making it a notable real-world public-sector deployment of open-weight LLMs. Comments were split between jokes and practical support: one technical commenter argued the system could be very useful if connected to government documents because LLMs perform well with retrieved context, while an Austrian commenter framed it as a strong proof-of-concept that can later swap in stronger or fine-tuned models.
A commenter argued the platform’s main value will come from retrieval/context grounding rather than the base model’s parametric knowledge: if Austria indexes “all the government documents behind it,” an LLM could help citizens navigate procedures and forms more effectively than relying on training data alone.
An Austrian commenter framed the rollout as a proof of concept for locally hostable/public-sector AI, noting that the backend could later be swapped for stronger or fine-tuned models. They emphasized that even a “modest model” may yi
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み