LTX-2.5、画像から 10 秒動画を 6.8 秒で生成するオープンウェイトモデル
本文の状態
日本語全文を表示中
詳細モードで約18分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
LTX は新モデル LTX-2.5 を発表し、NVIDIA 製スーパーチップ上で 10 秒の AI 動画生成を 6.8 秒で完了する速度を実現した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 22:07
AI深層分析
キーポイント
驚異的な推論速度の実現
同社によると、NVIDIA のスーパーチップ上で 10 秒の 720p 画像から動画への変換を 6.8 秒で完了する性能を達成し、これはリアルタイム処理よりも高速である。
生成パイプラインの完全再構築
新機能は既存モデルへの追加ではなく、生成パイプラインのほぼ全段階が再設計されており、高画質化とアーティファクト削減を両立する新しい拡散デコーダーを採用している。
マルチショットと物理 AI 対応
カット間でのキャラクターやシーンの一貫性を保つネイティブなマルチショット生成機能を追加し、さらにロボット工学向けに微調整可能な事前学習済みチェックポイントも提供している。
コスト削減とオープンウェイト戦略
圧縮モデルの改良により比較モデルに対し約 1/8 のコストで実行可能となり、Hugging Face や ComfyUI を通じてオープンウェイトとして公開され、年間収益 1000 万ドル以下の組織は無料で利用可能である。
ハードウェア要件と速度の差
6.8 秒という高速処理は NVIDIA GB200 2 基による自己ホスト環境での測定結果であり、同社 API を利用した場合は 1080p レンダリングで 23.7 秒かかる。
重要な引用
"We're trying to maintain the same efficiency and the inference speed that we're known for, but constantly pushing the quality up."
"LTX-2.5 generates a 10-second, 720p image-to-video clip in 6.8 seconds faster than real time."
"With video models, world models, there are so many different use cases that require people to get access to the weights and create flows that really work for them."
"Our answer is open weights with licenses that allow individuals and companies below a certain amount of revenue to use the model for free, and once they're successful, to come up with some kind of licensing agreement with us."
編集コメントを表示
編集コメント
LTX-2.5 の発表は、動画生成モデルの速度とコスト効率における新たな基準を示すものである。特にローカル環境での実用性を高める最適化と、オープンウェイトとしての戦略的展開は、開発者コミュニティに大きな影響を与えるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Lightricks からスピンアウトしたオープンワールドモデル企業の LTX は本日、オープンウェイトの動画および「世界」モデル最新バージョンである LTX-2.5 をリリースしました。このモデルは、両社の戦略的な同日ローンチパートナーシップにより、オープン生成メディアのプロトタイピング環境として事実上の標準となっているノードベースワークフローツール ComfyUI にネイティブ統合されています。
同モデルは現在、Hugging Face 上でオープンウェイトとして公開されており、ComfyUI 内や LTX API を通じて利用可能です。管理された生成を希望するチーム向けです。年間再発収が 1,000 ドル未満の組織は無料で利用でき、大企業はライセンス契約を個別に交渉します。LTX によると、同社のモデルは累計 3,300 万回のダウンロードを突破しており、「オープンワールド」モデルラインとしては市場で最も広く使われています。
今回のリリースに先立ち、VentureBeat は LTX の共同創業者兼 CEO ズー・ファルブマン氏と ComfyUI の共同創業者兼CEO ヨランド・ヤン氏に独占インタビューを行いました。リリース内容やパートナーシップの背景、そしてなぜ両社がクローズドな API ではなくオープンウェイトこそが動画および世界モデル市場で勝利すると確信しているのかについて話を伺いました。
「私たちはこれまで培ってきた推論速度と効率性を維持しつつ、常に品質を向上させるよう努めています」とファルブマン氏は語りました。「今回のリリースでは多ショット対応や、高画質化のための拡散デコーダー、新しい条件付けモードなどを導入しました。また、リアルタイムユースケースやロボティクスに不可欠な自己回帰モデルへのサポートも強化しています。」
LTX-2.5 の新機能
同社の発表によると、LTX-2.5 は既存のコアに新機能を追加するのではなく、生成パイプラインのほぼ全工程を再構築しています。今回の主な変更点は以下の通りです。
高画質のモーション映像における視覚的アーティファクトを低減し、文字や顔などの細部を復元しながらも、LTX 特有の高い圧縮率を維持する新しい拡散ビデオデコーダーを採用しました。
カット間を通じてキャラクター、シーン、声を一貫して保ち、個別に生成されたショットをつなぎ合わせるのではなく、シーケンス全体を単一の出力としてレンダリングするネイティブ・マルチショット生成機能を搭載しています。
複雑な多被写体プロンプトへの対応精度を高めるため、カスタム設計の Gemma 4 言語バックボーンと専用プロンプトエンハンサーを導入しました。
物理 AI やロボティクス向けに事前学習されたチェックポイントを提供し、映画のような映像とは異なるドメインデータでのファインチューニングを可能にする基盤となっています。
コストを抑えつつ推論速度を向上させ、ほぼフルモデル並みの品質を実現する大幅に改善されたディストillation モデルを採用しました。また NVIDIA との最適化協力により、メモリ要件を削減して NVIDIA RTX GPU 上でのローカル実行も実現しています。
同社は、同等モデルと比較してコストが約 8 分の 1、レンダリング時間が約 7 分の 1 に短縮され、データセンター用 GPU から Mac まで幅広いハードウェアで動作できると主張しています。
LTX が謳う速度と品質について
LTX の発表資料の中で最も目を引く数値は速度です。同社によると、LTX-2.5 は 10 秒間、720p の画像から動画への変換を、6.8 秒というリアルタイムより速い時間で生成できるとのことです。
ただし、この数値にはハードウェアの条件が伴います。これは、NVIDIA の最上位チップである GB200 を 2 基使用した環境で、安定稼働状態を維持しながら自己ホスト測定した結果です。このような構成は、多くのチームが用意できるものよりも遥かに上級のものであり、LTX が提供する管理型 API を通じて同じタスクを実行した場合、1080p というより高い解像度でレンダリングされるため 23.7 秒を要しました(なお、API には 720p のプランは存在しません)。
同社が競合他社の API を同一タスクで測定したエンドツーエンドの結果によると、Google の Gemini Omni Flash は 52 秒、xAI の Grok 1.5 は 63 秒、Google の Veo 3.1 は 8 秒クリップ生成で 70 秒、MiniMax H3 は 180 秒、ByteDance の Seedance 2.5 は 317 秒、Kuaishou の Kling 3.0 Pro は 398 秒となりました。
画質については、LTX が盲検の並列比較テストの結果を公開しています。これは、どのモデルが生成したか評価者に知らされない状態で、同じプロンプトから作成された動画を評価者が投票する形式です。
このテストで LTX-2.5 は 67% の勝率を記録し、Seedance 2.5 の 65% をわずかに上回りました。その他は Gemini Omni Flash が 55%、MiniMax H3 が 50%、Seedance 2.0 が 44%、Wan 2.6 が 42%、FLUX 3 が 28% です。
これらの数値はいずれもベンダー報告値か LTX 自身が測定・委託したものであり、独立機関による検証は行われていません。同社は評価結果について「評価範囲が拡大するにつれて変化していくもの」と位置付けており、現時点では暫定的なデータとしています。これは購入者が自社のワークロードで実際にテストすべき方向性の示唆であり、確定されたランキングではありません。
発表資料では、純粋な性能数値よりも「導入のしやすさ」を強調しています。LTX-2.5 は VRAM 16GB 以上を搭載した GPU であればどこでも動作し、オンプレミスやエッジ環境での展開も可能。API を通じた利用も可能です。出力結果にブランドロゴの付与は必須ではなく、顧客独自のデータや知的財産に基づいたファインチューニングも自由に行えます。この柔軟性は、クローズドな API 専業モデルとの競合他社や、米国・欧州で重み(ウェイト)が入手できない、あるいはファインチューニングをライセンスで制限されているオープンライセンスのライバルたちとは対照的です。
API ビジネスモデルへの賭けに背いて
ファーブマン氏にとって、今回のリリースは業界がクローズドなモデルへ集約する動きに対する反動として始まった戦略の一環です。「Sora が登場した頃、大手企業がこぞって自社のモデルをクローズド化し始めたことに気づきました。API を通じて利用するだけでは、私たちが構築したいような業務には多くの企業で対応できないと判断しました」とファーブマン氏は語ります。
技術的な観点から彼は、動画や世界モデルは言語モデルよりもはるかに広い「ユースケースの範囲(サーフェスエリア)」を持っていると説明します。「LLM の API が扱うのは、質問への回答という狭い領域です。単語を入力して単語を受け取るだけですが、動画モデルや世界モデルでは、重みにアクセスし、自社の業務に最適化したワークフローを構築する必要があるユースケースが非常に多岐にわたります」と彼は述べています。
彼は、オープン化は慈善活動ではないと率直に語った。「私たちは明らかに慈善のためにこのことをしているわけではありません」と彼は述べた。「私たちの答えは、特定の収益規模以下の個人や企業が無料でモデルを利用でき、成功した後は当社とライセンス契約を結ぶことができるオープンウェイトの提供です。」
「私たちは、開発者が自信を持って構築できる基盤となるモデルを作ろうとしています」と彼は付け加えた。「我々はこう言っています。オープンウェイトは、私たちにとって一度きりの慈善活動の奇跡ではありません。これは戦略なのです。このモデルを届けるための正しい方法だと信じており、今後もこれを続けていきます。」
Facetune から世界モデルへ、そして標準となったノードグラフ
LTX は、Jerusalem に本社を置く Lightricks から生まれた同社が最もよく知られているのは、Facetune や Videoleap といった消費者向けクリエイティブアプリです。自己資金で立ち上げられ、黒字経営だった Lightricks は 2022 年に基盤モデルへ転換し、2024 年初頭に映画制作プラットフォーム「LTX Studio」をローンチしました。その後、2024 年 11 月に初のオープンウェイトモデルである LTX Video(LTXV)をリリースし、2025 年 5 月にはパラメータ数 130 億のバージョンを追加しました。
ファルブマンは、CTO のヤロン・インガーと CMO のニール・ポッテルと共に同社を共同設立しました。現在、LTX ブランドは世界モデル事業の前面に立ち、ニューヨーク、ロンドン、シカゴにオフィスを構えています。
ComfyUI は 2023 年 1 月、"comfyanonymous" と名乗る匿名の開発者によるオープンソースのサイドプロジェクトとして始まりました。Stable Diffusion のためにノードベースのグラフィカルインターフェースを開発し、ユーザーがモデルや処理ステップを連鎖させて、繰り返し使えるビジュアルワークフローを構築できるようにしたのです。以来、生成メディア分野で最も急速に成長するオープンソースプロジェクトの一つとなり、新しい画像・動画モデルのテストや組み合わせ、そして実装のための標準環境として定着しました。現在では Comfy Org という企業が支援しており、同社は開発継続のために 1700 万ドルを調達しています。共同創設者の Yan が CEO を務めています。
なぜ ComfyUI が企業導入の入り口となっているのか
モデル企業とツール企業がなぜ手を組んで発表するのかと疑問に思う読者に対し、Farbman は非常に率直な答えを返しました。ComfyUIこそが LTX の有料顧客の源泉だからです。
「当社の顧客の多くは Comfy での利用から始めます」と彼は語りました。「すでに業界で極めて人気のあるプロトタイピングシステムであり、多くの潜在顧客が Comfy 内のワークフローを確立した後に当社にやってきます。Comfy 上で既に動作しているため、Comfy 統合に対するゼロデイサポートを提供することは、いわば顧客獲得チャネルとしての重要性から当然の判断です。」
ヤン氏は、ComfyUI の役割をオープンエコシステムにおける接続層として説明しました。「Comfy はコアとなるレイヤーの上に位置し、ユーザーがローカルマシンでオープンウェイトモデルを推論したり、パートナーノードシステムを通じてクローズドなモデルにもアクセスしたりできるようにするものです。最終的には、これらすべてを統合したワークフローを構築することで、クリエイティブな分野からデータパイプライン、さらにはロボット工学のシナリオに至るまで、多様な用途を可能にします」と語りました。
ヤン氏はこのフライングホイール(好循環)こそが、オープンモデルを商業的に持続させる鍵であると主張しました。「私たちはこれらのモデルを広め、世界中へ普及させる手助けをしています。人々はその上でさまざまなワークフローやモデルの革新を行い、それがさらにスタジオやロボット研究所へと波及します。そうした企業はライセンスを取得し、得られた価値の一部を LTX や残りのエコシステムに還元することになります」。
企業が知っておくべきこと
両名の執行役員は、「動画生成モデルは映像クリップを作るためだけのもの」という前提に対して異議を唱えました。ファルブマン氏は、映画のような映像とは無関係な企業向けの実装事例を次々と挙げました。
「ハードウェアベンダーの中には、拡散モデルを用いた計算写真撮影を実現しようとしているところもあります。例えば、センサーから出力される生データ(通常はノイズが多い)のストリームを処理し、ノイズ低減の方法を探っているケースです」
「あるいは、VFX制作や水シミュレーション、昼間のシーンを夜景に変換する手法を検討している映像スタジオも同様です。アニメーションスタジオでは、アニメーターがキーフレームを作成し、システムがそれらを補間して動画化するパイプラインを効率化する方法を探っています」
「どこから着手すべきか迷っている企業にとっての推奨ルートは、自社の従業員がすでに経験した道です。ファルブマン氏は『多くの企業がすでに Comfy を導入しており、さらに追随するところも増えるでしょう』と述べています。
「Comfy は適切なレベルの構造を提供し、細部まで調整可能でありながら、複雑な部分を抽象化して隠蔽してくれます。企業は通常、社内でモデルや Comfy の操作を試した後に相談に来るものです」
ヤン氏は、すでに LTX やその他のオープンモデルを生産環境で運用しているスタジオや企業の間に見られる、一貫した 2 つの軌道パターンを説明しました。「彼らには研究部門や R&D、クリエイティブパイプラインがあり、何でも試すことができます」と彼は語りました。「時々、これらのパイプラインが十分に成熟して、何らかの生産環境へと昇格します。そしてその過程で、企業向けの議論が始まります。私たちが注力するのはツール作りであり、LTX の側ではライセンス問題が中心となります。」
重み(weights)がオープンであるため、この実験段階はすべて自社のハードウェア上で行うことができ、生成ごとの課金も不要です。また、データや知的財産が社外に流出することもないため、機密性の高い映像や独自キャラクター、規制対象データを扱う企業にとっては大きな違いとなります。商業的な導入のトリガーとなるのはスケールであり、年間収益(ARR)が 1,000 万ドルを超える組織にはライセンスが必要となります。
ヤン氏は、動きの遅い企業のリスクをより鮮明な言葉で示しました。「これはクリエイティブ業界全体を根本的に揺るがすトレンドです」と彼は語りました。「スタジオは必死にロードマップを探り、どうすれば先を行けるか、あるいは少なくとも AI の普及波に取り残されないかを模索しています。」
開発者、リアルタイムアプリ、そしてエッジ
ソフトウェア開発者にとって、今回のリリースは「リアルタイム」の物語が現実味を帯びてきたことを象徴するものです。ComfyUI と並んで、LTX は 2 つの新たなパートナーを発表しました。1 つ目は LTX を活用してオリジナル映画や動画を制作する AI フィルムスタジオ「Asteria」、もう 1 つは低遅延推論インフラ上で LTX-2.5 を実行し、インタラクティブアバターやライブワールド、リアルタイムロボティクスワークロードを動かす開発者向けプラットフォーム「Reactor」です。これにより、開発者は自前でインフラを整えることなく、本番環境で使えるリアルタイム体験を構築できるようになります。
ヤン氏は、オープンウェイトと低遅延がもたらす可能性を示す viral な事例として「Flipbook」を紹介しました。これは Reddit で話題となったインタラクティブな体験で、クリック可能な世界全体がその場で生成されるものです。「このインターフェースに表示されているすべては、LTX モデルを使ってリアルタイムストリーミングで生成されています」とヤン氏は説明します。「どこをクリックしても、そこですぐに新しいインタラクションが生成される環境、あるいは世界です。このような体験や実験は、オープンウェイトモデルが存在せず、LTX 特有のパフォーマンスがなければ実現しなかったでしょう。」
ファルブマン氏は、エッジでの効率性は副次的な結果ではなく、意図的な設計目標であると述べています。「私たちにとって重要なのは、消費者向けハードウェアや物理 AI の現場に近い場所で動作する、極めて効率的なモデルを作ることです」と彼は語り、業界の急速な進展にも言及しました。「最近は休暇を取るのも難しいほどで、あるモデルをリリースしている間に、すでに次のトレーニングに没頭しており、新しい論文が毎日発表されています。」
映画制作者向け:今すぐバーチャル・プロダクション、後から楽になる
プロの映画製作者やスタジオにとって、ヤン氏はリアルタイム・ワールドモデルが生産プロセスそのものの形を変えると見ています。撮影とポストプロダクションの間のギャップを縮めるのです。「最近では、リアルタイム・モデル、あるいはワールドモデルがバーチャル・プロダクションの一部としてスタジオで採用されるようになりました。これにより、撮影した直後にポストプロダクションの結果に近いものを見ることができます」とヤン氏は説明します。「プロデューサーや監督にとって、『これが欲しい』『これは違うからもう一度』と明確に指示できる体験を提供できます。かつてのハリウッドのプロセスは、複数の部署間を往復する必要があり、私には巨大な混乱に見えました。」
また、アニメーション、フォトリアリズム、ゲーム、3D、ロボット工学などにおいてモデルの専門性が大きく異なることを踏まえ、モデル同士の直接比較を過度に厳密に読み取るべきではないとも注意を促しました。「各モデルには単に異なる特性があるのです」と彼は語りました。「マイケル・ Phelps とマイケル・ジョーダンを比べるようなものです。どちらが優れたアスリートかという比較ではなく、ここでは専門分野が違うだけなのです。」
一方、ComfyUI の famously 陡峭な学習曲線に畏縮するアマチュアやインディークリエイターについては、ヤン氏はこのツールが彼らに寄り添う部分はあるが、あくまで一部に限られると率直に述べています。
「スキーに似ていますね」と彼は説明します。「Comfy を使えば降りやすいスロープはありますし、今後さらにそうしたスロープを増やしていければと考えています。しかし、真の技術者やプロのクリエイターが不可欠として必要とし、それなしでは生きられないような『ダブルブラックダイヤモンド』級の難易度のコースを廃棄することはありません。これが、モバイルアプリ型のクリエイティブツールとの決定的な違いなのです。」
LTX-2.5 は本日、Hugging Face で利用可能となり、ComfyUI にもネイティブ対応しています。また、LTX API を通じての利用も可能です。
原文を表示
LTX, the open world model company spun out of Lightricks, today released LTX-2.5, the newest version of its open-weights video and "world" model and it arrives natively integrated into ComfyUI, the node-based workflow tool that has become the de facto prototyping environment for open generative media, through a strategic day-one launch partnership between the two companies.
The model is available now as open weights on Hugging Face, inside ComfyUI, and through the LTX API for teams that want managed generation. It is free to use for organizations under $10 million in annual recurring revenue; larger companies negotiate a license. LTX says its models have passed 33 million downloads, making the LTX family the most-used "open world" model line on the market.
Ahead of the launch, VentureBeat spoke exclusively with LTX co-founder and CEO Zeev Farbman and ComfyUI co-founder and CEO Yoland Yan about the release, the partnership, and why both companies are betting that open weights — not closed APIs — will win the video and world model market.
"We're trying to maintain the same efficiency and the inference speed that we're known for, but constantly pushing the quality up," Farbman said. "We are introducing many cool things in this release: multi-shot support, a diffusion decoder for better quality, new conditioning modes, better support for autoregressive models that are critical for real-time use cases and robotics."
What's new in LTX-2.5
According to the company's announcement, LTX-2.5 rebuilds nearly every stage of the generation pipeline rather than bolting new capabilities onto an older core. The headline changes:
A new diffusion video decoder that reduces visual artifacts in high-motion footage and reconstructs fine detail like text and faces, while preserving LTX's high compression ratio.
Native multishot generation that renders a full sequence as a single output, holding character, scene, and voice consistent across cuts rather than stitching individually generated shots together.
A custom Gemma 4 language backbone and dedicated prompt enhancer for more accurate handling of complex, multi-subject prompts.
A pretrained checkpoint tuned for physical AI and robotics giving teams a base to fine-tune on domain data that looks nothing like cinematic video.
A substantially improved distilled model that delivers near-full-model quality at lower cost and faster inference, and, through an optimization effort with NVIDIA, runs locally on NVIDIA RTX GPUs with reduced memory requirements.
The company claims roughly one-eighth the cost and one-seventh the render time of comparable models, with output that runs on hardware ranging from data center GPUs down to a Mac.
How fast and how good LTX says it is
The most eye-catching number in LTX's launch materials is speed: the company says LTX-2.5 generates a 10-second, 720p image-to-video clip in 6.8 seconds faster than real time.
The caveat is the hardware behind it. That figure was measured self-hosted on two of NVIDIA's top-end GB200 chips at steady state, a configuration far beyond what most teams have racked; the same job through LTX's own managed API took 23.7 seconds, albeit rendered at the higher 1080p resolution (the API has no 720p tier).
By the company's end-to-end measurements of competing APIs on the same task, Google's Gemini Omni Flash came in at 52 seconds, xAI's Grok 1.5 at 63 seconds, Google's Veo 3.1 at 70 seconds (for an 8-second clip), MiniMax H3 at 180 seconds, ByteDance's Seedance 2.5 at 317 seconds, and Kuaishou's Kling 3.0 Pro at 398 seconds.
On quality, LTX shared results from blind, side-by-side human preference tests, in which evaluators voted on videos generated from the same prompt without knowing which model produced which.
LTX-2.5 recorded a 67% win rate, narrowly ahead of Seedance 2.5 at 65%, with Gemini Omni Flash at 55%, MiniMax H3 at 50%, Seedance 2.0 at 44%, Wan 2.6 at 42%, and FLUX 3 at 28%.
All of these figures are vendor-reported measured or commissioned by LTX itself, not independently verified and the company labels the preference results preliminary, noting it expects them "to evolve as evaluation expands." They are directional claims a buyer should test against their own workloads rather than settled rankings.
The launch materials also lean on deployment terms rather than raw performance: LTX-2.5 runs on any GPU with a minimum of 16GB of VRAM, deploys on-premises, at the edge, or via API, carries no mandatory branding on output, and can be fine-tuned on a customer's own data and IP flexibility the company contrasts with closed API-only rivals and with open-licensed competitors whose weights are unavailable in the U.S. and Europe or whose licenses restrict fine-tuning.
Betting against the API business model
For Farbman, the release is another installment in a strategy that began as a reaction to the industry's consolidation around closed models.
"We started with our own models out of necessity, because around the time that Sora came out, we realized that all the big guys are trying to close their models, and working through APIs just doesn't work for many businesses, including the kind of stuff that we wanted to build," he said.
The technical argument, he explained, is that video and world models have a fundamentally wider "surface area" of use cases than language models.
"With LLMs, the surface area of the API is pretty narrow, we're typically asking some kind of question, passing words and getting words back," Farbman said. "With video models, world models, there are so many different use cases that require people to get access to the weights and create flows that really work for them."
He was blunt that the openness is not charity. "We're definitely not doing this as philanthropy," he said. "Our answer is open weights with licenses that allow individuals and companies below a certain amount of revenue to use the model for free, and once they're successful, to come up with some kind of licensing agreement with us."
"We're trying to build a model that builders can confidently build upon," he added. "We're coming and saying: guys, open weights is not some kind of one-time philanthropic fluke for us. It's the strategy. We believe this is the right way to serve these models, and we're going to keep doing that."
From Facetune to world models and the node graph that became a standard
LTX grew out of Lightricks, the Jerusalem-headquartered company best known for consumer creative apps including Facetune and Videoleap. Bootstrapped and profitable, Lightricks pivoted to foundation models in 2022, launched its LTX Studio filmmaking platform in early 2024, and released its first open-weights LTX Video model (LTXV) in November 2024, following it with a 13-billion-parameter version in May 2025. Farbman co-founded the company alongside CTO Yaron Inger and CMO Nir Pochter, and the LTX brand now fronts its world model business, with offices in New York, London, and Chicago.
ComfyUI began in January 2023 as an open-source side project by a pseudonymous developer known as "comfyanonymous," who built a node-based graphical interface for Stable Diffusion that let users chain models and processing steps into repeatable visual workflows. It has since become one of the fastest-growing open-source projects in generative media the standard environment where new image and video models are tested, combined, and pushed into production and is now backed by a company, Comfy Org, which raised $17 million to keep developing the tool. Yan, a co-founder, serves as its CEO.
Why ComfyUI is the front door for enterprise adoption
For readers wondering why a model company and a tooling company are launching arm-in-arm, Farbman's answer was unusually candid: ComfyUI is where LTX's paying customers come from.
"A whole lot of our customers are starting their journey with Comfy," he said. "It's already this prototyping system that's extremely popular in the industry, and a lot of the potential customers are coming to us after they already figured out the flow inside Comfy. It's already working, so for us it's a no-brainer that we have to provide zero-day support for the Comfy integration, because it's basically our customer acquisition channel."
Yan described ComfyUI's role as the connective layer of the open ecosystem. "Comfy at the core is sitting as a layer on top, giving people accessibility to the open-weight models that people can inference on their local machine, or tap into closed models as well through our partner node system," he said. "In the end, [they] combine everything together into a workflow that empowers various things, from the creative side all the way to data pipeline and robotics type of scenarios."
That flywheel, Yan argued, is what sustains open models commercially: "We help promote and push these models into the world... people do all sorts of workflow and model innovation on top of it, and that further propagates these models into studios or robotics labs. Those companies would end up acquiring licenses and then contribute a part of the value gained back to LTX and the rest of the ecosystem."
What enterprises should know
Both executives pushed back on the assumption that a video model is only for generating videos. Farbman rattled off a list of enterprise deployments that have little to do with cinematic clips.
"We have hardware customers that are trying to figure out how to do computational photography with diffusion models, for example, taking a stream of raw pixels that are coming from the sensors, which is typically very noisy, and trying to figure out how to reduce noise there," he said. "Or think about the production studios that are trying to figure out how to do VFX, how to do water simulation, how to turn day into night. Or think about animation studios: they're trying to figure out how to streamline their pipeline, where animators are creating keyframes and then the system uses them as interpolation."
For enterprises weighing where to start, the recommended path is the one their own employees have probably already taken. "A lot of enterprises have already adopted Comfy, and I think many others will follow," Farbman said. "It gives this right level of structure, where you can tweak things a lot, but it still abstracts a lot of things away... Enterprises are typically reaching out after people internally have already played with the model, played with Comfy."
Yan described a consistent two-track pattern among studios and companies already running LTX and other open models in production. "They have their research, or R&D, creative pipeline, anything goes," he said. "Once in a while, some of these pipelines get good enough that they graduate into some kind of production environment. And somewhere along the line, the enterprise conversation gets started. On our end, it's more around tooling, and on the LTX side, it's more around the licensing."
Because the weights are open, that entire experimentation phase can happen on a company's own hardware, with no per-generation billing and no data or IP leaving its systems, a meaningful distinction for enterprises with sensitive footage, proprietary characters, or regulated data. The commercial trigger only arrives with scale: organizations above $10 million in ARR need a license.
Yan framed the stakes for slower-moving companies in starker terms. "This is a trend that is just fundamentally going to disrupt the entire creative industry," he said. "Studios are heavily trying to figure out what is the roadmap and how do we get ahead, sometimes not even get ahead, just how do we avoid falling behind the AI adoption wave."
Developers, real-time apps, and the edge
For software developers, the release leans into a growing real-time story. Alongside ComfyUI, LTX named two other launch partners: Asteria, the AI film studio producing original film and video on LTX, and Reactor, a developer platform that runs LTX-2.5 on low-latency inference infrastructure to power interactive avatars, live worlds, and real-time robotics workloads, so developers can build production-grade real-time experiences without standing up that infrastructure themselves.
Yan pointed to a viral example of what open weights plus low latency makes possible: Flipbook, an interactive experience that spread on Reddit in which an entire clickable world is generated on the fly. "Everything people see on that interface is generated using an LTX model, live-streamed," he said. "It's an environment, or a world, where anywhere you click, it just generates a brand-new interaction... That type of experience and experimentation wouldn't exist without an open-weight model, without LTX's type of performance."
Farbman said efficiency at the edge is a deliberate design target, not a side effect. "For us, it's very important to create an extremely efficient model that people can run on edge devices, both on consumer hardware and close to the edge with physical AI," he said, while acknowledging the relentless pace of the field: "These days, it's almost hard to take a vacation. Things are progressing so quickly that while you're releasing one model, you're already deeply into training another one, and new papers are coming on a daily basis."
Filmmakers: virtual production now, easier slopes later
For professional filmmakers and studios, Yan sees real-time world models changing the shape of production itself, collapsing the gap between shooting and post. "These days you see real-time models, or world models, getting adopted in studios as part of what's called virtual production, meaning you can shoot and then immediately get close to what the post-production result looks like," he said. "You give a much better experience to the producer or director to say, 'okay, this is what I want,' or 'this is not what I want let me actually reiterate.' Whereas before, the entire Hollywood pipeline is, in my opinion, a giant mess where it has to constantly go between multiple departments."
He also cautioned against reading head-to-head model comparisons too literally, given how differently models specialize across animation, photorealism, gaming, 3D, and robotics. "Various models have simply different characteristics," he said. "It's like comparing Michael Phelps with, I don't know, Michael Jordan. It's not really a comparison of who's a better athlete, there are just different specialties here."
As for amateur and indie creators intimidated by ComfyUI's famously steep learning curve, Yan was direct that the tool will meet them partway, but only partway.
"It's kind of like skiing," he said. "There are easy slopes that you can go down using Comfy, and hopefully we can create more and more of these easy slopes overall. But we'll never sacrifice the existence of the double-black-diamond type of lanes, because the real technical, professional creatives actually need and couldn't live without that type of core power. That's actually our core differentiator compared to a mobile-app type of creative tool."
LTX-2.5 is available today on Hugging Face, natively in ComfyUI, and through the LTX API.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み