LTX、NVIDIA GPU対応のオープンウェイト世界モデル「LTX-2.5」を公開
本文の状態
日本語全文を表示中
詳細モードで約10分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
LTX は NVIDIA RTX GPU で動作するオープンウェイト世界モデル LTX-2.5 を発表し、クラウド依存からローカル生成への転換を促す。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 05:20
AI深層分析
キーポイント
ローカル環境での高品質生成の実現
LTX-2.5 は NVIDIA RTX GPU や DGX Spark で動作するように最適化され、VRAM 要件を削減して既存のハードウェアでフロンティア級のモデルを実行可能にした。
制作ワークフローの変革
キャラクターの一貫性を保つネイティブマルチショット生成や、高運動シーンでのアーティファクト低減により、スタジオやクラウドなしでデスクトップ上で本格的な動画制作が可能になった。
圧倒的な生成速度の向上
2x NVIDIA GB200 環境では 10 秒分のクリップを 6.8 秒で生成し、競合クローズドモデルよりも最大約 58 倍高速な処理を実現した。
コストと柔軟性の向上
追加の生成費用やクレジット制限がないため、クリエイターは多様なバリエーションを低コストで試作し、広告疲労を防ぐための迅速なリフレッシュが可能になる。
生成速度の劇的な向上
LTX-2.5はオンプレミス環境で10秒分の動画を6.8秒で生成し、競合他社よりも最大58倍高速である。この速度差により、夜間バッチ処理や迅速なA/Bテストが理論上の話から実用的な運用へと移行する。
重要な引用
The entire production stack now fits on a single desktop.
LTX-2.5 puts something in creators’ hands that used to sit behind a studio door: real consistency.
On-prem, LTX-2.5 generates faster than the clip’s own runtime.
On-prem, LTX-2.5 generates faster than the clip's own runtime, 7.6x faster than the nearest closed alternative and roughly 58x faster than the slowest.
編集コメントを表示
編集コメント
LTX-2.5 の発表は、生成 AI の利用を「クラウドの専用リソース」から「個人が所有するデスクトップ」へと民主化する重要な転換点となる。特に速度と一貫性の向上により、広告クリエイティブや短期動画制作の現場におけるワークフロー再構築が加速すると予想される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
動画制作の現場では、ソーシャルメディア用のクリップ、広告クリエイティブ、映画のプロットビジュアライゼーションなどがクラウドからローカルの GPU へ移行しつつあります。LTX はこの変化に対応するため、「LTX-2.5」をリリースしました。これは動画生成やリアルタイムアプリケーション、物理 AI 向けのオープンウェイト型ワールドモデルです。
LTX は同モデルを NVIDIA RTX GPU や NVIDIA DGX Spark でのローカル推論に最適化し、VRAM の必要量を削減しました。これにより、最先端のワールドモデルがクリエイターがすでに所有するハードウェア上で動作可能になりました。今回のリリースは、NVIDIA が同日にオープンウェイトの「Nemotron 3.5 Lightning」エージェントモデルを発表したことに合わせて開始された、1 ヶ月間のローカル AI シリーズの柱となっています。両者の発表から読み取れるのは、「オープンモデルをローカルで加速させること」が、もはや標準的な制作インフラになりつつあるという事実です。
ローカル生成がクリエイターにもたらす変化
LTX-2.5 は、これまでスタジオの奥深くにしか存在しなかったものをクリエイターの手に届けました。それが「一貫性」です。ネイティブなマルチショット生成機能により、シーケンス全体を一つのまとまりとしてレンダリングできます。これによってキャラクターの外見がカットごとに維持され、以前のオープンモデルで問題となっていた不具合(グリッチ)が解消されました。その結果、キャンペーン制作に耐えうる品質が実現しています。
さらに、より鋭敏な「Gemma 4」言語バックボーンと、高運動量のショットにおけるアーティファクトを削減する新しいデコーダーを搭載したことで、出力はポストプロダクションの仕上げに近いレベルに達しました。これらすべてが、コンシューマー向けの NVIDIA RTX GPU 上で ComfyUI 内に直接実行可能です。
デスクに一人いるクリエイターでも、素早い LoRA の微調整を行うだけで、ブランドキャラクターや独自のスタイルを確立できます。スタジオは不要です。クラウドも不要です。知的財産が機械から外部へ流出することもありません。
これが本当の転換点です。制作に必要なすべての工程が、今やデスク一台に収まるようになりました。かつては撮影クルー、撮影日、レンダリングファーム、そしてクラウド利用料が必要だったものが、すでに搭載されている RTX グラフィックボード上で完結するのです。
追加のクリップ生成にも、従来型の課金や使用量制限はありません。これによりクリエイターの働き方が根本から変わります。さまざまな実験を大胆に行い、一つの安全なアイデアに賭けるのではなく、十数通りの方向性を追求できます。GPU でバッチ処理して一晩で週分のコンテンツを生成すれば、朝には選択肢が並んだフォルダが開かれているでしょう。
常に新しい素材が必要とされるショート動画クリエイターや広告チームにとって、これは画期的な変化です。広告疲労は通常 7〜10 日で発生するため、ボトルネックだったのはアイデアの欠乏ではなく、十分な量を制作するためのコストと時間でした。ローカルでの生成ならそれが解消されます。同じブリーフからバリエーションを瞬時に作り出し、十通りのフックを検証し、五つの市場向けにローカライズして、疲労が訪れる前にクリエイティブを更新できるのです。
ソロクリエイターや小規模チームでも、デスク上の RTX GPU 一台で、フルスタジオ並みのアウトプット量を達成できるようになりました。
速度:数字から見るその背景
生成速度が遅ければ意味がありませんが、LTX-2.5 はその点で圧倒的です。LTX が公開した画像から動画への生成ベンチマークによると、2 台の NVIDIA GB200 を使用してオンプレミスで実行した場合、10 秒分のクリップ生成に要する時間はわずか 6.8 秒です。一方、LTX API を経由する場合でも 23.7 秒と高速です。
これに対し、最も速いとされるクローズドソースの競合モデルである Omni Flash、Grok 1.5、Veo 3.1 でさえも、生成には 52 秒から 70 秒を要します。さらに遅いシステムになると、Seedance 2.0 が 196 秒、FLUX 3 が 259 秒、Seedance 2.5 が 317 秒、Kling 3.0 Pro はなんと 398 秒にも及びます。オンプレミス環境における LTX-2.5 の生成速度は、生成対象のクリップ自体の再生時間よりも速く、最も近い競合モデルより約 7.6 倍、最速の遅いシステムからは約 58 倍も高速です。この圧倒的な差により、夜間の一括生成や迅速な A/B テストが理論上の話ではなく、現実的な運用として可能になります。
NVIDIA のローカル AI モメンタム
8 月を通じて、NVIDIA はローカル AI エコシステム全体におけるモデル、アプリケーション、ツールに注目を集めています。本日同時にリリースされた「Nemotron 3.5 Lightning」は、常時稼働するエージェント向けのオープンな 30B マルチエキスパートモデルです。また、各エージェントのワークフローステップを最適なモデルへルーティングするオープンソースライブラリ「NeMo Switchyard」も加わりました。これらの共通点はハードウェアの選択にあります。NVIDIA エコシステムに属するオープンモデルは、RTX 搭載 PC からワークステーション、データセンター、クラウドまで幅広くスケーラブルです。LTX-2.5 は、クリエイター、開発者、ロボットチーム向けの NVIDIA アクセラレーション対応の世界モデルとして、このストーリーに直接組み込まれます。
LTX-2.5 とは何か?
言語モデルが次の単語を予測するのに対し、世界モデルは「次の瞬間」を予測します。環境を生成し、その振る舞いをシミュレーションし、ユーザーがその中で行動できるようにします。この基盤は映画制作、広告、ゲーム、シミュレーション、そして倉庫や工場のロボット制御を支えています。
LTX は同シリーズを「最も利用されているオープンな世界モデル」と位置づけ、ダウンロード数は 3,300 万回を超えると発表しています。その中でも LTX-2.5 は、これまでのリリースの中で最も能力が高いとされています。重み(weights)が公開されているため、チームはハードウェアの選定からカスタマイズ、知的財産権に至るまで完全なコントロールを握ることができます。
アーキテクチャにおける新機能
LTX は既存のコアに機能を追加するのではなく、生成パイプラインのほぼすべての段階を再構築しました。
- 新しい拡散ビデオデコーダー: 高運動シーンでの視覚的アーティファクトを低減しつつ、LTX の高い圧縮率を維持し、既存の映像素材にも忠実な表現を実現します。
- ネイティブ・マルチショット生成: 一連のシーケンスを一つの出力としてレンダリング。カットを跨いでキャラクター、シーン、音声の一貫性を保ちます。複雑で複数の被写体を含むプロンプトの理解度を高めるため、カスタムされた Gemma 4 の言語バックボーンと専用のプロンプト強化機能が採用されています。
- 拡散忠実度レンダリング: 8 倍に時間圧縮された潜在空間で動きと構造を構築し、視覚的な詳細を固定するための高忠実度キーフレームを生成します。キーフレームの数は、シーンの複雑さと計算リソースの予算に応じて動的に調整されます。
- 物理 AI チェックポイント: ロボティクス向けに事前学習されたチェックポイントは、映画映像とは異なるドメインデータでのファインチューニングのための基盤を提供し、チームが活用できます。
より強力な蒸留モデル:生産環境での展開に必要なコストを削減し、推論速度を向上させながら、同等の品質を実現します。
LTX-2.5 が向いているユーザー
映画・映像スタジオ:複数ショットの一貫性と洗練されたデコーダーにより、実際の制作現場で使えるシーケンスが生成可能です。Asteria 社のようなスタジオではすでに LTX を活用してオリジナルフィルムや動画の制作を行っています。
ショートフォームクリエイターと広告チーム:1 クリップごとの課金がないローカル生成により、A/B テストを戦略的に実施できます。夜間にバッチ処理でバリエーションを作成し、週次でクリエイティブを更新し、予算をかけずに市場ごとにローカライズが可能です。
リアルタイムアプリケーション開発者:Reactor は低遅延インフラ上で LTX-2.5 を実行し、インタラクティブアバター、ライブワールド、リアルタイムロボティクスワークロードを駆動します。
ロボット工学および物理 AI チーム:物理 AI チェックポイントは、映画以外のドメインデータに対するファインチューニングのベースとなります。Markov Robotics は LTX を活用して、物理システムが世界をどのように知覚し、移動するかを開発しています。
利用可能性とライセンス
LTX-2.5 は、Hugging Face でオープンウェイトとして提供され、ComfyUI にもネイティブ対応しており、管理された生成には LTX API を通じて利用可能です。データセンターの GPU から Mac まであらゆる環境で動作し、年間収益が 1,000 ドル以下の組織は無料で利用できます。コードとドキュメントは GitHub に公開されています。
主なポイント
LTX-2.5 は、動画・リアルタイムアプリ・ロボット工学向けのオープンウェイト型ワールドモデルです。
オンプレミスでの生成では、10 秒のクリップを 6.8 秒で処理可能ですが、競合するクローズドなモデルでは 52〜398 秒かかります。
NVIDIA の最適化により、RTX GPU や DGX Spark 上でのローカル推論に必要な VRAM が削減されました。
Gemma 4 をバックボーンに採用したネイティブのマルチショット機能により、キャラクターの一貫性が保たれ、ローカル環境でもキャンペーン品質の出力が可能になりました。
重み(Weights)は Hugging Face で無償公開されており、年間収益が 1,000 万ドル未満の企業であれば利用可能です。また、リリース初日から ComfyUI への対応も完了しています。
本記事では NVIDIA チームのリーダーシップとリソース提供に感謝いたします。なお、本記事は NVIDIA のスポンサーシップにより作成されました。
この記事は MarkTechPost に掲載された「The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model」の翻訳です。
原文を表示
Video production is shifting as social clips, ad creative and film pre-visualization move from cloud to local GPUs. LTX today released LTX-2.5, an open weights world model for video generation, real-time applications, and physical AI, built for exactly that shift. LTX optimized the model for local inference on NVIDIA RTX GPUs and NVIDIA DGX Spark, cutting VRAM requirements so a frontier world model runs on hardware creators already own. The release anchors NVIDIA’s month-long local AI series, launched the same day as its open Nemotron 3.5 Lightning agent model. The signal from both: open models, accelerated locally, are becoming default production infrastructure.
What Local Generation Changes for Creators
LTX-2.5 puts something in creators’ hands that used to sit behind a studio door: real consistency. Native multishot generation renders a whole sequence as one coherent piece, holding a character’s look shot to shot, fixing the glitching that made earlier open models unusable for campaigns. Add a sharper Gemma 4 language backbone and a new decoder that cuts artifacts in high-motion shots, and the output is close to post-ready. It all runs on a consumer NVIDIA RTX GPU, straight inside ComfyUI. One person at a desk can lock a branded character or signature style with a quick LoRA fine-tune. No studio. No cloud. No IP leaving the machine.
That is the real shift: the entire production stack now fits on a single desktop. What used to take a crew, a shoot day, a render farm, and a cloud bill now happens on the RTX card already in the machine. Additional clips carry no per-generation fees or metered credits. That rewires how creators work: experiment widely, chase a dozen directions instead of betting on one safe idea, and let the GPU batch-generate a week of content overnight. You wake up to a folder full of options.
For short-form creators and ad teams on constant refresh, that is transformational. Ad fatigue commonly sets in within 7 to 10 days, so the bottleneck was never ideas; it was the cost and time of producing enough of them. Local generation erases it: spin up variations on the same brief, test ten hooks, localize for five markets, and refresh creative before fatigue arrives. Solo creators and small teams can now match the output volume of a full studio with one RTX GPU on a desk.
Speed: The Numbers Behind the Story
None of this matters unless generation is fast, and it is. In LTX’s published image-to-video benchmark, a 10-second clip takes 6.8 seconds on-prem running on 2x NVIDIA GB200 and 23.7 seconds via the LTX API. The fastest closed alternatives listed, Omni Flash, Grok 1.5, and Veo 3.1, land at 52 to 70 seconds. Slower systems stretch far beyond:Seedance 2.0 at 196, FLUX 3 at 259, Seedance 2.5 at 317, and Kling 3.0 Pro at 398. On-prem, LTX-2.5 generates faster than the clip’s own runtime, 7.6x faster than the nearest closed alternative and roughly 58x faster than the slowest. That gap makes overnight batch generation and rapid A/B iteration practical, not theoretical.
NVIDIA’s Local AI Momentum
Throughout August, NVIDIA is spotlighting models, applications, and tools across the local AI ecosystem. Nemotron 3.5 Lightning, also released today, is an open 30B mixture-of-experts model for always-on agents, joined by NeMo Switchyard, an open source library that routes each agent workflow step to the best-fit model. The common thread is hardware choice: NVIDIA-ecosystem open models scale from RTX PCs to workstations, data centers, and cloud. LTX-2.5 slots directly into that story as an NVIDIA-accelerated world model for creators, developers, and robotics teams.
What is LTX-2.5?
Where large language models (LLMs) learn to predict the next word, world models learn to predict the next moment. They generate environments, simulate how they behave, and let users act inside them. That foundation supports film, advertising, gaming, simulation, and robots in warehouses and factories. LTX describes the LTX family as the most used open world model, with more than 33 million downloads, and positions LTX-2.5 as its most capable release yet. Open weights give teams full control of hardware, customization, and IP.
What’s New in the Architecture
LTX rebuilt nearly every stage of the generation pipeline rather than bolting features onto an older core:
New diffusion video decoder: Reduces visual artifacts in high-motion scenes while preserving LTX’s high compression ratio and staying true to existing footage.
Native multishot generation: Renders a full sequence as one output, holding character, scene, and voice consistent across cuts. A custom Gemma 4 language backbone and dedicated prompt enhancer improve comprehension of complex, multi-subject prompts.
Diffusion Fidelity Rendering: Builds motion and structure in an 8x temporally compressed latent space, then generates high-fidelity keyframes to anchor visual detail. Keyframe count adapts to scene complexity and compute budget.
A physical AI checkpoint: A pretrained checkpoint tuned for robotics gives teams a base for fine-tuning on domain data unlike cinematic video.
A stronger distilled model: Delivers the same quality at lower cost and faster inference for production-volume deployment.
Who LTX-2.5 is For
Film and video studios: Multishot consistency plus the cleaner decoder make sequences usable in real productions. Studios like Asteria already produce original film and video on LTX.
Short-form creators and ad teams: Local generation with no per-clip fees turns A/B testing into a strategy: batch variations overnight, refresh creative weekly, and localize across markets without a production budget.
Real-time application developers: Reactor runs LTX-2.5 on its low-latency infrastructure to power interactive avatars, live worlds, and real-time robotics workloads.
Robotics and physical AI teams: The physical AI checkpoint provides a fine-tuning base for non-cinematic domain data. Markov Robotics uses LTX to develop how physical systems perceive and move through the world.
Availability and Licensing
LTX-2.5 ships as open weights on Hugging Face, natively in ComfyUI, and through the LTX API for managed generation. It runs on anything from data center GPUs to a Mac and is free for organizations under $10M in annual recurring revenue. Code is on GitHub, with documentation.
Key Takeaways
LTX-2.5 is an open weights world model for video, real-time apps, and robotics.
On-prem generation hits 6.8 seconds for a 10-second clip, versus 52 to 398 seconds for closed rivals.
NVIDIA optimization cuts VRAM requirements for local inference on RTX GPUs and DGX Spark.
Native multishot with a Gemma 4 backbone holds characters consistent, making campaign-grade output possible locally.
Weights are free on Hugging Face under $10M ARR, with day-one ComfyUI support.
Thanks to the NVIDIA team for the thought leadership / resources for this article. This article is sponsored by NVIDIA.
The post The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み