中国のMiniMax、初のオープンAI動画モデル「H3」を公開しランキング首位に
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
中国のMiniMaxが330億パラメータのオープンモデルH3を公開し、Artificial Analysisのランキングで動画編集分野で初となる首位を獲得した。
AI深層分析を開く2026年8月4日 23:07
AI深層分析
キーポイント
オープンモデルとしての歴史的快挙
MiniMaxが公開したH3は、Video Editing分野で初めてオープンモデルとして1位にランクインし、Text-to-Videoでも2位、Image-to-Videoでも3位を獲得した。
多様なメディア統合と生成能力
同モデルはテキスト、画像、動画、音声を同時に処理でき、4〜15秒のステレオサウンド付きクリップを生成する。単一のプロンプトで最大9枚の参考画像や3本の動画・音声クリップを組み込める。
ローカル利用とライセンス制限
ComfyUIでのローカル実行は768pに制限され、2K解像度モジュールや中間形式変換機能(H3-Context-IR)は非公開である。商用利用は年間収益2,000万ドル未満の企業に限られる。
重要な引用
MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time.
Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video.
編集コメントを表示
編集コメント
MiniMax H3の登場は、動画生成領域におけるオープンソースモデルの実用性と性能が飛躍的に向上したことを示す明確な証拠である。ただし商用利用における収益制限や機能の一部非公開という点は、導入を検討する際に必ず確認すべき重要な制約条件となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
中国のテック企業 MiniMax が、動画生成モデル「H3」の重み(ウェイト)を公開し、オープンソースモデルとして初めて動画生成ランキングで首位を獲得しました。評価サイト「Artificial Analysis」によると、H3 は「Video Editing(動画編集)」部門で 1 位、「Text-to-Video(テキストから動画へ)」で 2 位、「Image-to-Video(画像から動画へ)」で 3 位にランクインしています。
このパラメータ数 330 億のモデルは、テキスト、画像、動画、音声を同時に処理し、ステレオサウンド付きで 4〜15 秒のクリップを生成します。モデルカード(MiniMax-H3)によると、プロンプトには最大 9 枚の参考画像、3 クリップの動画、3 クリップの音声を組み込むことが可能です。
*Video by MiniMax H3
一方で、2 つの機能は非公開のままです。解像度 2K のモジュールと、プロンプトや参照素材を構造化された中間形式に変換する「H3-Context-IR」は含まれていません。ComfyUI で H3 をローカル実行する場合、最高解像度は 768p に制限され、コンテキストの事前準備は MiniMax が公開しているプロンプティングガイドを参照してユーザー自身で行う必要があります。ただし、オープンウェイトであるため、独自の映像やキャラクター、特定のビジュアルスタイルでファインチューニング(微調整)することは可能です。
ライセンスに関する注意点として、年間収益が 2,000 万ドル未満の企業に限り商用利用が認められています。
同日、ByteDance は非公開モデル「Seedance 2.5」をリリースしました。こちらは内蔵オーディオ機能を備え、30 秒のクリップ生成が可能です。
AI News Without the Hype – Curated by Humans
THE DECODER に購読すると、広告なしで記事を読めるほか、週刊の AI ニュースレターや年6回の独占レポート「AI Radar」、アーカイブへのフルアクセス、コメント欄の利用が可能になります。
原文を表示
MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time. Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video. The 33-billion-parameter model processes text, images, video, and audio together, generating four- to 15-second clips with stereo sound. According to the model card, a single prompt can include up to nine reference images, three video clips, and three audio clips.
*Video by **MiniMax H3*
Two pieces remain closed, though. The 2K resolution module and H3-Context-IR, which translates prompts and reference material into a structured intermediate format, aren't included. Running H3 locally in ComfyUI tops out at 768p, and users will need to handle context prep themselves using MiniMax's published prompting guides. The open weights do allow fine-tuning on custom footage, characters, or a specific visual style. One catch on the license side: commercial use is only permitted for companies making under $20 million in revenue.
ByteDance released its closed Seedance 2.5 the same day, which generates 30-second clips with built-in audio.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
同じ出来事を3媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み