MiniMax、テキスト・画像・動画・音声を統合処理するオムニモーダル動画モデル「H3」を公開
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
MiniMax は汎用マルチモーダル生成モデル「MiniMax H3」をリリースし、テキストや画像、動画、音声を統合コンテキストとして処理して 2K・ステレオ音声付きの 15 秒動画を生成する API とアプリを提供した。
AI深層分析を開く2026年8月1日 18:05
AI深層分析
キーポイント
汎用マルチモーダルモデルへの転換
MiniMax H3 は従来の個別の専門家モデル(テキストから動画、画像から動画など)を統合し、自然言語による参照や編集関係を表現する単一の事前学習パラダイムを採用した。
高度な生成仕様と制約
同モデルは 2K 解像度で 4〜15 秒の整数秒動画を生み出し、ネイティブステレオ音声を付与するが、入力ファイル数は最大 12 個に制限される。
技術的革新とアーキテクチャ
H3-VAE による高圧縮トークナイザーと H3-Omni Transformer の採用により、推論コストを削減しつつ 2K 生成を実現し、コンテキストの理解と生成タスクを分離した。
API と市場展開
同モデルは 2026 年 7 月 31 日に API および「Hailuo AI」アプリを通じて利用可能となり、広告やゲーム、映像制作などの業界向けに提供される。
H3-Omni Transformer アーキテクチャの最適化
理解と生成のワークロードを分離しハードウェア利用率を調整することで、エンドツーエンドのトレーニングスループットが約30%向上した。
重要な引用
MiniMax H3 is not a text-to-video model with add-ons.
reference the camera movement from Video 1, have the character in Image 2 sing, match the vocals to Audio 3
H3-VAE: A full tokenizer overhaul. Its high compression ratio delivers a stated 4× gain in effective sequence length
"H3 unifies text, image, video, and audio into one generation model — 2K, 4–15s, native stereo."
編集コメントを表示
編集コメント
2026 年という未来の日付でのリリース発表は、同社のロードマップや予測的な文脈を示唆している可能性がある。技術的には「自然言語による参照と編集」というアプローチが、複雑な動画制作ワークフローを簡素化する鍵となりそうだ。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
MiniMax は汎用的なマルチモーダル生成モデル「MiniMax H3」を発表しました。これは単なるテキストから動画への変換機能に付随機能を追加したモデルではありません。MiniMax によれば、このモデルはテキスト、画像、動画、音声を一つの統合された文脈として読み込み、ネイティブのステレオサウンドを備えた動画を生成する汎用マルチモーダル生成モデルです。
主な仕様は以下の通りです:出力解像度は 2K、生成時間は 4〜15 秒、かつ秒数は整数値のみとなります。
従来の動画生成パイプラインでは、テキストから動画への変換、画像から動画への変換、最初と最後のフレーム指定、被写体参照、モーション参照、動画編集など、それぞれが個別の専門モデルとして分かれていました。これに対し MiniMax H3 は、それらを一つの事前学習パラダイムに統合し、参照関係や編集指示を自然言語で記述できるようにしました。MiniMax が提示した例文は、その点をよく表しています。「Video 1 のカメラワークを参照し、Image 2 のキャラクターに歌わせ、Audio 3 の音声と歌唱を同期させる」といった指示が可能です。
実用化は可能なのか?
現状では、API を通じての利用は可能です。しかし、自前のハードウェアで動かすことは現時点ではできません。MiniMax は 2026 年 7 月 31 日に H3 の公開を開始し、モデル ID「MiniMax-H3」としてプラットフォーム API に実装されるとともに、一般消費者向けアプリ「Hailuo AI」でも利用可能になりました。
産業別用途:MiniMax は、広告、ブランディング、EC サイト、製品デザイン、UI/UX デザイン、ゲーム分野に加え、映画のプレビジュアライゼーションや小売カタログメディアなどにも MiniMax H3 を位置付けています。
具体的な応用例:広告バリエーションの自動生成、商品紹介動画やリスト作成動画、アニメーションポスター、映画タイトルシーケンス、ウェブサイトのヒーローループ、キャラクターの一貫性を保ったゲームのカットシーン、そして動画から動画へのモーション転送などです。
動画生成ガイドには、3 つの主要な入力モードが記載されています。テキストから動画を生成する「テキスト・トゥ・ビデオ」、最初のフレームまたは最後のフレーム画像から開始する「先頭/末尾フレーム・トゥ・ビデオ」、そして参照映像を利用する「リファレンス生成」です。
API のエンドポイントは 1 つに統合されており、非同期の 3 ステップフローで動作します。まずタスクを作成し、次にタスク ID をポーリングしてステータスを確認、最後に生成されたコンテンツ URL からファイルをダウンロードします。
設計時に考慮すべき入力制限は以下の通りです。
参照画像:最大 9 枚まで。
参照動画:最大 3 クリップ(各クリップ 2〜15 秒、合計 15 秒以内)。
参照音声:最大 3 クリップ。ただし、画像または動画の添付なしでは音声ファイルを送信できません。
混合入力の場合、ファイル総数は 12 個までです。プロンプト文字数は 7,000 字以内、リクエストボディは 64 MB 以内とされています。大規模なアセットを扱う場合は URL 入力が推奨されます。
各ファイルのサイズ制限は以下の通りです。
動画:50 MB 以内
画像:30 MB 以内
音声:15 MB 以内
対応フォーマットは、動画が H.264/H.265、画像が JPG/PNG/WEBP/HEIC/HEIF、音声が WAV/MP3 です。
このモデルを支える 4 つの技術的要素について解説します。
コンテキスト・オムニ・レプレゼンテーション(文脈統合表現)
MiniMax はキャプション生成機能を再構築しました。従来の「対象となる動画そのもの」を記述するだけでなく、「文脈と対象動画の関係性」までを含めて記述できるようにしています。これにより、ソース素材の多くは推論に約 10 万トークン必要でしたが、平均して約 4,000 トークンに圧縮(ディストillation)されました。言語が橋渡しとなり、固定されたタスクセットから、より柔軟で記述的なオープンな形式へと変換されています。
H3-VAE(高効率トークナイザー)
これはトークナイザーの完全な刷新です。高い圧縮率を実現し、有効なシーケンス長を 4 倍に引き上げる成果を出しています。これによりトレーニングと推論のコストが削減され、ネイティブ 2K 解像度に対応できる基盤技術となっています。
H3-Omni Transformer:MiniMax はここでは Hailuo-02 のアーキテクチャを意図的に外しました。マルチモーダルなコンテキストのシーケンス長変動が 3 倍に増えたため、トレーニング用アーキテクチャは理解と生成のワークロードを分離し、それぞれに対してハードウェア利用率を最適化しています。その結果、エンドツーエンドのトレーニングスループットは約 30% 向上しました。
インコンテキスト再生成:ボルトオン型の超解像モジュールを追加するのではなく、ベースモデルが低解像度の出力を自身でインコンテキスト内で再生成し、元のマルチモーダルなコンテキストを読み直します。これにより、従来のアップスケーラーが推測に頼っていた小さな文字や細部の描写が復元され、ブランドや製品のレンダリングにおいて直接的な価値を持ちます。
価格と立ち位置について
MiniMax 自身の主張によると、H3 は 2K 解像度での 1 秒あたりの料金が主流モデルの 3 分の 1 未満、768p では主流の 720p モデルの半額以下です。同社は X(旧 Twitter)で発表と価格設定の両方を強調しました。第三者の追跡サイトや報道によると、2K の従量課金料金は 1 秒あたり 0.13 ドル(15 秒のクリップで約 1.95 ドル)ですが、執筆時点では MiniMax の従量課金ページにはまだ Hailuo 2.3 のティアしか表示されていなかったため、この数値は報告されたものとして扱い、公式情報とは区別して捉える必要があります。
SCMP は Artificial Analysis のレポートを引用し、MiniMax H3 が動画編集においては首位に立つ一方、テキストから動画への生成では Google の Gemini Omni Flash に、画像から動画への生成では Seedance 2.0 および Gemini Omni Flash にそれぞれ劣っていると報じています。
主なポイント
H3 はテキスト、画像、動画、音声を統合した単一の生成モデルです。解像度は 2K で、4〜15 秒のクリップをネイティブステレオ音声付きで生成できます。
オープンウェイト版は「今後数日中に公開される予定」ですが、現時点ではまだリリースされておらず、利用可能な唯一の方法は API です。
H3-VAE が持つ 4 倍の有効シーケンス長拡張が、ネイティブ 2K 生成を経済的に実現可能にしています。
スーパーリゾリューションの代わりにコンテキスト内再生成を採用することで、小さな文字やブランドロゴの品質も維持されています。
Artificial Analysis のランキングでは、H3 は動画編集で 1 位ですが、テキストから動画、画像から動画への生成では競合他社に後れを取っています。
感情分析
MiniMax が「MiniMax H3」を発表。15 秒間の 2K クリップを生成し、ネイティブのステレオ音声も対応するオムニモーダルな動画モデルです。
この発表は、MarkTechPost で最初に紹介されました。
原文を表示
MiniMax releases MiniMax H3, a general-purpose multimodal generation model. MiniMax H3 is not a text-to-video model with add-ons. MiniMax describes it as a general-purpose multimodal generation model that reads text, images, video, and audio as one unified context and returns video with native stereo sound. The mains specs include: 2K output, 4–15 seconds, integer durations only.
Previous video stacks split into text-to-video, image-to-video, first-and-last-frame, subject reference, motion reference, and video editing, each often a separate expert model. MiniMax H3 folds those into one pretraining paradigm where reference and editing relationships are expressed in natural language. MiniMax’s example prompt makes the point: reference the camera movement from Video 1, have the character in Image 2 sing, match the vocals to Audio 3.
Is it deployable?
Today: yes, through the API and no, on your own hardware. MiniMax launched H3 on July 31, 2026 with the model live in the platform API under the model ID MiniMax-H3 and in the consumer Hailuo AI app.
Industries: MiniMax positions MiniMax H3 for advertising, branding, e-commerce, product design, UI/UX, and gaming along with film pre-visualization and retail catalog media.
Applications: Ad variant generation, product and listing videos, animated posters, film title sequences, website hero loops, character-consistent game cinematics, and video-to-video motion transfer.
The API surface
The video generation guide documents three entry modes: text-to-video, first/last-frame image-to-video, and reference generation. Behind one endpoint and an asynchronous three-step flow: create a task, poll task_id, download content.url.
Input limits worth designing around:
Reference images: up to 9. Reference videos: up to 3 clips, 2–15 s each, ≤15 s total. Reference audio: up to 3 clips, and audio cannot be sent without an accompanying image or video.
Mixed input caps at 12 files total. Prompt length ≤7,000 characters; request body ≤64 MB, with URL input recommended for large assets.
File sizes: video ≤50 MB, image ≤30 MB, audio ≤15 MB, per asset.
Formats: H.264/H.265 video, JPG/PNG/WEBP/HEIC/HEIF images, WAV/MP3 audio.
Four technical pieces doing the work
Contextual Omni Representation: MiniMax rebuilt captioning so it describes the relationship between context and target video, not just the target. Most source material requires roughly 100K tokens of inference, distilled to about 4K tokens on average. Language is the bridge that turns a fixed task set into an open, descriptive one.
H3-VAE: A full tokenizer overhaul. Its high compression ratio delivers a stated 4× gain in effective sequence length, cutting training and inference cost and it is the enabling technology for native 2K.
H3-Omni Transformer: MiniMax explicitly set aside the Hailuo-02 architecture here. Multimodal context tripled sequence-length variance, so the training architecture separates understanding and generation workloads and tunes hardware utilization for each. Reported result: end-to-end training throughput up nearly 30%.
In-Context Regeneration: Instead of a bolt-on super-resolution module, the base model regenerates its own low-resolution output in-context, re-reading the original multimodal context. That is what recovers small text and fine detail that conventional upscalers guess at — directly relevant to brand and product rendering.
Price and standing
MiniMax’s own claim: at 2K, H3’s per-second price is less than a third of mainstream models; at 768p, less than half the price of mainstream 720p. The company amplified both the launch and the pricing framing on X (1, 2). Third-party trackers and launch coverage put the 2K pay-as-you-go rate at $0.13 per second, about $1.95 for a 15-second clip, but MiniMax’s pay-as-you-go page still listed only Hailuo 2.3 tiers at the time of writing, so treat that figure as reported, not primary.
On placement: SCMP reports, citing Artificial Analysis, that H3 leads in video editing while trailing Google’s Gemini Omni Flash in text-to-video and sitting behind both Seedance 2.0 and Gemini Omni Flash in image-to-video.
Key Takeaways
H3 unifies text, image, video, and audio into one generation model — 2K, 4–15s, native stereo.
Open weights are promised “in the coming days,” not shipped; the API is the only path today.
H3-VAE’s 4× effective sequence-length gain is what makes native 2K economically viable.
In-context regeneration replaces super-resolution, preserving small text and brand marks.
Artificial Analysis ranks H3 first in video editing, behind rivals in text-to-video and image-to-video.
Sentimental Analysis
Check out the Technical details. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio appeared first on MarkTechPost.
AI算出
主要ニュースainew評価標準
AI モデルの核心機能(オムニモーダル処理、2K/15 秒生成、ネイティブ音声)と技術的革新(コンテキスト統合表現、高効率トークナイザー)が詳細に報じられており、新規性の高い主要ニュースである。ただし、日本企業との直接的な関わりや日本固有の情報が含まれていないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み