ComfyUI、MiniMax H3 のオープンウェイト動画モデルをネイティブサポート
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
MiniMax は自社初のオープンウェイトモデル H3 を公開し、翌朝には ComfyUI で Day-0 サポートが開始された。
AI深層分析を開く2026年8月3日 12:01
AI深層分析
キーポイント
オープンウェイトと ComfyUI 統合
MiniMax は自社初のオープンウェイトモデル H3 を公開し、翌朝には ComfyUI で Day-0 サポートが開始された。
マルチモーダル入力と生成能力
テキスト、画像、動画、音声を入力として受け付け、これらを統合して 2K 解像度・15 秒の動画を生成する。
ネイティブステレオ音声生成
音声を後処理で付加するのではなく、動画生成と同じパス内でネイティブなステレオ音声として出力する。
高度な編集・制御機能
最初と最後のフレーム指定や参照映像によるモーション転送など、詳細な演出制御を可能にする機能を備える。
コミックスタイルの視覚効果と音声同期
赤と青黒のパレットを用いた重厚なインク画風で、キャラクターのセリフに同期して漫画風のテキストが画面に表示される。
重要な引用
MiniMax H3 dropped today with open weights, and it's natively supported in ComfyUI as of this morning.
Audio is a property of the model, not a post-process. Every audio output is native stereo.
"GET READY TO" - "MEET" — "YOUR" — "MAKER"
"Editorial tech product film."
編集コメントを表示
編集コメント
次世代動画モデルがオープンウェイト化され、ComfyUI での利用が可能になったことは、ローカル環境での高度なコンテンツ制作を民主化する重要な一歩である。特に音声と映像の統合生成機能は、従来の後処理ワークフローに比べて実用的な価値が高いと言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
MiniMax H3 が本日、オープンウェイトモデルとしてリリースされ、本朝より ComfyUI でネイティブサポートが開始されました。まさに Day Zero です。
これは次世代のオープンウェイト動画生成モデルです。テキスト、画像、動画、音声を入力すると、リアルなステレオサウンドを備えた 2K 解像度、最大 15 秒の動画を生成します。Hailuo 01 や Hailuo 02 に続く MiniMax の第 3 世代モデルであり、同社がオープンウェイトで公開した初の動画モデルとなります。
モデルの特徴
- テキストから動画へ:プロンプトのみで生成可能。
- 画像から動画へ:静止画に命を吹き込む。
- 最初と最後のフレーム指定:開始フレーム、終了フレーム、あるいは両方を指定し、残りの部分をモデルが補完します。
- リファレンスから動画へ:参照画像、動画、音声を供給して、被写体や動き、声を一貫して維持したまま動画を生成できます。
出力は最大 2K、15 秒まで対応。音声も動画生成と同じ処理パスでステレオとして生成され、後付けされるものではありません。
マルチモーダルな文脈理解
これは MiniMax が最も力を入れている機能であり、5 つの別タスクを単一のモデルに統合する鍵となります。実際の業務では、単一のモダリティ(情報様式)だけで完結することは稀です。H3 は画像、音声、動画を同時に処理し、それらの関係性を説明するプロンプトに基づいて統合・解釈します。入力と出力のショットとの関係を記述すれば、モデルがマルチモーダル間の連携を自動的に処理してくれます。
ネイティブステレオオーディオ
音声は後処理ではなく、モデル自体の機能です。すべての音声出力はネイティブなステレオ音質で生成されます。
編集とモーション転送
グラフワークにおいて最も重要なのは、モーション転送です。参照動画からカメラの動きや演技、カットのリズムといった動きを供給し、被写体とスタイルは別のソースから取得します。これにインプレース編集を組み合わせることで、ショットごとの反復的な作業が可能になります。
出力例
コミックブックのインク画風。太い線画、赤と青黒のパレット、夜の街を舞台に。 と をリファレンスフレームとして使用し、 はそのまま利用してください。
カット 1:屋上にいる少年スーパーヒーローの真上からのアングル。赤いマフラーが風になびき、両手を腰に当ててカメラを見上げている。そばかすと自信満々な笑顔が特徴だ。カメラは彼がセリフを言う間、ゆっくりと彼の方へ降りていく。話しながら、コミック風のグラフィックオーバーレイテキストが音声に合わせて単語ごとに現れる。「GET READY TO」「MEET」「YOUR」「MAKER」。巨大な不規則なコミック書体の文字で、白地に太い黒のアウトラインと赤いドロップシャドウを施し、荒々しい角度に傾ける。3 つの言葉が彼の顔とレンズの間、空中に積み重なるように浮かび上がる。
トランジション:屋上から激しく whip pan( Whip Pan)して、浮いていた文字も一緒に塗りつぶされるように消え去る。モーションブラーがかかり、動きの跡が残る。
カット 2:スカイラインを圧倒する巨大な黒いメカカイジュウを低めのヒーローアングルで捉える。背後に身を反らし、恐ろしい大音量の ROAR(咆哮)を放つ。口は大きく開き牙がむき出し、赤い目と胸部のコアが眩しすぎるほど輝いている。頭部からは青い稲妻が走っており、その咆哮の衝撃波が塵を巻き上げ、ビル街の窓を揺らす。音のインパクトからコミック風のスピードラインやインクの飛び散りが爆発的に広がる。咆哮がピークに達する瞬間、カメラへと身を乗り出す。その咆哮をしばらく保持する。
編集者によるテック製品の映像。透明なゲーミングマウスが、元のシーンで映し出されています。背景は真っ黒なスタジオの虚無空間で、暗く微妙に反射する表面があります。ドラマチックな二色調の鮮やかな青と温かみのあるネオンオレンジのリムライティングが当たっており、深い柔らかい影が純粋な黒へと溶け込んでいます。モノクロームのダークパレットに、電気のような青とアンバーのアクセントを効かせています。素材モチーフは、発光する内部の金属マイクロコンポーネントと、光沢のあるアクリルの屈折です。この環境設定は映像全体を通じて一定です。
ショット 1:シーンは画像 1 の瞬間に開始します。マウスが自信を持って暗い表面に置かれています。青とオレンジのライトがゆっくりと明るさを増し、透明なアクリルシェルを深く透過して屈折します。カメラはゆっくりと意図的にズームインし、複雑な回路を明らかにしていきます。
ショット 2:リブ状のスクロールホイールと層状になった内部マイクロコンポーネントの極端なマクロプロファイルにカットされます。カメラが側面を滑らかに移動する間、温かみのあるオレンジ色の光の鋭いビームが金属質感を横断します。これは深い青のアンビエントグロウに対して完璧なコントラストを生み出しています。
ショット 3:ローアングルからのビューティショットにカットされます。マウスは暗く反射する表面から数センチ浮遊し、重力を感じさせずにゆっくりと精密な軌道を描いて回転します。二色調の照明がガラスのような透明な縁に沿って優しくフレアを起こした後、スリムなシルエットへとゆっくりとフェードアウトしていきます。
音声:深く脈打つサブベースのルームトーン、鋭く触覚的な機械的なクリック音、カット時のガラスのような掃引音が響き、最終的にフェードアウトする際にほぼ無音に収束する電子音の盛り上がり。
ハイファッションのエディトリアルフィルム。全体に豪華なスローモーション、柔らかいグラデーションのスタジオ空。
音楽と効果音:深い和太鼓、きらめく琴の撥弦、現代的なサブベースが融合したシネマティックなスコアが映画を駆動する。
ショット 1: その隣にはマスクが「壊れた」状態で吊るされている。それは浮遊する破片の形成へと砕け散り、各々の金継ぎ(きんつぎ)のピースは宙に浮かび、ゆっくりとその場で回転している。それらの間の金色の継ぎ目はくすんでおり、何かを待っているようだ。彼女がその方向へ目を向ける。
ショット 2: 「組み立て」の瞬間。膨大なエネルギーと共に、金色の継ぎ目が「点火」する。溶けた光のアークが破片から破片へと溶接火のように飛び移り、ピースは一つずつ snapping(結合)していく。ゆっくりから連射速度へと加速し、各々の snapping は金色に閃光を放ち、溶けた液滴が回転して飛び散る。周囲の液体リボンも衝撃波のリップルで震える。ついに最後の破片が完全に収まり、マスク全体が融合する。金継ぎの脈管は燃え上がる。
ショット 3: 「金色の龍」が巨大な蛇行飛行でフレームを横断する。赤いガラスの角が先頭となり、その体躯は彼女とマスクを取り巻く空間を巻き込む。鱗からは金色の光が放たれ、その尾跡はクリムゾンの液体を螺旋状に引きずりながら後方に残す。
ショット 4: 龍の尾跡の中で、マスクは磁気的に空中から彼女の顔へと「引き裂かれる」。速く、強く、完璧に一直線。深くflare(閃光)しながら着座し、すべての金色の亀裂が点灯する。そして輝く金継ぎの脈管がマスクの縁から首筋へ、さらにサンセットジャケットへと広がり、刺繍は糸ごとに点火されていく。
SHOT 5: 彼女が降りて暗い液体の波の上に静かに着地し、構えた戦士の姿勢へと瞬時に切り替わる。その姿はルックブックのフレームのように静止する。彼女の肩の後ろには龍が巻きつき、両側の液体が螺旋を描きながら彼女の周りを二重らせん状に上昇していく。カメラが落ち着くまで、編集されたポスターフレームとして保持される。
参照画像として、, , を使用してください。
鮮やかな魚眼レンズで捉えた製品コマーシャル。ハイパーな彩度を持つ夏の光の中で、の女性(黄色いレインコートを着てジャングルの滝のそばにしゃがみ込み、虹色のグラデーションがかかったソーダ缶をレンズに向けている)が登場し、水滴が滴り落ちている。
MUSIC: 全体を通してアップテンポなトロピカル・ハウストラックが流れる。パンチの効いたキックドラム、明るいスチールドラムの音、温かみのあるベースのグルーヴ。
カット 1:魚眼レンズのヒーローショット。彼女がカメラのレンズを見つめる瞬間、背景に巨大な太字タイポグラフィがビートに合わせて次々と現れます。「STAY」そして「HYDRATED」。画面全体を覆うような大ぶりな白抜きのブロック文字が、魚眼レンズ特有の歪みに沿って曲がりながら配置されています。彼女はその背後にありつつも滝の手前に位置し、反対の手で缶に向かって指先をフックのように引っ掛け、タブの下にかけます。
トランジション:タブの超接写。カチッと乾いた音とヒスという空気の音が同時に鳴り響き、その瞬間に魚眼レンズの絞りシャッターが黒く閉じ、まるでカメラが瞬きをするかのような演出になります。
カット 2:絞りが再び開き、新しい視点(POV)へと切り替わります。前景には極端に歪んだ缶が大きく映し出され、魚眼効果で大きく変形しています。彼女は笑顔を見せると、缶の中の液体を床へ注ぎ込みます。液滴は重力を感じないかのように飛び散り、光がその流れの中で虹色に屈折します。背景の滝は彼女の後方で柔らかく広がっています。
トランジション:缶を下ろすと、太い一滴がカメラレンズに向かって落下し、画面を埋め尽くします。
カット 3:その水滴を通して映し出される最終的なワイドショット。ターコイズブルーの滝の水たまりの中に、ラベルがカメラに向いたまま、静かに浮いている虹色の缶があります。霧の中で優しく揺れながら、背後では滝が穏やかに轟いています。水面には「STAY COMFY」という文字がきらめく反射として浮かび上がっています。製品のヒーローショットをこの状態でキープします。
シャープで陽気、プレミアムな製品広告のエネルギー。すべてのショットに魚眼レンズによる歪みを取り入れています。
ComfyUI でのローカル推論用に最適化
コンシューマー向けハードウェアで H3 を円滑に動作させるには、相当な機械学習エンジニアリングの工夫が必要でした。モデルのモジュレーション重み(全パラメータの約 40%)を剪定し、機能的に同等なルックアップテーブルに置き換えることで、出力品質を一切損なうことなくメモリ使用量を劇的に削減できました。
さらに、重みには正確で効率的な int8 畳み込み量子化(int8 convrot quantization)が適用されており、カスタムカーネルによって推論時のピーク VRAM 使用量も抑制されています。

その結果、フル精度時の 123.6 GB から最小モデル版では 42.5 GB へ、総メモリ使用量を 66% 削減することに成功しました。これに動的な VRAM オフロード機能を組み合わせることで、次世代の 2K ビデオモデルを RTX 3060 などのローカル GPU で動作させることが可能になります。
はじめに
ComfyUI を最新バージョン 0.30.0 に更新してください。
以下のワークフローをダウンロードするか、テンプレートライブラリから選択してください。
- MiniMax H3 I2V ワークフローのダウンロード
- MiniMax H3 R2V ワークフローのダウンロード
- MiniMax H3 T2V ワークフローのダウンロード
ワークフロー内の注意書きに従い、モデルをダウンロードして適切なモデルディレクトリに保存してください。
プロンプトを入力し、フレームや参照入力をつなぎ合わせて実行します。
モデル重み:珞 Comfy-Org/MiniMax-H3
いつものように、創作を楽しんでください!
原文を表示
MiniMax H3 dropped today with open weights, and it’s natively supported in ComfyUI as of this morning. Day zero.
This is a next-generation open-weights video model. Feed it text, images, video, or audio and it generates video with real stereo sound, up to 2K, up to 15 seconds a clip. It is MiniMax’s third-generation video model, following Hailuo 01 and Hailuo 02, and the first the company has released with open weights.
Model Highlights
Text-to-video — prompt only.
Image-to-video — bring an image to life.
First-and-last-frame — control the opening frame, the closing frame, or both, and let the model fill in the rest.
Reference-to-video — supply reference images, video, or audio and carry a subject, a motion, or a voice through the clip.
Output runs to 2K and up to 15 seconds. Audio is generated with the video in the same pass, in stereo, not bolted on afterward.
Multimodal context understanding
This is the capability MiniMax leads with, and it’s what collapses five separate tasks into one model. Real work rarely draws on one modality. H3 takes images, audio, and video together and resolves them against a prompt that explains how they relate. Describe the relationship between your inputs and the shot you want, and the model handles the cross-modal work itself.
Native stereo audio
Audio is a property of the model, not a post-process. Every audio output is native stereo.
Editing and motion transfer
Motion transfer is the one that matters most for graph work. A reference video can supply movement — a camera move, a performance, a cutting rhythm — while the subject and style come from elsewhere. Combined with in-place editing, that means iterating on a shot.
Example Outputs
Bold comic-book ink style, heavy linework, red and blue-black palette, night city. Use <Picture 2> and <Picture 1> as reference frames and <Audio 1> exactly as it is.
CUT 1: top-down view of the little boy superhero on the rooftop — red cape fluttering in the wind, hands planted on his hips, freckles and a cocky grin as he looks straight up into the camera. The camera slowly descends toward him as he delivers his line — as he speaks, comic-book graphic overlay text word by word in sync with his voice: "GET READY TO" - "MEET" — "YOUR" — "MAKER" — huge jagged comic lettering, white with heavy black outlines and red drop shadows, tilted at scrappy angles, until the three words hang stacked in the air above him between his face and the lens.
TRANSITION: a violent WHIP PAN off the rooftop that SMEARS the floating words away with it, motion-streaked —
CUT 2: low hero angle on the colossal black mech-kaiju towering over the skyline as it rears back and unleashes a GIANT terrifying ROAR — jaws wide with fangs, red eyes and chest-core flaring blinding bright, blue lightning arcing off its head, the roar's shockwave rippling dust and rattling windows down the buildings, comic-style speed-lines and ink splatter bursting from the impact of the sound. It leans INTO the camera as the roar peaks. Hold on the roar.
Editorial tech product film. The transparent gaming mouse from <Picture 1> in its original scene: a pitch-black studio void with a dark, subtle reflective surface, lit by dramatic duotone vibrant blue and warm neon orange rim lighting, deep soft shadow falloff into pure black. Monochromatic dark palette with electric blue and amber accents. Material motif: glowing internal metallic micro-components and glossy acrylic refractions. The environment is constant throughout.
SHOT 1: The scene opens exactly on image 1, the mouse resting confidently on the dark surface; the blue and orange lights slowly pulse brighter, refracting deeply through the transparent acrylic shell as the camera executes a slow, deliberate push-in to reveal the intricate circuitry.
SHOT 2: Cut to an extreme macro profile of the ridged scroll wheel and layered internal micro-components; the camera glides slowly along the side as a sharp beam of warm orange light sweeps across the metallic textures, contrasting perfectly against the deep blue ambient glow.
SHOT 3: Cut to a low-angle beauty shot: the mouse levitates weightlessly a few centimeters above the dark reflective surface, rotating in a slow, precise orbit; the duotone lighting flares gently along the glassy transparent edges before fading slowly into a sleek silhouette.
Audio: deep pulsing sub-bass room tone, sharp tactile mechanical clicks, a sweeping glassy whoosh on cuts, and a rising electronic swell that resolves to near-silence on the final fade.
High-fashion editorial film, luxurious slow motion throughout, soft gradient studio sky.
MUSIC & SFX: a cinematic score fusing deep taiko drums, shimmering koto plucks and modern sub-bass drives the film
SHOT 1: beside her, the mask hangs BROKEN — shattered into the floating shard formation of <Picture 2>, every kintsugi piece suspended and slowly rotating in place, the gold seams between them dim and waiting. She turns her eyes to it.
SHOT 2: THE ASSEMBLY, with enormous energy — the gold seams IGNITE, arcs of molten light leaping shard to shard like welding fire, and the pieces snap together one by one, accelerating from slow to rapid-fire, each snap flaring gold, molten droplets spinning off, the surrounding liquid ribbons shuddering with shockwave ripples — until the final shard slams home and the whole mask fuses, its kintsugi veins blazing.
SHOT 3: the golden dragon of <Picture 3> SWOOPS through the frame in one huge serpentine fly-through — red glass antlers first, its coils wrapping the space around her and the mask, scales throwing golden light, its wake dragging the crimson liquid into a spiral behind it.
SHOT 4: in the dragon's wake the mask magnetically RIPS across the air onto her face — a fast, hard, perfectly straight pull — seating with a deep flare as every gold crack lights, and glowing kintsugi veins spread from the mask's edge down her neck and across the sunset jacket, embroidery igniting thread by thread.
SHOT 5: she descends and lands softly ON the dark liquid wave, snapping into a poised warrior stance and holding it like a lookbook frame — the dragon coiled behind her shoulder, both liquids spiraling upward around her into a double helix. Held editorial poster frame as the camera settles.
Use <Picture 1>, <Picture 2>, <Picture 3> as reference images.
Vibrant fisheye product commercial, hyper-saturated summer light, the woman from <Picture 1> in a yellow raincoat crouched by a jungle waterfall holding a rainbow-gradient soda can toward the lens, condensation dripping.
MUSIC: an upbeat tropical house track drives the entire film — punchy kick drum, bright steel-drum plucks, warm bass groove.
CUT 1 : the fisheye hero frame — as she looks into the lens, GIANT BOLD TYPOGRAPHY stamps across the background behind her, one word per beat: "STAY" then "HYDRATED" — massive clean white block letters spanning the whole scene, curving with the fisheye distortion, sitting behind her but in front of the waterfall. She reaches her opposite hand towards the can and hooks a finger under the tab.
TRANSITION: extreme close-up of the tab — it OPENS with a crisp CLICK-hiss, and exactly on the click the fisheye lens iris shutters closed to black, like a camera blinking.
CUT 2: the iris reopens on a new POV — the can EXTREMELY distorted in the foreground, huge and warped by the fisheye, she smiles and dumps the liquid out of the can onto the floor, droplets scattering weightlessly, sunlight refracting rainbow through the stream, the waterfall soft behind her.
TRANSITION: she lowers the can and one fat droplet falls toward the lens, filling the frame —
CUT 3: through the droplet into the final wide: the rainbow can floating upright and serene in the turquoise waterfall pool, label facing camera, bobbing gently in the mist, the waterfall thundering softly behind — and "STAY COMFY" shimmering as a reflection on the water's surface beside it. Hold the product hero frame.
Crisp, joyful, premium product-ad energy. Fisheye distortion in every shot.
Optimized for local inference in ComfyUI
Getting H3 to run well on consumer hardware took significant machine learning engineering. We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.
On top of that, the weights ship with an accurate and efficient int8 convrot quantization, and custom kernels reduce the peak VRAM use during inference.

The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.
Getting started
Update ComfyUI to the latest version 0.30.0
Download the workflows below, or find them in the template library.
Download MiniMax H3 I2V Workflow
Download MiniMax H3 R2V Workflow
Download MiniMax H3 T2V Workflow
Follow the note in the workflow to download the models and save them in the correct model directory.
Write your prompt, connect any frame or reference inputs, and run.
Model weights: 珞 Comfy-Org/MiniMax-H3
As always, enjoy creating!
AI算出
主要ニュースainew評価標準
記事は MiniMax の第 3 世代モデル「H3」が Day-0 で ComfyUI にネイティブ対応したことを報じており、2K 解像度やネイティブステレオ音声といった具体的な新機能を含んでいるため、新規性と関連性は高い。ただし、日本企業との直接的な関わりや日本固有の情報が含まれていないため、日本の文脈での関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
同じ出来事を3媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み