ComfyUI に動画生成モデル「HappyHorse 1.1」追加
ComfyUI のパートナーノードとして登場した動画生成モデル「HappyHorse 1.1」は、ネイティブ音声同期機能と多様なプロダクション要件への対応により、実用レベルの動画制作ワークフローを大幅に強化しました。
キーポイント
ネイティブ音声同期の実現
ダイアログ、効果音、BGM を単一のレンダリングパスで生成する機能を搭載し、追加工程なしで完全同期された動画を作成可能にしました。
プロダクション品質の向上
動的な動きの一貫性、キャラクターの固定、プロンプトへの忠実度、テキストレンダリング、そして映画のような構図を強化し、商業利用に適した品質を実現しました。
高度な参照機能と柔軟性
最大 9 枚の画像参照をサポートし、キャラクターとシーンを分離して組み合わせることで、背景が変わってもキャラクターの一貫性を保ちつつ複雑なストーリーを生成できます。
3 つの専用ノードによる多様な用途
テキストから動画を作成する T2V、静止画像からアニメーション化する I2V、そして複数キャラクターとシーンを参照して演出する R2V の 3 つモードを提供し、720p/1080p に対応しています。
3 つの動作モードと出力仕様
Text-to-Video、Image-to-Video、Reference-to-Video の 3 モードから選択可能で、720p または 1080p で音声付きの動画が生成されます。
ワークフローのダウンロード提供
各モード(I2V, T2V, R2V)に対応した ComfyUI ワークフローファイルのダウンロードリンクが用意されています。
重要な引用
It produces dialogue, sound effects, and background music in one single render pass without extra steps.
Version 1.1 targets five core production-critical capabilities: dynamic, expressive motion; consistent character rendering; reliable prompt adherence; stable text rendering; and authentic cinematic framing.
Better long-context retention handles prompts beyond 2,500 characters, and a single prompt can describe 6–8 consecutive scenes with the model autonomously allocating time and switching camera angles.
Output arrives with audio baked in at 720p or 1080p.
影響分析・編集コメントを表示
影響分析
このリリースは、動画生成 AI が単なる実験的なツールから、商用・実用レベルのプロダクションワークフローに直接組み込める本格的なソリューションへと進化したことを示しています。特にネイティブ音声同期機能と高度なキャラクター一貫性の確保により、制作コストの削減とクオリティの向上が同時に達成され、クリエイターや企業の動画制作プロセスに大きな変革をもたらすでしょう。
編集コメント
HappyHorse 1.1 の登場は、動画生成 AI が「映像を作る」だけでなく「音声を伴う物語を構築する」段階へと進化したことを象徴しています。特に ComfyUI 上でノードとして提供されることで、既存のワークフローへの導入ハードルが下がり、実務での活用加速が期待されます。
ComfyUI にて、パートナーノードとして「HappyHorse 1.1」の利用が可能になりました。この動画生成モデルは、短編シリーズ、EC サイト向けのコマーシャル、ブランドマーケティングコンテンツ、ゲームのカットシーンなど、実務での本番運用を想定して設計されています。
同モデルの目玉機能は、ネイティブな音声同期生成です。セリフ、効果音、BGM を追加の手順なしで、1 回のレンダリングパスで同時に生成できます。
バージョン 1.1 では、以下の 5 つの本番運用に不可欠な機能を強化しました。それは、ダイナミックで表現豊かな動き、一貫性のあるキャラクター描画、プロンプトへの忠実な追従、安定した文字のレンダリング、そして映画のような本格的な構図です。
ComfyUI ニュースレターをお読みいただきありがとうございます。新着投稿の受け取りや活動支援のため、無料で購読できます。
ワークフローをダウンロードする
1.1 版の新機能
ダイナミックな表現力:滑らかな動きとフレームの一貫性により、v1.0 で見られた硬直した動きや遅鈍さが解消されました。
多画像参照から動画へ(R2V)の強化:入力された詳細を忠実に保持し、生成時に最大 9 枚までの参照画像をサポートします。
複数キャラクターの一貫性:複数のキャラクター参照を使用しても、各キャラクターの外見が明確に保たれ、視覚的な混同は発生しません。
キャラクターとシーンの柔軟な組み合わせ:キャラクターと背景を別々の参照として読み込めます。背景環境が変わっても、キャラクターの描写は一貫して維持されます。
指示の追従能力を強化:2,500 文字を超える長いプロンプトも正確に理解し、1 つのプロンプトで 6〜8 シーン連続の描写が可能になりました。モデルが自動的に時間を配分し、カメラアングルを切り替えるため、複雑な構成もスムーズに実現できます。
自然な肌質感とクローズアップ対応:テカりや過度なシャープネスの問題を解消し、シリーズ物や CM にも使えるリアルなテクスチャを実現しました。
映画言語の完全サポート:"ショット・リバース・ショット" や "トラッキング・ショット" といった専門用語も完全に理解。カット間のトランジションとペース配分が格段にスムーズになりました。
音声機能の向上:セリフや効果音の再現精度が高まり、緊密な映像との同期を保ちながら、感情表現豊かなパフォーマンスを付加しました。
3 つのノード、1 つのモデル
HappyHorse 1.1 は、それぞれ異なる用途に最適化された 3 つのノードとして提供されます:
Text-to-Video (T2V):ゼロから完全なシーンを構築。スタイル、ショットサイズ、照明、アクション、音声まで、すべてプロンプトで制御できます。
Image-to-Video (I2V):静止画の最初のフレームをアニメーション化。画像自体にデザインや雰囲気が含まれているため、動きとカメラワークだけを記述すれば OK です。
Reference-to-Video (R2V):複数キャラクターによる舞台劇のような演出が可能。キャラクターやシーンを参照画像にマッピングし、タイムスタンプ付きのストーリーボードで各キャラクターのセリフを指示できます。
すべてのモデルは 720p と 1080p の出力に対応。動画長さは 3〜15 秒まで柔軟に設定でき、アスペクト比も 16:9、9:16、1:1、4:3、3:4、21:9 など多岐にわたります。エクスポートされるすべての動画には、完璧に同期された音声が付属します。
使い始め
ComfyUI を最新バージョンに更新してください。
HappyHorse のノードは、Node Library から「HappyHorse」と検索するか、Templates Library に用意されたテンプレートを読み込むことで利用できます。
モードを選択します。Text-to-Video(テキストから動画へ)、Image-to-Video(画像から動画へ)、Reference-to-Video(参照画像から動画へ)のいずれかを選び、プロンプトと必要に応じて参照画像を接続して実行すれば完了です。生成される出力には音声も含まれており、720p または 1080p の解像度で提供されます。
Workflow をダウンロードする(I2V)
Workflow をダウンロードする(T2V)
Workflow をダウンロードする(R2V)
ComfyUI ニュースレターをお読みいただきありがとうございます。新しい投稿を受け取り、私の活動を支援するために、無料で購読してください。
原文を表示
HappyHorse 1.1 is now available in ComfyUI as a Partner Node. This video model is engineered for real-world production use cases, including short episodic series, e-commerce commercials, brand marketing content, and game cutscenes.
A standout feature of the model is native synchronized audio generation. It produces dialogue, sound effects, and background music in one single render pass without extra steps.
Version 1.1 targets five core production-critical capabilities: dynamic, expressive motion; consistent character rendering; reliable prompt adherence; stable text rendering; and authentic cinematic framing.
Thanks for reading ComfyUI Newsletter! Subscribe for free to receive new posts and support my work.
Download Workflow
What’s new in 1.1
Dynamic expressiveness: Smoother motion and frame consistency eliminate the stiff, sluggish movement from v1.0.
Enhanced multi-image reference-to-video (R2V): Faithfully preserves input details, supporting up to 9 reference images per generation.
Multi-character consistency: Multiple character references keep a distinct look with no visual cross-contamination.
Flexible character × scene combinations: Feed characters and scenes as separate references. Characters stay fully consistent even as the background environment changes.
Upgraded instruction following: Better long-context retention handles prompts beyond 2,500 characters, and a single prompt can describe 6–8 consecutive scenes with the model autonomously allocating time and switching camera angles.
Natural skin and close-up viability: Fixes shiny skin and over-sharpening issues, with lifelike texture for series and commercials.
Cinematic language: Full support for terms like shot-reverse-shot and tracking shot, with far more cohesive transitions and pacing between shots.
Upgraded audio: More accurate dialogue and sound-effect rendering, with emotional performance layered on top of tight audio-video synchronization.
Three nodes, one model
HappyHorse 1.1 ships as three nodes, each tuned to a different job:
Text-to-Video (T2V): Build a complete scene from scratch. You control style, shot size, lighting, action, and audio entirely through the prompt.
Image-to-Video (I2V): Animate a static first frame. The image already carries the look, so you just describe the motion and the camera move.
Reference-to-Video (R2V): Orchestrate a multi-character stage play. Map characters and scenes to reference images, then direct them through a timestamped storyboard with per-character dialogue.
All three models support 720p and 1080p output, video lengths ranging from 3 to 15 seconds, plus flexible aspect ratios including 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and more. Every exported video comes with perfectly synced audio.
Getting Started
Update ComfyUI to the latest version
Find the HappyHorse nodes via the Node Library (search “HappyHorse”) or load a ready-made template from the Templates Library.
Pick your mode: Text-to-Video, Image-to-Video, or Reference-to-Video, wire in your prompt and any reference images, then run. Output arrives with audio baked in at 720p or 1080p.
Download Workflow (I2V)
Download Workflow (T2V)
Download Workflow (R2V)
Thanks for reading ComfyUI Newsletter! Subscribe for free to receive new posts and support my work.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み