ByteDance Seed、動画生成モデル「Seedance 2.5」を発表
本文の状態
日本語全文を表示中
詳細モードで約19分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ByteDance Seed Blog
ByteDance は Seedance 2.5 を正式に発表し、30 秒の単発生成や複数回拡張機能、多様な参照素材の活用、タイムスタンプレベルでの編集制御を強化した次世代動画作成モデルとしてリリースした。
AI深層分析を開く2026年7月31日 23:40
AI深層分析
キーポイント
長時間ストーリーテリングの実現
1 回の生成で最大 30 秒の高品質なオーディオビジュアルクリップを作成可能であり、複数回の拡張機能により一貫性のある多分間のコンテンツを一つのテイクで制作できる。
高度なマルチモーダル参照機能
1 回のパスで最大 30 枚の画像、10 本の動画クリップ、10 個の音声クリップを参照素材として入力でき、クレイレンダリングやモーションなど多様な参照能力を強化している。
精密かつ安定した編集機能
オーディオとビデオコンテンツのターゲット編集にタイムスタンプレベルでの制御を提供し、グリーンスクリーンやカメラ視点変更などの高度な編集機能をプロフェッショナルな用途に対応させる。
30秒間の単一パスによる長編ストーリーテリング
Seedance 2.5は生成可能時間を15秒から30秒に延長し、準備から展開、転換点、結末に至るまで論理的につながった複数のショットを構成して物語を展開する。
多段階拡張による一貫性の維持
既存の動画出力に対して後続のショットを滑らかに追加できる多ラウンド拡張機能を備え、主要キャラクターや環境、物語のリズムの一貫性を保つ。
重要な引用
Today, we are officially launching Seedance 2.5, the new-generation video creation model.
Seedance 2.5 can generate high-quality, 30-second audio-video clips in a single pass and supports multiple rounds of extension.
Users can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass.
Within 30 seconds, the model can organize multiple logically connected shots so that a story unfolds through setup, development, turning points, and resolution, rather than simply extending a single moment.
編集コメントを表示
編集コメント
Seedance 2.5 の発表は、生成 AI が単発の動画クリップ作成から、一貫性のある長編ストーリーやプロフェッショナルな編集ワークフローへと進化する重要な転換点を示している。特に参照機能とタイムスタンプ制御の強化は、実務レベルでの活用可能性を大きく高める要素である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本日、次世代の動画生成モデル「Seedance 2.5」を正式にリリースします。
Seedance 2.0 の公開以降、ユーザーが求める動画生成モデルの役割は変化しました。単なるクリップの生成から、一つの作品としての完成へと期待が高まっているのです。Seedance 2.0 が持つ統合型マルチモーダル音声・映像共同生成アーキテクチャを継承しつつ、Seedance 2.5 は「基礎的な生成」と「参照に基づく生成」に焦点を当て、長編ストーリーテリング、マルチモーダル参照、そして編集機能において大きな飛躍を遂げました。実世界のユースケースに基づいて設計された Seedance 2.5 は、クリエイティブな想像力と制御性をさらに広げ、生産性の向上も実現します。
主な特徴は以下の通りです。
1 回の生成で最大 30 秒、複数回の延長が可能
Seedance 2.5 は、一度の処理で高品質な 30 秒間の音声付き動画クリップを生成でき、さらに複数回にわたる延長にも対応しています。より長い動画における連続性を高めるため、カットのつなぎ目やシーン変更も改善され、画像・音声・動きの質が大幅に向上しました。その結果、AI 動画でよく見られる不自然さを排し、より滑らかで洗練された視覚品質を実現します。これにより、ユーザーは統一された映像言語で高品質な数分間のコンテンツを、ワンカットで生み出すことが可能になります。
多モーダル参照機能の完全強化: 1 回の入力で最大 30 枚の画像、10 クリップの動画、そして 10 クリップの音声を参照素材として利用できるようになりました。粘土レンダリングやモーション、クリエイティブなスタイルなど、幅広い参照対応も強化され、複数の被写体やシーン、カット割りをまたぐ複雑なアイデアを、制作者の意図をより深く理解した上で実現可能です。
精度と安定性の向上した編集機能: タイムスタンプレベルでの制御が可能になり、音声や動画コンテンツの特定部分へのターゲット編集が効率化されました。さらに、グリーンバック処理、カメラアングルの変更、参照に基づく編集といった高度な機能を強化し、映画や広告制作などプロフェッショナルで複雑な現場の厳しい要件にも対応します。
長編ストーリーテリング、多モーダル参照、そして編集機能の進化により、Seedance 2.5 は単に「1 回で生成できる動画の長さ」を超えた存在となりました。制作者のクリエイティブな意図をより深く理解し、アイデアから完成した映像までのプロセスを、高い制御性のもとで実現します。そこで今回は、Seedance 2.5 でゼロからエンドツーエンド制作された短編映画をご覧いただくことにしましょう。
あなたのブラウザでは動画再生がサポートされていません。今日より Seedance 2.5 は「Jimeng AI」や「Doubao Pro」などのプラットフォームで公開を開始し、今後は BytePlus ModelArk を通じて API アクセスも提供予定です。ぜひご体験いただき、フィードバックをお寄せください。
プロジェクトホームページ:
https://seed.bytedance.com/seedance2_5
アクセス方法:
Jimeng Web → 動画生成 → Seedance 2.5 を選択
Doubao Pro → 動画生成 → Seedance 2.5 を選択
1 回で完結する 30 秒の長尺ストーリーテリング:複数回の拡張機能で完全な物語を一度に描く
Seedance 2.5 は、単発での動画生成時間を 15 秒から 30 秒に延長し、より長い動画におけるストーリーテリング能力を強化しました。この 30 秒の枠内で、モデルは論理的につながった複数のショットを構成し、単なる瞬間の延長ではなく、「設定」「展開」「転換点」「解決」という物語の流れを自然に描き出します。
例えば、歌手のステージパフォーマンスを捉えたワンカット映像では、 dressing room(楽屋)でスタッフと交流する様子から始まり、バックヤードの廊下を歩き、ダンサーたちと合流し、最後に彼らと共にステージへと立ち上がるまでの一連の流れを、単にステージに立つ瞬間だけを切り取るのではなく、物語全体として表現します。
ブラウザで動画が再生できません。
T2V プロンプト: 1 ショットで手持ちジンバルによるトラッキングショット。カメラは重厚な赤いカーテンの隙間からゆっくりと押し込み、温かみのあるバックステージの着替え部屋へと入っていく。カメラに背を向けた若い女性歌手がイヤーピースを調整しているところへ、スタッフから出演時間だと告げられる。彼女はカメラの方を向き、シティポップを歌い始める。カメラは後退しながら彼女を追跡し、カーテンをくぐって薄暗いバックステージの通路へと進む。その間、踊り子たちと自然に交流し、あるスタッフがマイクを手渡す。やがて彼女と踊り子たちはステージへ上がり、カメラは弧を描いて背後へと回り込み、赤と黒のステージデザイン、LED スクリーン、スポットライト、ハザード(霧)、反射する床を徐々に映し出す。最後にカメラは引き、アリーナ全体を捉えたワイドショットとなり、満員の観客、ライティングボード、光るスティック、歓声上げる群衆が映し出される。若々しく自由な精神に溢れるコンサートのクライマックスを捉えている。
モデルの複数回拡張機能により、ユーザーは既存の動画出力に対してその後のショットをスムーズに追加できます。拡張プロセスを通じて、主要キャラクター、環境、物語のリズムの一貫性が維持されます。これにより、数分間の動画を一度に生成できるようになり、クリップの分割や映像の繰り返し結合、トランジションの修正にかかる手間を大幅に削減できます。
ブラウザで動画が再生できません。
*R2V プロンプト:動画を拡張してください。@Video 1 の映像と登場人物を引き継ぎ、30 秒間の新しいクリップを生成します。キャラクター、シーン、ビジュアルスタイル、効果音はすべて一貫性を持たせてください。少年がサッカーボールを抱えて列車の車内を走ります。地下鉄が停車すると側扉が開き、彼はすぐに外へ飛び出します。男性主人公が彼を追いかけます。二人はホームを横断して路上へと駆け出し、通行人や車両に驚きを与えます。最後に男性主人公が少年を捕まえて止めます。少年は悔しそうな表情で上を見上げます。すると、怒っていた男性主人公の表情が次第に和らぎ、少年の頭を撫でて、あきれ顔で微笑みます。
視覚的な表現においては、カメラワーク間の遷移がより滑らかになりました。主要な被写体は複数のカットを通じて安定して描かれ、音声と映像も同期しているため、非常に一貫性のある長尺動画の生成が可能となっています。例えば、京劇のシーンでは、主人公の流れるような袖に合わせて優雅に円を描くようにカメラがパンします。この際、被写体と背景は完全に統一され、袖の揺れは空中で自然な弧を描きながら、現実世界の物理法則にも忠実に表現されています。
ブラウザで動画が再生できません。
*R2V プロンプト: 16:9 ワイドスクリーン、映画のような質感、単一の連続撮影、滑らかなカメラワーク、カットなし。シーン参照:@画像 4*
0–5秒:@画像 2 の「オーバーロード」のアップショットから始まります。カメラは彼の上半身をゆっくりと円を描くように回り込み、ミディアムショットへと移行します。オーバーロードが回転し、体や背中の旗がレンズを素早く通り過ぎることで自然なオクルージョン(遮蔽)を作り出し、カメラは@画像 1 の「コンソート・ユー」の側へ追従します。
6–10秒:@画像 1 のミディアムショットから、コンソート・ユーを安定して円を描くように撮影。彼女の水流袖(ウォータースリーブ)が弧を描く様子を捉えます。彼女は腕を上げ、手首を振って袖を広げ、半回転します。その後、袖を引き寄せポーズを決め、横目でオーバーロードの方を見ます。
11–20秒:@画像 3 の男性戦士が空中宙返りから登場。オーバーロードが中央に立ち、戦士は対角線上で攻防を繰り広げながら前後します。コンソート・ユーはオーバーロードの背後やや横に位置し、水流袖の動きで力強さの中に柔らかさを加えます。カメラは戦士のミディアムクローズアップからゆっくりと引き、舞台全体を捉えるフルショットへと移行します。最後は3人が一斉に観客の方を向き、京剧(ペキンオペラ)特有の決定的なポーズで幕を閉じます。
また、AI 生成動画特有の不自然さを解消するため、Seedance 2.5 はオブジェクトの質感や肌・目のディテール、照明、色彩の彩度といった細部を体系的に最適化します。さらに、字幕や背景音楽における制御不能な発生も最小限に抑え、実写映画のようなクオリティに近い完成品を提供します。
多モーダル参照機能の包括的アップグレードで、複雑な創作タスクへの制御力を強化
Seedance 2.5 は、多モーダル参照生成能力をさらに高めています。一度の入力で最大 30 枚の画像、10 クリップの動画、そして 10 クリップの音声ファイルを参照素材として読み込めるようになりました。より多くの量と幅広い種類の参照情報を活用することで、ユーザーの意図をより正確に捉え、複数の被写体や豊かな背景、柔軟なカメラワークを持つ複雑な動画を生成できます。
このモデルは、視覚的な構成、シーン、スタイル、キャラクター、小道具といった要素をすべての素材から包括的に理解し、指示に従って動画生成プロセスに適用します。複数人物のショットやグループでの物語展開といった複雑なシナリオにおいても、各キャラクターの外見と声を維持しつつ、それぞれの特性を安定して保つことが可能です。
ブラウザで動画が再生できません。
R2V プロンプト: 16:9 の横長フォーマットで、30 秒間のコンサートシーンを生成します。映像は映画のようなリアリティを持ち、本物のコンサートホールの照明と影、温かみのある黄金色のステージライト、厳かなクラシックコンサートの雰囲気を表現してください。
会場は @Image 1 を参照し、ピアニストは @Image 2、チェロ奏者は @Image 3、バイオリン奏者は @Image 4 をそれぞれ基準にしてください。リードボーカリストは @Image 5 の姿を厳密に再現します。オーケストラの残りのメンバーには @Images 6 から 10 を、合唱団には @Images 11 から 14 を、観客席には @Images 15 から 18 をそれぞれ参照させてください。
演出としては、リードボーカリストがステージ中央から手前へ歩み出る様子を描きます。ピアニストはピアノのそばに配置し、オーケストラは左右と後方に配列します。合唱団はステージ奥に立ちます。
映像はまず、コンサートホール全体を捉えた高角度のワイドショットで始まります。ピアニストが鍵盤を弾くと同時に、リードボーカリストがスポットライトの中へ踏み込み、歌唱を開始します。カメラは自然な動きでバイオリン、チェロ、そしてオーケストラへと移り、彼らが一体となって演奏する様子を捉えます。バイオリンの音色は明るく、チェロの音色は温かみのあるものとして表現してください。
後半では合唱団も加わります。リードボーカリストは手前の観客と一瞬視線を交わし、観客からは笑顔と軽い頷きで応答が返ってきます。最後のショットではカメラが引いていきます。歌唱が終わると、観客から拍手が起こります。
Seedance 2.5 は、クレイレンダリングやモーション、クリエイティブな参照といった特定の参照機能を強化し、フレーム内の被写体、動作、カメラワークに対するより細やかな制御を可能にします。例えば、クレイレンダリング参照機能を使えば、テクスチャのない 3D モデルでシーンの空間構造、キャラクターのポーズ、動きの軌道、カメラアングルなどを構築できます。モデルはこの構造情報に基づいて動画を生成するため、複雑なショットの構図や配置がクリエイターの意図に忠実に反映されます。
さらに Seedance 2.5 はライティング制御も改善しました。クレイレンダリングから得られる空間情報を活用することで、光源の方向、色温度、強度、影の投影など物理法則に従ったリアルな照明効果を生成します。その結果、最終的な映像で光と影がより自然に表現されるようになります。
ブラウザで動画が再生できません。
*R2V プロンプト:カメラの動き、ペース配分、ショットサイズの変化、被写体の軌跡、およびブロック構成については「@Clay Render 1」を参照してください。キャラクターデザイン、シーン設定、素材、照明、色彩、そして童話のような雰囲気については「@Image 2」を参照し、白いモデルを夢見がちな温かみのある 3D アニメーションショートとしてレンダリングしてください。物語は以下のように展開します:幻想の空を飛行する → 雲海を飛ぶ伝説の獣たちと並走する → 海へとダイブする → マンタレイと共に深海を泳ぎ抜く → 時空の鏡のような裂け目を通過する → 宇宙から星をつかみ取る → ベッドルームへ戻る → お父さんが毛布をかけてくれる → 絵本が閉じられ、最後のフレームで静止します。
より精密かつ信頼性の高い編集機能によるクリエイティブ効率の向上
動画制作において、ユーザーは生成プロセス中にペースを制御するとともに、後工程で詳細を調整する必要があります。アクションが発生する正確な秒数、カメラカットのタイミング、特定のクリップ内でのキャラクター動作に修正が必要かどうかといった要素は、最終的な作品に大きな影響を与えます。
より精密で信頼性の高い編集機能を備えることで、クリエイターは自らのアイデアを的確に具現化でき、作業効率の向上と反復生成にかかるコストの削減が可能になります。
Seedance 2.5 はタイムスタンプを活用した精密なコンテンツ編集をサポートします。生成プロセス中、ユーザーは特定の時間枠に対してプロンプトを用いて物語の展開、カメラアングル、動き、全体のテンポを細かく制御できます。これにより、出力結果がクリエイターの意図にさらに近づきます。生成後でも、キャラクターやアクション、プロットの要素など、特定のクリップ内でのターゲットを絞った修正が可能で、編集前後の連続性とリアリティを保ちながら作業を進められます。
また Seedance 2.5 は、グリーン画面編集、カメラ視点の変更、参照画像に基づく編集といった複数の機能を強化し、映画や広告といったプロフェッショナルな現場の厳しい要件にも対応します。例えばグリーン画面編集では、主役をそのままに保ちつつ背景を差し替え、全く異なる物語を紡ぎ出すことが可能です。さらに、新しい環境における物理法則への反応も見事に描き出せます。衣服が揺れる方向や髪の状態、歩行のリズム、照明との相互作用など、被写体がシーンと調和して溶け込むよう細部まで忠実に再現します。
ブラウザで動画が再生できません。*R2V プロンプト:@Video 1 を使用し、グリーンスクリーンの背景、障害物、衣装、サポートキャラクターをレンダリングしてください。0〜4 秒:屋外でのトレーニング、障害物を岩やレンガ、タイヤ、木製の箱に置き換える。4〜10 秒:ロッカールームで友人が励ます様子。10〜15 秒:国際試合、トレーニング用のポールを元のディフェンダーとゴールキーパーに置き換え、主人公が得点するシーン。全体的にフォトリアルで映画のようなクオリティにする。
ブラウザで動画が再生できません。*R2V プロンプト:@Video 1 を編集します。キャラクター、アクション、ビジュアルスタイルはそのままに、カメラの動きのみを調整してください。15 秒間のセグメント化されたカメラプラン:0〜4 秒、マイクロ FPV モーションでパンの横をすれ違い、ポップするトーストとホイップパンでコーヒーへと追跡。4〜7 秒、パンの縁に沿って押し込みながら横方向にトラッキングし、ひっくり返って元に戻るフライドエッグを追う。7〜11 秒、素早くトップダウンビューへ上昇し、一定の速度で下降して皿とキーを sweeping する。11〜15 秒、ハンドヘルドのクローズアップで手を追跡し、高速な横方向のホイップをかけ、朝食に押し込み、ミディアムツーショットへと引き抜く。シーケンス全体は滑らかで連続的かつ安定したものに保つ。
より広範な産業シナリオへ、現実世界での価値を追求し続ける
モデルの能力が進化するにつれ、Seedance 2.5 は教育や製造業など、より幅広い産業分野へとその適用範囲を広げています。特に教育現場では、すでに実際の学習環境への導入が始まっています。
例えば、Seedance 2.5 は授業の背景にある歴史的文脈や登場人物、ストーリーを、より生き生きとした没入感のある映像に変換します。また、教師が指導用ビデオを作成する際の効率も向上させます。科学原理や歴史的出来事、実験手順といった抽象的な内容を、ダイナミックなデモンストレーションへと変えるのです。これにより教育コンテンツ制作のハードルが下がるだけでなく、きめ細やかなカスタマイズも可能になります。
*An example of Doubao Learning app's "Doubao Classroom" scenario*
*R2V プロンプト:表現豊かな東洋画風のスタイル。南宋時代の臨安の街並み。賑やかな通りを子供たちが走り回り、「振り返ればそこにいる、明かりが薄れていく」と唱えています。カメラは子供たちを追って通りを sweeping します。その後、カメラは上へと傾き、@Image 1 から Xin Qiji が現れます。Xin Qiji は首をかしげ、遠くには消えゆく提灯の光の中に一人の男が立っています。このショットは通して連続しています。*
産業製造、エンボディド・インテリジェンス(具身知能)、自動運転などの分野において、Seedance 2.5 は極めて具体的な生産ワークフローへの統合が進んでいます。このモデルは、ロボットの知覚能力や操作スキルを訓練するための高品質な合成動画データを生成可能です。また、産業用シミュレーション、プロセス研修、機器デモンストレーションにも活用されています。
自動運転分野では、極端な気象条件や複雑な交通状況といったロングテール(稀な事象)のシナリオをシミュレートでき、システムテストや訓練のための多様なサンプルを提供します。
あなたのブラウザは動画再生に対応していません。*R2V プロンプト:@Clay Render 1 のカメラワーク、構図、ショットスケール、空間関係、部品位置、モデル構造、組み立て順序、モーションパスを参照し、@Image 1 の素材、照明、色彩、反射、雰囲気を参考に、粘土レンダリングを高級で写実的な自動車組み立てシーケンスに変換してください。
まとめと今後の展望
Seedance 2.5 は、現実世界の理解と描画において大きな一歩を踏み出し、動画生成を単なるクリップレベルの出力から包括的なクリエイティブワークフローへと進化させました。同時に、複雑な動作における物理的妥当性や、複数の主体が関わるシーンの安定性など、まだ改善の余地があることも認識しています。
今後、Seed チームはより一貫性のある物語作りを追求し、直感的な生成・編集体験の提供に努めるとともに、モデルが現実世界の物理法則をさらに深く理解できるよう取り組んでいきます。Seedance モデルが、より生き生きとした表現が可能になり、制御性を高め、ユーザーの意図をより正確に理解するものへと進化することで、多くのユーザーが創造的なアイデアを具現化できる支援を行いながら、幅広い業界ニーズにも応え続けていくことを目指しています。
原文を表示
Today, we are officially launching Seedance 2.5, the new-generation video creation model. Since the release of Seedance 2.0, we have noticed a shift in what users expect from video creation models: from merely generating a clip to completing a creative work. Building on the unified multimodal audio-video joint-generation architecture of Seedance 2.0, Seedance 2.5 centers on foundational generation and reference-based generation, delivering major breakthroughs in long-form storytelling, multimodal reference, and editing. Grounded in real-world use cases, it opens up greater creative imagination and control, and further unlocks productivity.
Key highlights include:
Up to 30 seconds per generation, with multi-round extensions: Seedance 2.5 can generate high-quality, 30-second audio-video clips in a single pass and supports multiple rounds of extension. It also improves shot transitions and scene changes for stronger continuity in longer videos, and delivers notable gains in image, audio, and motion quality, resulting in a more natural, polished visual quality than commonly seen in AI-generated video. As a result, users can produce high-quality multi-minute content with a consistent audiovisual language, bringing a complete story to life in one take.
Fully upgraded multimodal referencing: Users can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. The model also strengthens a range of reference capabilities, including clay render, motion, and creative references, enabling it to better grasp the creator's intent and realize complex ideas that span multiple subjects, scenes, and shot changes.
More precise and stable editing capabilities: Seedance 2.5 offers timestamp-level control for targeted editing of audio and video content, notably improving efficiency and controllability. The model also enhances advanced editing features, such as green screen, camera perspective, and reference-based editing, to meet the rigorous demands of professional, complex fields like film and advertising.
With advancements in long-form storytelling, multimodal reference, and editing, Seedance 2.5 goes beyond longer single-pass video generation. The model better understands creative intent and delivers the journey from idea to finished video with greater control. Now, we'd like to invite you to watch a short creative film, produced end-to-end by Seedance 2.5.
您的浏览器不支持视频播放。Today, Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro, and other platforms, with API access coming soon via BytePlus ModelArk. We invite you to give it a try and share your feedback.
Project homepage:https://seed.bytedance.com/seedance2_5Access:Jimeng Web -> Video Generation -> Select Seedance 2.5Doubao Pro -> Video Generation -> Select Seedance 2.5
30-second long-form storytelling with multi-round extensions: presenting complete stories in a single pass
Seedance 2.5 extends single-pass video generation from 15 to 30 seconds and further strengthens its storytelling in longer videos. Within 30 seconds, the model can organize multiple logically connected shots so that a story unfolds through setup, development, turning points, and resolution, rather than simply extending a single moment. For example, in a one-take clip of a singer's stage performance, the model portrays the full story of the singer interacting with staff in the dressing room, then walking through the backstage corridor, meeting the dancers, and stepping onto the stage with them for the performance, instead of only the moment of walking on stage.
您的浏览器不支持视频播放。
T2V prompt: One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.
Thanks to the model's multi-round extension capability, users can smoothly append subsequent shots to existing video outputs. Throughout the extension process, it maintains the consistency of main characters, environments, and narrative pacing. This allows users to output videos lasting several minutes at once, reducing the effort required to split clips, repeatedly splice footage, and fix transitions.
您的浏览器不支持视频播放。
R2V prompt: Extend the video. Continue from the visuals and subjects in @Video 1 and generate another 30-second clip, keeping the character subjects, scene, visual style, and sound effects consistent. The little boy runs along the train carriage holding a soccer ball. When the subway stops, the side door opens and he immediately dashes out, with the male lead chasing after him. The two run across the platform and out onto the street, startling passersby and vehicles along the way. The male lead finally catches up and grabs him. The boy looks up, aggrieved. The male lead's anger slowly fades; he pats the boy's head and shows a helpless smile.
In terms of visual presentation, the model achieves smoother transitions between camera movements. The main subject remains stable across multiple cuts, and the audio and visuals remain in sync, resulting in highly coherent long-form videos. For example, in a Peking Opera scene, the camera executes a graceful circular pan following the lead actor's flowing sleeves, while the main subject and background remain entirely consistent. The swinging of the sleeves forms natural arcs in the air, closely adhering to real-world physics.
您的浏览器不支持视频播放。
R2V prompt: 16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot from @Image 1, following her water sleeves through the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side in a combat exchange. Consort Yu stands slightly behind and to the side of the Overlord, weaving in water-sleeve movements to set softness against strength. The camera slowly pulls back from a medium-close shot of the warrior to a full stage view. At the end, all three face the audience and strike a synchronized Peking opera finale pose.
Additionally, to address the overly artificial look often seen in AI-generated videos, Seedance 2.5 systematically optimizes details such as object textures, skin and eye features, lighting, and color saturation. The model also minimizes uncontrolled occurrences in subtitles and background music, delivering final products that closely resemble the cinematic quality of live-action footage.
Comprehensive upgrades to multimodal reference, bringing greater control to complex creative tasks
Seedance 2.5 further strengthens its multimodal reference generation capabilities. It allows users to input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass. A larger volume and wider variety of references can better capture the user's intent, producing complex videos with more subjects, richer scenes, and more flexible camera work.
The model comprehensively understands elements such as visual composition, scenes, styles, characters, and props across all materials, applying them to the video generation process as instructed. Even in complex scenarios like multi-character shots or group storytelling, it can preserve the appearances and voices of multiple characters while keeping each subject's characteristics stable.
您的浏览器不支持视频播放。
R2V prompt: A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting and shadows, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly follow @Image 5. Reference @Images 6 to 10 for the rest of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the audience seating. The lead vocalist walks from center stage toward the front edge. The pianist is positioned by the piano. The orchestra is arranged on both sides and toward the rear. The choir stands at the back of the stage. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera naturally moves across the violin, cello, and orchestra as they perform together, with the violin feeling bright and the cello warm. In the latter part, the choir joins in. The lead vocalist briefly makes eye contact with front-row audience members, who respond with a smile and a slight nod. In the closing shot, the camera pulls back. The singing ends, and the audience joins in the applause.
Seedance 2.5 also enhances specific reference capabilities including clay render, motion, and creative referencing, giving users finer control over subjects, actions, and camera work in the frame. For instance, with clay render referencing, users can build a scene's spatial structure, character poses, motion paths, and camera angles using textureless 3D models. The model then uses this structure to generate the video, ensuring that the composition and blocking of complex shots closely match the creator's expectations. Additionally, Seedance 2.5 improves lighting control. By leveraging the spatial information from the clay render, it generates realistic lighting effects that follow physical laws, such as light source direction, color temperature, intensity, and shadow projection. This results in more natural light and shadow in the final output.
您的浏览器不支持视频播放。
R2V prompt: Refer to @Clay Render 1 for camera movement, pacing, shot-size transitions, subject trajectory, and blocking. Refer to @Image 2 for character design, scene, materials, lighting, color, and fairy-tale atmosphere, and render the white model as a dreamy, warm, 3D animated short with a childlike fantasy feel. The story unfolds as follows: flight through a fantasy sky → mythical beasts flying alongside through a sea of clouds → a dive into the ocean → weaving through the deep with manta rays → passing through a mirrored rift in spacetime → picking stars from the cosmos → transforming back into the bedroom → father tucking in the blanket → the picture book closes and holds on the final frame.
More precise and reliable editing for higher creative efficiency
In video creation, users typically need to control pacing during the generation process and also refine details afterward. The exact second an action occurs, the precise timing of a camera cut, and whether a character's movement in a specific clip requires adjustment all profoundly impact the final result. More precise and reliable editing capabilities allow creators to accurately bring their ideas to life, improving efficiency and reducing the cost of repetitive generation.
Seedance 2.5 supports precise content editing via timestamps. During the generation phase, users can use prompts to control the narrative, camera perspective, movement, and overall rhythm for a specific time frame, aligning the output more closely with their creative intent. After generation, users can make targeted modifications to characters, actions, or plot elements within specific clips, all while maintaining continuity and realism before and after the edits.
Seedance 2.5 also elevates multiple editing features, such as green screen editing, camera perspective editing, and reference-based editing, to meet the rigorous demands of professional fields like film and advertising. In green screen editing, for example, the model can replace backgrounds and tell entirely different stories while keeping the main subject intact. Furthermore, it excels at rendering how the subject responds to the physical rules of the new environment. This includes the fluttering direction of clothes, the state of hair, gait rhythm, and lighting interaction, ensuring the subject blends harmoniously with the scene.
您的浏览器不支持视频播放。
R2V prompt: Using @Video 1, render the green-screen background, obstacles, wardrobe, and supporting characters. 0–4s: outdoor training, replace the obstacles with rocks, bricks, tires, and wooden crates. 4–10s: locker room, friends offering encouragement. 10–15s: international match, replace the training poles with original defenders and a goalkeeper, and the protagonist scores. Overall photorealistic, cinematic quality.
您的浏览器不支持视频播放。
R2V prompt: Edit @Video 1. Keep the characters, actions, and visual style unchanged. Adjust only the camera movement. A 15-second segmented camera plan: 0–4s, a micro-FPV move skims tightly past the pan, then follows the popping toast and whip-pans to the coffee; 4–7s, push in and track laterally along the rim of the pan, following the fried egg as it flips up and lands back in place; 7–11s, rapidly rise to a top-down view, then descend at a steady pace, sweeping across the plate and keys; 11–15s, use a handheld close-up to follow the hands with a fast lateral whip, then push in on the breakfast and pull back to a medium two-shot. Keep the entire sequence smooth, continuous, and stable.
Going deeper into broader industry scenarios, continuously exploring real-world value
As the model's capabilities evolve, Seedance 2.5 is reaching deeper into broader industry scenarios such as education and manufacturing. In education, the model has begun to enter real learning settings. For example, Seedance 2.5 can turn the historical context, characters, and storylines behind a lesson into more vivid and immersive visuals. It also helps teachers produce instructional videos more efficiently, turning abstract content — scientific principles, historical events, experimental procedures — into dynamic demonstrations. This not only lowers the barrier to producing educational materials but also allows for highly flexible content customization.
您的浏览器不支持视频播放。*An example of Doubao Learning app's "Doubao Classroom" scenario*
R2V prompt: Expressive Eastern painterly style. A street scene in Lin'an during the Southern Song dynasty. Several children run and shout through the bustling street, chanting, "I turn around, and there he is, where the lantern lights grow dim." The camera follows the children as they run, sweeping past the lively street. The camera then tilts up to reveal Xin Qiji from @Image 1. Xin Qiji turns his head, and in the distance stands a man among the fading lantern lights. The shot stays continuous throughout.
In sectors like industrial manufacturing, embodied intelligence, and autonomous driving, Seedance 2.5 is becoming integrated into highly specific production workflows. The model can generate high-quality synthetic video data that helps train robots' perception and manipulation skills. It is also being utilized for industrial simulations, process training, and equipment demonstrations. For autonomous driving, the model can simulate long-tail scenarios, such as extreme weather and complex traffic conditions, providing more diverse samples for system testing and training.
您的浏览器不支持视频播放。
R2V prompt: Reference the camera work, composition, shot scale, spatial relationships, part positions, model structure, assembly order, and motion paths from @Clay Render 1. Reference the materials, lighting, color, reflections, and atmosphere from @Image 1, and turn the clay render into a high-end, photorealistic car assembly sequence.
Summary and looking forward
Seedance 2.5 marks a significant step forward in understanding and rendering the real world, elevating video generation from clip-level outputs to comprehensive creative workflows. At the same time, we recognize there is still room for improvement, particularly regarding the physical plausibility of complex motions and the stability of scenes involving interactions among multiple subjects.
Looking ahead, the Seed team will continue to explore more coherent storytelling, deliver a more intuitive generation and editing experience, and further deepen the model's grasp of real-world physics. We hope the Seedance models will become more vivid, more controllable, and better at understanding users' intent, helping more users express their creative ideas while continuing to explore and serve broader industry needs.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み