ByteDance Seed、シーン単位音声生成モデル「Seed Audio 1.0」を発表
本文の状態
日本語全文を表示中
詳細モードで約10分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ByteDance Seed Blog
Seed Audio 1.0 は音声、効果音、環境音などの要素を個別に扱うのではなく、一つの統一されたフレームワーク内で協調して生成する「シーン」単位のアプローチを採用している。
AI深層分析を開く2026年8月1日 04:41
AI深層分析
キーポイント
シーン単位の統合生成アプローチ
Seed Audio 1.0 は音声、効果音、環境音などの要素を個別に扱うのではなく、一つの統一されたフレームワーク内で協調して生成する「シーン」単位のアプローチを採用している。
表現の幅と一貫性の維持
単独での音声生成においても、抑制されたナレーションから劇的な演技まで幅広い表現範囲をカバーしつつ、キャラクターの声のアイデンティティを一貫して保つことが可能である。
クリエイティブワークフローの変革
各レイヤーを手動で作成・調整する従来の作業から脱却し、対話、雰囲気、アクション、感情を一つの統合された体験として提供することで、サウンドシーンディレクションに近い感覚を実現する。
BytePlus 経由での利用開始
同モデルは現在 BytePlus のコンソールを通じて利用可能となっており、プロジェクトページと試用リンクが公開されている。
協調された音景とタイミング制御
1 つのプロンプトで対話、環境音、効果音を同一の文脈下でモデル化し、自然なタイミングと音色を生成する。キャラクターのセリフの開始時間を 100 ミリ秒単位で指定可能である。
重要な引用
Seed Audio 1.0 is built around that unit of creation: the scene.
It is an audio creation model that generates speech, sound effects, ambience, and other scene-level audio elements within one unified framework.
The goal is to make audio generation feel less like assembling clips, and more like directing a sound scene.
Seed Audio 1.0 models them under a shared scene context, allowing dialogue, ambience, and sound cues to align more naturally in timing, tone, and acoustic space.
編集コメントを表示
編集コメント
音声生成技術が個別の機能統合から「シーン」という文脈単位へ進化するのは、クリエイティブ領域における重要な転換点である。実用化に向けた BytePlus へのアクセス提供は、開発者やコンテンツ制作者にとって即座に検証可能な環境を提供しており、今後の業界標準となる可能性を秘めている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
クリエーターのためのフルシーン音声生成
クリエイターが音を想像する際、それは単なる孤立したファイルの集合として捉えられることはめったにありません。
台詞が語られる前に、その部屋の空気感を感じ取ります。沈黙における緊張感、フレームの外で聞こえる足音、セリフの下に埋もれたアラーム音、場所をリアルにする背景音、そして正確なタイミングで現れる効果音。それらすべてが、クリエイターの頭の中では一体のものとして鳴り響いています。
Seed Audio 1.0 は、まさにその「シーン」という創作の単位を中心に設計されています。
これは、音声、効果音、環境音、その他のシーンレベルの音響要素を一つの統合されたフレームワーク内で生成する音声作成モデルです。プロンプトを入力するだけで、キャラクターのパフォーマンス、タイミング、背景の質感、そしてサウンドデザインが調和した一貫性のある音声を生成できます。
この「シーンレベル」のアプローチは、音声そのものの扱い方にも変化をもたらします。音声単体を生成する場合でも、Seed Audio 1.0 はより広い表現範囲を実現可能です。抑制されたナレーションから高揚する感情まで、自然な会話から劇的な演技までを、声のアイデンティティを保ちながら表現できます。
目指しているのは、音声を断片的に組み立てる作業ではなく、まるでサウンドシーンを演出するかのような体験を提供することです。
Seed Audio 1.0 は現在、BytePlus を通じて利用可能です。
プロジェクトページ: https://seed.bytedance.com/en/seedaudio1_0
今すぐ試す: https://console.byteplus.com/voice/new/setting/activate?projectName=default
シーンが重要な理由
音声生成技術は、声の合成、効果音、環境音、多言語対応など個々の機能において急速に進化を遂げてきました。しかし、クリエイティブな作品においては、これらの要素がいかに相互作用するかが成否を分けます。
セリフには適切な演技が求められます。効果音もタイミングが命です。環境音はシーンを包み込むべきであり、主役を圧迫してはいけません。背景のテクスチャは感情を補強しつつ、テンポを乱さないよう配慮が必要です。
現在でもこうした作業の多くは手動で行われています。各レイヤーを作成し、タイムラインに配置し、タイミングやミックスを調整して、シーンが納得いくまで繰り返すのです。
Seed Audio 1.0 は、より統合的なアプローチを探求するモデルです。複数の音声要素を「同じサウンドシーン」の一部としてモデル化することで、クリエイターは孤立した出力から、セリフ・雰囲気・アクション・感情を一つの整合性のある体験としてまとめた、完全なサウンドシーンの構築へと移行できます。
Seed Audio 1.0 の機能
コーディネートされたサウンドシーンとセリフタイミングの生成
プロンプトには、話者の特徴、感情的なトーン、セリフの話し方、周囲の環境、重要な効果音、そしてシーンの展開方法などを記述できます。
Seed Audio 1.0 はこれらを共通のシーンコンテキストの下でモデル化します。その結果、セリフ・環境音・音声キューが、タイミングやトーン、音響空間においてより自然に同期するようになります。
Seed Audio 1.0 は、単一のプロンプトから複雑なサウンドシーンを生成することが可能です。
ブラウザで動画が再生できません。これにより、Seed Audio 1.0 は物語的な音声、台本付きの対話、ショート動画、広告、ゲームコンテンツ、ポッドキャスト風の制作など、音声がシーンを支える必要があるあらゆるフォーマットで活用できます。
Seed Audio 1.0 はキャラクターのセリフに対してプロンプトレベルでのタイミング制御をサポートしており、現在の精度は 100 ミリ秒刻みです。クリエイターが各セリフの開始時間を指定できるため、動画吹き替え、再録音、広告、タイミングが重要な要素となる台本付き音声など、時間軸を厳密に管理する用途にも適しています。
ブラウザで動画が再生できません。
多彩で一貫性のある声を生成
Seed Audio 1.0 を使えば、テキストによる説明や許可された参照サンプル、あるいはその両方から声を構築できます。各話者ごとに個別のモデルを学習させる必要はありません。
ブラウザで動画が再生できません。ブラウザで動画が再生できません。しかし、声は単なる音色ではありません。実際の創作現場では、同じキャラクターでも状況に応じて、静かなナレーションから緊迫した表現へ、抑制された報告から劇的な対話へと柔軟に変化させながら、依然として同一人物であるという一貫性を保つ必要があります。
Seed Audio 1.0 は音声生成モデルとしてだけでなく、オーディオ作成モデルとして学習されているため、感情やペース配分、周囲の音響効果、物語の意図といった文脈の中で声を理解します。音声のみを生成する場合でも、この背景知識により、感情表現、抑揚、リズム、話し方など、より豊かな表現が可能になります。
より長いコンテンツの場合、Seed Audio 1.0 は一度に最大 2 分間の音声を生成でき、さらに継続的な拡張もサポートします。これにより、キャラクターの声を長く続くシーンや繰り返し拡張する際にも一貫性を持たせ、聴き手に認識されやすくします。
自然な表現で多言語オーディオを制作
グローバルなクリエイターにとって、多言語対応は単なる翻訳の問題ではありません。キャラクターの声は言語を超えて同一性を保ちつつ、それぞれの地域のリズムや発音、ペース、そして感情の表し方に適応する必要があります。
Seed Audio 1.0 は中国語、英語、日本語、韓国語、スペイン語、インドネシア語、ドイツ語、フランス語、タイ語、ベトナム語など 20 以上の言語に対応しています。このモデルは、キャラクターの声の特徴やパフォーマンススタイルを維持したまま、多言語間での一貫性を保ちながら生成が可能です。
これは、グローバル展開、ゲームのローカライズ、ブランドキャンペーン、多言語ポッドキャスト、ショート動画制作において特に役立ちます。一つのクリエイティブコンセプトやキャラクター声を基盤にすれば、各市場向けに音声制作パイプラインをゼロから再構築する必要なく、効率的にコンテンツを展開できます。
您的浏览器不支持视频播放。
音響シーン全体を扱う統合モデル
フルシーンオーディオの生成は困難です。なぜなら、同時に二種類の理解が必要となるからです。
言語レベルでは、モデルはシーン全体を理解する必要があります。誰が話しているのか、どのような感情を帯びているのか、どのタイミングでセリフや効果音が登場し、その瞬間がどのように展開していくのか——これらを把握することが求められます。
音響レベルでは、音声に説得力を持たせるための細部を維持しなければなりません。話者固有の特性、抑揚、質感、インパクト、環境音、そして空間的な連続性です。
Seed Audio 1.0 は、これらの二つのレベルを統合された生成フレームワークで結びつけます。言語モデルが創造意図をシーンレベルの制御パラメータとして構造化し、拡散ベースの音響生成器が高忠実度の潜在空間内で最終的な音声を作成します。この潜在表現は、話者や感情、タイミング、シーンの文脈といった意味指示によって制御可能でありながら、音響の詳細を保持するように設計されています。
これにより、Seed Audio 1.0 は、セリフ、効果音、環境音、その他の音響要素を別々のタスクとしてではなく、同一のシーン内の構成要素としてモデル化することが可能になります。モデルは、声がいかに環境に溶け込むか、効果音がどのようにセリフを支えるか、タイミングの変化がシーンの感情的な形状をどう変えるかを学習します。
評価
Seed Audio 1.0 の性能は、テキストによる音声生成、多様なシナリオでの音響作成、そして多言語対応の三つの領域で評価されました。
A/B 比較の主観的評価では、Seed Audio 1.0 はテキスト記述から声を生成する際に明確な優位性を示しました。また、使いやすさと高品質な出力率の両方が向上しており、音声デザインとパフォーマンスに対する制御力が強化されていることが伺えます。

映画・テレビ番組、TV プログラム、ショートドラマ、アニメーション、ポッドキャストの対話、ライブコマース、オンラインコンテンツ制作、舞台公演、そして音声合成(TTS)など、多様なシーンでの音声生成において Seed Audio 1.0 をテストしました。その結果、ほとんどのシナリオで実用可能な音声が生成される割合は 90% を超えました。

多言語生成においては、人間の評価者によって指示の遵守度と音質の自然さを測定しました。音質の自然さについては、ほとんどの言語で MOS(平均意見スコア)が 4.0 を達成しました。複雑な音声生成タスクにおける指示の遵守度では、ベトナム語を除くすべての評価対象言語で 3.5 を上回りました。

これらの結果は、Seed Audio 1.0 が、感情豊かな声の生成から複雑な多言語の音声シーンまで、幅広い実用的なクリエイティブタスクをサポートできることを示唆しています。
次のステップ
Seed Audio 1.0 は、より広範な音声創作への道程における最初の歩みに過ぎません。
現在、タイミング制御は主にキャラクターのセリフに焦点を当てています。今後は、効果音や環境音、音楽など、より細かな要素への制御も拡大していく予定です。また、長尺で複雑なシーンにおける音声の一貫性向上にも取り組み続けます。
将来に向けては、動画参照などの新たな入力形式に対応し、長尺作品やマルチトラックのオーディオ生成に関する高度な機能の探索を進めます。さらに、多言語翻訳の制御性を高めていくことで、クリエイターが異なる言語間での表現、タイミング、長さの管理をより容易に行えるようにしていきます。
中長期的な方向性は明確です。オーディオ生成は、より表現力豊かで、より制御しやすく、そしてクリエイターの思考プロセスに即したものへと進化していくべきです。
Seed Audio 1.0 は、その一歩となるモデルです。想像したサウンドシーンを完成度の高いオーディオへと変えることを支援します。
原文を表示
Full-scene audio generation for creators
Creators rarely imagine sound as a set of isolated files.
They hear the room before a line is spoken. They hear the pressure in a pause, the footstep just outside the frame, the alarm buried under the dialogue, the background that makes a place feel real, and the cue that arrives at exactly the right moment.
Seed Audio 1.0 is built around that unit of creation: the scene.
It is an audio creation model that generates speech, sound effects, ambience, and other scene-level audio elements within one unified framework. Given a prompt, Seed Audio 1.0 can produce coordinated audio with character performance, timing, background texture, and sound design shaped together.
That scene-level approach also changes what the model can do with speech itself. Even when generating voice alone, Seed Audio 1.0 can produce a broader expressive range — from restrained narration to heightened emotion, from natural conversation to dramatic performance — while keeping the voice identity consistent.
The goal is to make audio generation feel less like assembling clips, and more like directing a sound scene.
Seed Audio 1.0 is now available through BytePlus.
Project page: https://seed.bytedance.com/en/seedaudio1_0Try it now: https://console.byteplus.com/voice/new/setting/activate?projectName=default
Why scenes matter
Audio generation has made fast progress across individual capabilities: voice synthesis, sound effects, ambience, and multilingual speech. But creative work often depends on how these elements interact.
A voice line needs the right performance. A sound effect needs the right moment. Ambience needs to hold the space without overpowering the scene. Background texture needs to support emotion without breaking the pacing.
Today, much of that work still happens manually: creating each layer, placing it on a timeline, adjusting the timing and mix, and repeating until the scene feels right.
Seed Audio 1.0 explores a more integrated path. It models multiple audio elements as parts of the same sound scene, helping creators move from isolated outputs to complete sound moments — scenes that carry dialogue, atmosphere, action, and emotion in one coherent experience.
What Seed Audio 1.0 can do
Generate coordinated sound scenes and control dialogue timing
A prompt can describe the speaker, the emotional tone, the line delivery, the surrounding environment, key sound effects, and how the scene should unfold.
Seed Audio 1.0 models them under a shared scene context, allowing dialogue, ambience, and sound cues to align more naturally in timing, tone, and acoustic space.
Seed Audio 1.0 can create complex audio scenes from a single prompt.
您的浏览器不支持视频播放。This makes Seed Audio 1.0 useful for narrative audio, scripted dialogue, short-form video, advertising, game content, podcast-style production, and other formats where sound needs to carry the scene.
Seed Audio 1.0 supports prompt-level timing control for character dialogue, with current timing precision at 100 ms intervals. Creators can specify when lines should enter, making the model useful for video dubbing, re-voicing, advertising, and scripted audio where timing is part of the deliverable.
您的浏览器不支持视频播放。
Generate voices with range and consistency
Seed Audio 1.0 lets creators shape a voice from a text description, an authorized reference sample, or both — without training a separate model for each speaker.
您的浏览器不支持视频播放。您的浏览器不支持视频播放。But a voice is more than a timbre. In real creative work, the same character may need to move from calm narration to urgency, from restrained reporting to dramatic dialogue, while still sounding like the same person.
您的浏览器不支持视频播放。Because Seed Audio 1.0 is trained not only as a speech generator, but as an audio creation model, it learns voices in the context of scenes: emotion, pacing, surrounding sound, and narrative intent. Even in speech-only generation, this gives the model more room to shape delivery across emotion, prosody, rhythm, and speaking style.
For longer content, Seed Audio 1.0 can generate up to two minutes of audio in a single pass and supports further continuation, helping a character stay recognizable across longer scenes and repeated extensions.
Create multilingual audio with natural expression
For global creators, multilingual audio is not a translation problem alone. A character voice has to carry the same identity across languages while adapting to the rhythm, pronunciation, pacing, and emotional habits of each local audience.
Seed Audio 1.0 supports audio generation across 20+ languages, including Chinese, English, Japanese, Korean, Spanish, Indonesian, German, French, Thai, and Vietnamese. The model can adapt a character voice across languages while preserving its recognizable timbre and performance style.
This is especially useful for global content rollout, game localization, brand campaigns, multilingual podcasts, and short-form video. A team can start from one creative concept or character voice and extend it across markets more efficiently, without rebuilding the entire audio production pipeline for every language.
您的浏览器不支持视频播放。
A unified model for sound scenes
Full-scene audio generation is difficult because it requires two kinds of understanding at once.
At the language level, the model has to understand the scene: who is speaking, what emotion they carry, when each line or sound cue should enter, and how the moment should unfold. At the acoustic level, it has to preserve the details that make audio convincing: speaker identity, prosody, texture, impact, ambience, and spatial continuity.
Seed Audio 1.0 connects these two levels through a unified generation framework. A language model helps structure the creative intent into scene-level controls, while a diffusion-based acoustic generator produces the final audio in a high-fidelity latent space. This latent representation is designed to preserve acoustic detail while remaining controllable by semantic instructions such as character, emotion, timing, and scene context.
This is what allows Seed Audio 1.0 to model speech, sound effects, ambience, and other audio elements as parts of the same scene rather than as unrelated tasks. The model can learn how a voice sits in an environment, how a sound cue supports a line, and how timing changes the emotional shape of a scene.
Evaluation
We evaluated Seed Audio 1.0 across three areas: text-prompted voice generation, multi-scenario audio creation, and multilingual generation.
In A/B subjective evaluations, Seed Audio 1.0 showed clear preference gains in generating voices from text descriptions. It also improved both usability and high-quality output rates, suggesting stronger control over voice design and performance.

For multi-scenario audio creation, we tested Seed Audio 1.0 across a wide range of use cases, including film and television, TV programming, short drama, animation, podcast dialogue, live commerce, online content creation, stage performance, and text-to-speech. Across most scenarios, the usable audio rate exceeded 90%.

For multilingual generation, we used human evaluation to measure instruction following and audio naturalness. On audio naturalness, most languages achieved MOS scores above 4.0. On instruction following for complex audio generation tasks, all evaluated languages except Vietnamese scored above 3.5.

Together, these results suggest that Seed Audio 1.0 can support a broad range of practical creative tasks, from expressive voice generation to complex multilingual audio scenes.
What comes next
Seed Audio 1.0 is an early step toward a broader form of audio creation.
Today, timing control is focused mainly on character dialogue. We will continue extending fine-grained control to more sound elements, including sound effects, ambience, and music. We will also keep improving voice consistency in longer and more complex scenes.
Looking ahead, we plan to support more input modalities, including video references, and to explore more advanced capabilities for long-form and multitrack audio generation. We will also explore controllable multilingual translation, so creators can better manage expression, timing, and duration across languages.
The long-term direction is clear: audio generation should become more expressive, more controllable, and more aligned with how creators actually think.
Seed Audio 1.0 is a step in that direction — helping creators turn imagined sound scenes into finished audio.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み