ByteDance Seed、音声・映像を統合した双方向 LLM「SeedRealtime」を発表
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ByteDance Seed Blog
ByteDance Seed は、音声・映像・テキストを統合するネイティブな全双方向 LLM「SeedRealtime」を発表し、リアルタイムでの自然な対話体験と能動的な協働を実現した。
AI深層分析を開く2026年8月5日 13:57
AI深層分析
キーポイント
統合アーキテクチャによる深層理解
音声、映像、時間情報をネイティブに融合させることで、同音異義語の曖昧さ解消や視覚的文脈に基づく正確な解釈を可能にした。
能動的な対話と環境認識
ユーザーからの指示を待たずに視覚的な変化を検知してリマインダーを提供し、ツール呼び出しを自然に組み込むことで受動的反応から協働へ転換した。
自然な会話タイミングの制御
ユーザーの会話状態やペースをリアルタイムで感知して適切なタイミングで発言・停止し、雑音や傍聴者の会話を排除して滑らかな対話を維持する。
並行処理によるリアルタイム判断
知覚、理解、意思決定、発話を連続する音声・映像ストリーム上で並列実行し、聴覚と視覚情報を統合して即座に判断を下す。
会話のリズムと注意制御の課題
動画は常に変化するため、モデルは背景ノイズに反応せず、どの対象に注目し誰の声に耳を傾け、いつ応答すべきかを継続的に決定する必要がある。
重要な引用
SeedRealtime uses a unified architecture to natively fuse audio, video, and text
turning interaction from passive reaction into active collaboration
distinguishing bystander chatter and background noise so that it is not falsely triggered
it lets perception, understanding, decision-making, and expression run in parallel over continuous audio-visual streams
編集コメントを表示
編集コメント
従来のモジュール連結型アプローチからの脱却を示す重要な技術的転換点であり、実社会での自然な AI 対話実現への道筋を明確にした。発表元が業界における大規模展開の先駆者であることを強調しており、今後の市場動向に注目すべき内容だ。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本日、ネイティブな音声・映像対応のフルデュプレックス LLM「SeedRealtime」を正式にリリースします。
オムニモーダル(多様モード)な対話を実現するための重要な一歩として、SeedRealtime は音声、映像、テキストを統一的なアーキテクチャで融合させることで、連続するマルチモーダルストリーム上でのリアルタイム対話を可能にし、「見て、聞いて、話す」全新的な体験を提供します。同モデルは以下の3 つの核心的な突破を実現しています。
- 音声と映像の統合的理解: 音声、視覚情報、時間的要素を深く融合させるネイティブなサポートを備えています。このモデルは、視覚的な文脈を参照することで同音異義語の曖昧さを解消し、目にしたものの時間的な指示を正確に解釈します。これにより、「見たもの」「聞いたこと」「話したこと」が統合されます。
- 能動的な対話: 環境への継続的な意識と、自発的に発言する能力を兼ね備えています。視覚的な変化(例えば、特定のターゲットの出現など)を検知すると、モデルは促されなくてもリマインダーを提供できます。また、ツール呼び出しを応答に織り交ぜることで、受動的な反応から能動的な協力へと対話のパラダイムを変革します。
- 自然な会話タイミング: モデルはユーザーの会話状態とペースをリアルタイムで感知し、適切な瞬間に発言したり、一時停止したり、応答したりします。また、外部ノイズに対する耐性も高く、傍聴者の雑談や背景音を区別して誤作動を防ぐため、会話は滑らかで一貫性を保たれます。
エンドツーエンドの人間評価によると、SeedRealtime は従来のカスケード型モデルと比較して、音声対話のリズムに関する課題を半分に削減しました。このモデルは話すタイミングをより自然に判断できるため、途中での遮断や、沈黙後の遅れた応答、背景雑音による誤作動といった不自然な中断が大幅に減っています。同時に、1 回の会話をスムーズかつ完遂する確率も著しく向上しています。
現在、SeedRealtime は業界で初めて音声・視覚のフルデュプレックス技術の大規模展開を実現し、完全リリースされました。
ネイティブな音声・視覚対話へ:オムニモーダル化に向けて、同時に視聴し発話する
リアルタイムの音声・視覚対話は長らく 2 つのパラダイムの狭間で停滞していました。カスケード型システムは ASR(自動音声認識)、VLM(視覚言語モデル)、TTS(テキスト読み上げ)といった別々のモジュールを連鎖させるため、各段階間に遅延や情報の損失が生じます。一方、エンドツーエンド型のモデルはより流暢ですが、多くの手法では依然として外部の VAD(音声検出器)に話者交代を委ねており、本質的には半デュプレックスの「1 問 1 答」形式にとどまっています。
SeedRealtime の核心的な突破点は、音・視覚・タイミング・表情を単一のエンドツーエンドモデル内で統合した点にあります。すべてを聴いてから見て、最後に答えるのではなく、知覚・理解・判断・発話の各プロセスが連続する音声・映像ストリーム上で並列して動作します。これにより、「聞こえたこと」と「見たこと」がリアルタイムの判断を共同で支えるようになります。
最初の課題は会話のリズムです。自然な発話には間が含まれており、それが返答すべきタイミングを示す役割を果たします。一方、動画は常時オンで絶えず変化しています。モデルは背景ノイズに反応して頻繁に割り込まないようにしつつ、画面で何が起きているかを継続的に理解する必要があります。その際、どのオブジェクトに注目し、誰の声を聞き、今が返答すべきタイミングなのかを判断し続けることが求められます。
より深い課題は、音声と映像の統合モデル化および時間的な整合性です。「これをどうすればいいですか」という発言に対し、モデルは現在のシーン、ジェスチャー、視線、過去の行動を組み合わせて、「これ」が何を指すのかを特定しなければなりません。同音異義語や不明瞭な発話に遭遇した際も、映像情報を活用して曖昧さを解消する必要があります。音声、視覚、時間的情報を統合的にモデル化することで初めて、モデルは見るもの、聞くもの、話すものを真に結びつけることができます。視聴・聴取・発話を一つのリアルタイム意思決定システムに統合すれば、モデルは単に質問を待つだけでなく、シーンの変化を追跡し続け、適切なタイミングで発言できるようになります。
次に、SeedRealtime が実際の現場でどのように機能するか、7 つの具体例を通じて見ていきましょう。
音声・映像の統合的理解:シーンの把握と対象の特定
SeedRealtime は、混雑しノイズの多いオープンな実世界環境において、視覚・音声・時間的情報を整合させ、統合的に理解することが可能です。
4 人の友人が夕食を囲んでいます。周囲には多くの人がおり、会話も重なり合っています。ユーザーが出席者一人ひとりを紹介するにつれ、SeedRealtime は外見から名前と顔を一致させます。金髪の琪琪やメガネのジュリアを認識し、自ら乐乐にも挨拶します。その後の会話を通じて、各人の声をその人物のアイデンティティに結びつけたまま維持します。
次にグループは旅行について話し始めます。一人はビーチで写真を撮り水族館を訪れたいと望み、別の人は暑さと移動の手間を嫌がり、さらに別の人は魚介類アレルギーです。SeedRealtime は誰がどの発言をしたかを区別し、グループ全体の会話に基づき、全員が必要とする条件を満たす旅行プランを提供します。顔の認識、声の識別、そして各人のニーズの理解は、すべて同じ会話の中で単一のモデルによってリアルタイムで処理されます。
四川料理店では、中国語だけのメニューに外国籍の客が困っています。SeedRealtime はその場で料理を特定し、英語でおすすめを紹介します。さらに文化的背景にも触れ、「魚香肉絲(ユシアン・ロウスー)になぜ魚が入っていないのか」や「皮蛋(ピーダン)はどのように作られるのか」などを解説します。
サーバーが料理を運んできて「ご飯に合いますよ」と casually 言うと、モデルは画面に映る料理の文脈を理解し、それをゲストのために翻訳します。視覚情報を一度テキストに変換する必要はありません。同じモデル内で音声と同期されているため、モデルは眼前にあるものに基づいて回答できます。
能動的な対話:常時監視し、適切な瞬間に介入
このモデルは、ユーザーが毎回話し始めるのを待つのではなく、ライブシーンで起こる変化を自ら判断して、介入が必要かどうかを決定します。
河北省博物館の展示を見学している際、ユーザーが「金銀象嵌の虎鹿争食図屏風台が見つかったら教えて」と言いました。カメラが動き続ける中、SeedRealtime はシーンを監視し、その作品に焦点が当たった瞬間に声をかけてリマインダーを送ります。
その後、4 匹の龍と鳳凰が描かれた金銀象嵌方座や長信宮灯などの宝物について視覚的な詳細を基に説明し、鋳造、金銀象嵌、溶接といった工芸技術についても解説します。文脈の中でタスクを持ち続け、対象物が現れた瞬間にリマインダーを送る行為こそが、継続的な視覚知覚と能動的な対話の定義です。
複雑なエスプレッソマシンを操作している際にも、モデルは視覚状態の変化に基づいてリアルタイムでミスを修正し、フィードバックを与えることができます。
ユーザーがコーヒー豆をそのままポートフィルタに注ぐと、モデルは即座に「直接豆を入れるのではなく、まず挽いて粉状にする必要があります」と指摘します。抽出後には、クリームの色やカップ内の液体の量を視覚的に読み取り、「次回は抽出時間を 2〜3 秒短くしてください」と提案します。ユーザーが一つずつ手順を尋ねるのを待たずとも、実際の状況に応じてモデルが判断を下すことができます。
ResNet の論文を学習している際、モデルはネットワーク図を用いてスキップ接続が勾配消失を防ぐ仕組みを説明します。ユーザーが「私の代わりに見ていて、トレーニングパラメータのセクションに達したら教えて」と言うと、ページが素早くめくるのを画面を見守り続け、「3.4 実装」のセクションを正確に見つけると自ら停止し、学習率やモメンタム、重み減衰といった重要なトレーニング設定を読み上げます。連続するフレームの中からターゲットを見つけ、自ら停止することは、対話における真の能動的な協力を示しています。
自然な会話タイミング:雑音の中でいつ話すかを判断
ターンオーバーはもはや外部の VAD ルールに任せるのではなく、SeedRealtime がマルチモーダル情報に基づいて継続的に決定します。適切な時には躊躇せず発言し、沈黙すべき時には決して邪魔をしないため、複雑な環境でもより自然なリズムを維持します。
北京の大兴空港は混雑し、騒がしい場所です。同行者がふと「李さんのフライト」と無関係な雑談の中で口にした際も、モデルはその発言に反応して応答することはありません。
ユーザーが実際に質問したときは、すでに画面からスクロールしてしまった飛行機の情報であっても、以前に見た出発掲示板の情報を引き出してリアルタイムの到着時刻を伝え、さらにオンラインで荷物受取所の場所まで案内します。
その後、二人が歩きながら懐かしむ間も、モデルは目の前のシーンを常に監視し、配車サービスの看板を見つけると、自然にピックアップポイントへの道案内を追加します。騒がしい環境はもはや単なるノイズとしてフィルタリングされるだけのものではなく、雑談から得られる重要な情報も適切に保持され、必要に応じて引き出せるようになります。
お母さんが他の用事で忙しい中、娘に英語の動物の名前を教えるよう SeedRealtime に依頼します。背景ではお父さんが電話中で声が絶えず聞こえますが、モデルはそれらの無関係な会話に惑わされることなく、女の子が指差す先に合わせて発音の修正や例文の作成をリアルタイムで継続して行います。
話すべきタイミングで躊躇せず、干渉があるときも迷うことはありません。
お使いのブラウザでは動画が再生できません。
まとめと展望
SeedRealtime の登場は、音声・映像を統合した対話技術における重要な一歩です。このモデルは、音声・映像・テキストを単一のアーキテクチャ内でネイティブに融合させることで、シーン理解を継続し、リズムを把握し、複数のモダリティにまたがるリアルタイムな応答を実現します。これにより、「見て、聞いて、話す」という体験が、日常の自然な対話として確立されました。
しかし、現実世界の音声・映像コミュニケーションは複雑で連続的です。背景には雑音があり、声がかみ合い、シーンは次々と変化し、ユーザーはいつでも一時停止や追加発言、あるいは割り込みを行います。これらの実世界における課題に対応するため、私たちは引き続き複数の分野で研究開発を推進していきます。
低遅延でより自然な会話タイミングの実現
「聴く・理解する・応答する」という一連の処理を継続的に圧縮し、リアルタイムの会話が持つ微妙なリズム——割り込み、飛び入り、相槌、一時停止など——を単に高速化するだけでなく、適切なタイミングで再現できるようにします。
能動的な知覚と意思決定の強化
視覚と音声を継続的に理解する基盤の上に、いつリマインドすべきか、いつ情報を追加すべきか、ユーザーが本当に必要な情報をいつ提供すべきかを能動的に判断できる機能を構築します。これにより、モデルは「適切な時に話す」だけでなく、「適切な時に行動する」存在へと進化します。
複雑な複数人環境での堅牢性向上
騒音の多い環境や、画面内に複数の人物がいる状況、あるいは多人数による会話においても、誰が発話しているのか、何を注視しているのか、そして誰に反応すべきかを確実に識別できます。これらの能力は、より現実的な生活シーンや業務シナリオへの展開を可能にします。
「対話できる」から「行動できる」へ
ツール呼び出しと実世界を結びつけることで、モデルは単なるコミュニケーションや観察を超え、検索、予約、タスクの実行などを通じてユーザーの支援を行います。リアルタイムのマルチモーダル理解を具体的なアクションへと変換します。
AI の対話体験が、ターンベース型の質問応答に留まることなく、刻々と変化する現実世界の文脈を理解し、タイミングを敏感に捉えて、適切な瞬間にサポートを提供できるものになることを目指しています。
原文を表示
Today, we are officially launching SeedRealtime, a native audio-visual full-duplex LLM.
As a key step toward omni-modal interaction, SeedRealtime uses a unified architecture to natively fuse audio, video, and text, enabling real-time interaction over continuous multimodal streams and delivering a brand-new "watch, listen, and speak" experience. It achieves three core breakthroughs:
- Joint audio-visual understanding: Native support for the deep fusion of audio, visual, and temporal information. The model can resolve homophone ambiguity by drawing on the visual context, and it can accurately interpret temporal references in what it sees, uniting what is seen, heard, and said.
- Proactive interaction: Continuous environmental awareness paired with the ability to speak up on its own. When it notices a change in the visual (such as the appearance of a key target), the model can offer a reminder unprompted, and it can weave tool calls into its responses, turning interaction from passive reaction into active collaboration.
- Natural Conversational Timing: The model senses the user's conversational state and pacing in real time, chiming in, pausing, and responding at the right moments. It is also highly robust to interference, distinguishing bystander chatter and background noise so that it is not falsely triggered, keeping the conversation smooth and coherent.
End-to-end human evaluation shows that, compared with cascaded models, SeedRealtime reduces audio-visual conversational pacing issues by half: the model judges when to speak far more naturally, markedly reducing awkward breakdowns such as being cut off mid-sentence, responding sluggishly after a pause, or being falsely triggered by background noise and chatter. At the same time, the likelihood of completing a single conversation smoothly and fully has also improved significantly.
Currently, SeedRealtime has been fully rolled out, pioneering large-scale deployment of audio-visual full-duplex technology in the industry.
Native audio-visual interaction, toward omni-modal, watch, listen, and speak at once
Real-time audio-visual interaction has long been stuck between two paradigms. Cascaded systems chain together separate modules such as ASR, VLM, and TTS, introducing latency and information loss between stages. End-to-end models are more fluent, but many approaches still rely on an external VAD to decide turns, remaining essentially a half-duplex, one-question-one-answer interaction.
SeedRealtime's core breakthrough is unifying sound, vision, timing, and expression within a single end-to-end model. Rather than listening in full, then looking, and finally answering, it lets perception, understanding, decision-making, and expression run in parallel over continuous audio-visual streams, so that what is heard and what is seen jointly inform every real-time judgment.
The first challenge is conversational rhythm. Speech naturally contains pauses, which can help indicate when it is time to respond. Video, by contrast, is always on and constantly changing. The model must continuously understand what is happening on screen without jumping in too often because of background noise; instead, it keeps deciding which object to focus on, whose voice to listen to, and whether it should respond at this moment.
The deeper challenge lies in joint audio-visual modeling and temporal alignment. When a user says "how do I do this," the model must combine the current scene, gestures, gaze, and prior actions to determine what "this" refers to; when it encounters a homophone or unclear speech, it also has to use the visual scene to disambiguate. Only by modeling sound, visual, and temporal information together can the model truly connect what it sees, hears, and says. Once watching, listening, and speaking are brought into one real-time decision-making system, the model no longer merely waits for a question, but can keep tracking changes in the scene and speak up at the right moment.
Next, let's look at how SeedRealtime performs in real-world settings through seven concrete examples.
Joint audio-visual understanding, understanding the scene, identifying the subject
SeedRealtime can align and jointly understand visual, audio, and temporal information in crowded, noisy, open real-world settings.
Four friends are having dinner: plenty of people and overlapping chatter. As the user introduces everyone present one by one, SeedRealtime matches names to faces by appearance, recognizing light-haired Qiqi and bespectacled Julia and even greeting Lele on its own. It keeps each person's voice tied to their identity throughout the ensuing conversation.
The group then starts talking about travel, one after another: one wants to take photos at the beach and visit the aquarium, another dreads the heat and the effort, and yet another is allergic to seafood. SeedRealtime can tell which line comes from whom and, based on the group’s conversation, provide a travel plan that takes everyone’s needs into account. Recognizing people, telling voices apart, and understanding each person’s needs are all handled in real time by a single model within the same conversation.
您的浏览器不支持视频播放。At a Sichuan restaurant, a foreign diner is stumped by a Chinese-only menu. SeedRealtime identifies the dishes straight from the scene, recommends them in English, and even draws on cultural background to explain "why there is no fish in fish-fragrant pork (yuxiang rousi)" and "how century eggs are made."
As the server sets down a dish and casually says it "goes great with rice," the model understands the line in the context of the dish on screen and translates it for the guest. Visual information does not need to be turned into text first; it is aligned with speech within the same model, so the model can answer based on what is right in front of it.
您的浏览器不支持视频播放。
Proactive interaction, continuously observing, stepping in at the right moment
The model decides for itself whether it needs to step in, based on the ongoing changes in the live scene, rather than waiting for the user to speak first every time.
While touring an exhibition at the Hebei Museum, the user says, "Remind me when you see the gold-and-silver-inlaid bronze tiger-devouring-a-deer screen stand." As the camera keeps moving, SeedRealtime watches the scene, and when it pans across that piece, it speaks up with a reminder.
It then draws on visual details to explain treasures such as the gold-and-silver-inlaid bronze square table base with four dragons and four phoenixes and the Changxin Palace lamp, along with crafts including casting, gold-and-silver inlay, and welding. Holding a task in context and delivering a reminder the moment the target appears is exactly what defines continuous visual perception and proactive interaction.
您的浏览器不支持视频播放。While operating a complex espresso machine, the model can correct mistakes and give feedback in real time based on changes in visual state.
When it sees the user pour whole coffee beans straight into the portafilter, the model quickly points out: "You can't pour the beans in directly; you need to grind them into a fine powder first." After extraction, based on a visual read of the crema's color and the volume of liquid in the cup, it proactively suggests "shortening the extraction time by 2 to 3 seconds next time." Without the user asking step by step, the model can weigh in based on the actual situation.
您的浏览器不支持视频播放。While studying the ResNet paper, the model uses the network diagram to explain how skip connections mitigate vanishing gradients. When the user says, "Keep an eye out for me and remind me when you reach the training-parameters part," it keeps watching the screen as the pages flip quickly, accurately spots the "3.4 Implementation" section, and pauses on its own, then reads out key training settings such as learning rate, momentum, and weight decay. Spotting a target within a continuous stream of frames and stopping on its own demonstrates genuine active collaboration in interaction.
您的浏览器不支持视频播放。
Natural conversational timing, deciding when to speak amid interference
Turn-taking is no longer handed off to external VAD rules; instead, SeedRealtime decides continuously based on multimodal information, never hesitating to chime in when appropriate, and never interrupting when it's better to stay quiet, maintaining a more natural rhythm even in complex environments.
Beijing Daxing Airport is crowded and noisy. When a companion casually mentions "Old Li's flight" in passing chatter, the model is not triggered to reply by this unrelated remark. When the user actually asks, even though the flight information has already scrolled off screen, it can still draw on the departure-board information it saw earlier to give the real-time arrival time, and it goes online to provide the baggage-carousel location.
As the two then walk and reminisce, the model keeps watching the scene in front of it, and when it spots a ride-hailing sign, it naturally chimes in with directions to the pickup point. A noisy environment is no longer merely interference to be filtered out; key information from casual chatter can also be properly retained and drawn on when needed.
您的浏览器不支持视频播放。With mom busy with something else, she asks SeedRealtime to help her daughter learn animal names in English. In the background, dad is on the phone and voices are constant, yet the model is not led astray by this unrelated conversation, consistently following where the girl points to correct her pronunciation in real time and make up example sentences. The model does not hesitate when it should speak, and does not drift when there is interference.
您的浏览器不支持视频播放。
Summary and Outlook
The launch of SeedRealtime marks a key step forward for audio-visual interaction. By natively fusing audio, video, and text within a unified architecture, the model can continuously understand the scene, judge the rhythm, and respond across multimodal streams. With that, "watch, listen, and speak" truly becomes an everyday interactive experience.
Yet audio-visual communication in the real world is complex and continuous: backgrounds are noisy, voices overlap, scenes change, and users pause, add on, or interrupt at any time. To meet these real-world challenges, we will continue to push forward on several fronts:
- Lower latency and more natural conversational timing: Continuously compressing the end-to-end "hear, understand, respond" latency so that the subtle rhythms of real conversation, such as interrupting, jumping in, backchanneling, and pausing, are reproduced naturally, not just faster, but better timed.
- More proactive perception and decision-making: Building on continuous understanding of visual and sound to proactively judge when to remind, when to add information, and when to hand the user exactly the information they need, so that the model not only "speaks when appropriate," but also "acts when appropriate."
- More robust in complex multi-person scenarios: Reliably identifying who is speaking, what they are looking at, and who to respond to in noisy environments, with multiple people in frame and multi-party conversations, bringing these capabilities into more real-life and work scenarios.
- From "able to converse" to "able to act": Connecting tool calls with the real world so that the model can not only communicate and observe, but also help users complete lookups, bookings, and tasks, turning real-time multimodal understanding into concrete action.
We hope AI interaction will no longer be limited to turn-based question answering, but instead understand context in continuously changing real-world situations, grasp timing sensitively, and offer help at the right moment.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み