ByteDance、音声・映像・テキストを統合したリアルタイム LLM「SeedRealtime」を発表
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
ByteDance は音声・映像・テキストを単一アーキテクチャで統合したネイティブ全二重 LLM「SeedRealtime」の正式導入を発表し、従来のカスケード型モデルと比較して対話の遅延や誤作動を大幅に削減する技術的進展を示した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月7日 22:34
AI深層分析
キーポイント
単一アーキテクチャによる統合
知覚、理解、意思決定、応答生成を一つのモデル内で完結させるエンドツーエンドの統一フレームワークを採用し、多段階カスケードシステムに生じる情報損失や誤差蓄積を排除する。
全二重対話の実現
対話状態とタイミングを連続的にモデル化することで、視聴覚ストリーム上でのリアルタイムな双方向インタラクションを可能にし、自然で滑らかな会話を実現する。
性能評価の向上
エンドツーエンドの人間評価により、対話のリズム課題が半減し、割り込みや遅延、誤トリガーが大幅に減少したことが確認された。
多話者環境での同時認識と追跡
複数の話者が重なる会話において、人物の識別、音声の区別、内容の理解を同時に実行し、誰が話しているかを追跡して適切なタイミングで提案を行う。
環境変化への能動的対応
重要なオブジェクトの出現などのシーン変化を検知すると、受動的な応答から能動的な協働へ転換し、ツールを呼び出して即座に対応する。
重要な引用
SeedRealtime, a native audio-visual full-duplex LLM.
natively unifies audio, video, and text within a single architecture
reduces conversational pacing issues by half, significantly decreases interruptions, latency, and false triggers
SeedRealtime can simultaneously identify people, distinguish voices, and understand content, allowing it to track who is speaking, capture the key points of the discussion, and step in with suggestions at the right moment.
編集コメントを表示
編集コメント
ByteDance が単一アーキテクチャによる全二重対話を実現した SeedRealtime の実装を発表したのは、マルチモーダル AI の次世代標準を確立する重要な一手である。従来のカスケード型システムの限界を打破し、より自然な人間と機械の対話を可能にする技術的転換点と言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
テックブログ

概要
本日、私たちはネイティブなオーディオ・ビジュアル双方向 LLM「SeedRealtime」を正式に発表します。
オムニモーダル(多様な感覚)への相互作用を実現するための重要な一歩として、SeedRealtime は音声、映像、テキストを単一のアーキテクチャ内で統合しています。これにより、連続するマルチモーダルストリーム上でのリアルタイムな対話が可能になり、同時に視聴・聴取・発話を行う新しい体験を提供します。このモデルは、オーディオとビジュアルの共同理解能力、能動的な相互作用、そして会話のタイミング制御において大きな進歩を遂げました。
エンドユーザーによる評価では、従来のカスケード型モデルと比較して、SeedRealtime は会話のリズムに関する課題を半減させ、不意の割り込みや遅延、誤作動(false triggers)を大幅に削減しました。また、単発の対話における使いやすさについても、滑らかさと完結性の観点から顕著な改善が見られました。
現在、SeedRealtime は完全に展開が完了しており、業界においてオーディオ・ビジュアル双方向技術の大規模実装を先導しています。
技術的実装
SeedRealtime は、知覚・理解・意思決定・応答生成を単一のモデル内で完結させるエンドツーエンドの統合型オーディオビジュアルモデリングフレームワークを採用しています。これにより、多段階のカスケードシステムに生じる情報損失や誤差の蓄積を低減し、ネイティブなフルデュプレックス対話もサポートします。会話の状態とタイミングを継続的にモデル化することで、より滑らかで自然なリアルタイムでのやり取りを実現します。
エンジニアリング面では、断片化されたオーディオビジュアル入力とストリーミング生成を通じて低遅延のリアルタイム対話を継続的に最適化し、効率的な量子化や推論最適化によりサービング効率も向上させています。
統合型オーディオビジュアル理解
SeedRealtime は、視覚・音声・時間情報をネイティブに統合しています。単に生きているシーンを利用して同音異義語や曖昧な発話を解決するだけでなく、視覚的な変化と参照を継続的に追跡することで、ユーザーの意図をより正確に理解し、見ること・聞くこと・話すことの融合をさらに強化します。

複数の話者が重なる会話においても、SeedRealtime は同時に人物を特定し、声を区別し、内容を理解できます。これにより、誰が発言しているかを追跡し、議論の要点を捉え、適切なタイミングで提案を行うことが可能になります。
能動的な対話
SeedRealtime は、持続的な環境認識と能動的なコミュニケーションを融合させたシステムです。シーンにキーとなるオブジェクトの出現といった変化が検知されると、受動的な対応から能動的な協働へと転換し、必要なツールを呼び出して対話の一部として機能します。

美術館巡りなどのモバイル利用シーンでは、ユーザーの目標を把握し続け、適切なタイミングでリマインドを提供。その場で何が起こっているかを自然に説明しながら、不要な雑音をフィルタリングしてスムーズな対話を維持します。
自然な会話のリズム
SeedRealtime は会話の状態とリズムをリアルタイムで感知し、参加するタイミングや一時停止、応答のタイミングを自然に判断できます。また、周囲の干渉に対して強く、雑談や背景ノイズを識別して誤作動を防ぎ、円滑で一貫性のあるコミュニケーションを実現します。

駅や空港など、複数の話者がいる騒がしい環境でも、SeedRealtime は不要な干渉をフィルタリングし、会話の要点を継続的に理解・保持。必要な際に正確に recall することができます。
今後の展望
現実世界における音声と映像の相互作用は複雑で連続的です。背景にはノイズがあり、声がかぶり、シーンが切り替わり、ユーザーはいつでも一時停止したり情報を追加したり、割り込んだりします。これらの実世界の課題に対応するため、私たちは以下の方向性で取り組みを続けていきます。
低遅延化とより自然なタイミングの実現: 聴取から理解、応答までのエンドツーエンドの遅延をさらに削減し、会話中の割り込み、かぶり、バックチャネル(相槌)、一時停止などを、単に速くだけでなく、タイミングも適切に処理できるようにします。
能動的な知覚と意思決定: 音声と映像の両方で何が起きているかを継続的に理解し、いつリマインドし、いつ補足し、ユーザーが必要とする情報を正確に提供すべきかを判断します。これにより、モデルは単に話すだけでなく、適切な時に行動できるようになります。
複雑な多者間シナリオでの堅牢な性能: 雑音の多い環境や複数人がいるシーン、多者間の会話において、誰が話しているか、何を見ているのか、そして誰に話すべきかを確実に識別します。これらの機能をより多くの実生活や職場の場面で活用できるようにします。
「対話」から「行動」へ: ツール使用と現実世界を結びつけ、モデルがコミュニケーションや観察だけでなく、検索や予約、タスク完了などを実行できるよう支援します。リアルタイムのマルチモーダル理解を具体的なアクションへと変換するのです。
AI の対話体験が、単なるターンベースの質問応答に留まるのではなく、常に変化する現実世界の文脈を理解し、タイミングを敏感に捉え、適切な瞬間に支援を提供できるようになることを願っています。
原文を表示
Tech Blog

Overview
Today, we are officially introducing SeedRealtime, a native audio-visual full-duplex LLM.
As a key step toward omni-modal interaction, SeedRealtime natively unifies audio, video, and text within a single architecture, enabling real-time interaction over continuous multimodal streams and a new experience of watching, listening, and speaking simultaneously. The model advances audio-visual joint understanding, proactive interaction, and conversational timing.
End-to-end human evaluations show that, compared with cascaded models, SeedRealtime reduces conversational pacing issues by half, significantly decreases interruptions, latency, and false triggers, and also markedly improves the usability rate of single-turn conversations in terms of smoothness and completeness.
Currently, SeedRealtime has been fully rolled out, pioneering large-scale deployment of audio-visual full-duplex technology in the industry.
Technology Implementation
SeedRealtime adopts an end-to-end unified audio-visual modeling framework in which perception, understanding, decision-making, and response generation happen within one model, reducing the information loss and error accumulation introduced by multi-stage cascaded systems. It also natively supports full-duplex interaction, continuously modeling conversational state and timing to enable smoother and more natural real-time exchanges.
On the engineering side, we optimize continuously for low-latency real-time interaction through chunked audio-visual input and streaming generation, while improving serving efficiency with efficient quantization and inference optimization.
Joint Audio-Visual Understanding
SeedRealtime natively integrates visual, audio, and temporal information. It not only uses the live scene to resolve homophones and ambiguous speech, but also continuously tracks visual changes and references, enabling more accurate understanding of user intent and a tighter fusion of what is seen, heard, and said.

In overlapping multi-speaker conversations, SeedRealtime can simultaneously identify people, distinguish voices, and understand content, allowing it to track who is speaking, capture the key points of the discussion, and step in with suggestions at the right moment.
Proactive Interaction
SeedRealtime combines persistent environmental awareness with proactive communication. When it detects changes in the scene, such as the appearance of a key object, it can respond proactively and invoke tools as part of the interaction, turning passive response into active collaboration.

In mobile scenarios such as museum visits, SeedRealtime can keep track of the user's goal, provide timely reminders, and naturally explain what is happening in the live scene, while filtering out irrelevant distractions to maintain smooth interaction.
Natural Conversational Timing
SeedRealtime can perceive conversational state and rhythm in real time, naturally deciding when to join in, pause, or respond. It is also resilient to interference, distinguishing side conversations and background noise to avoid false triggers and keep communication smooth and coherent.

In noisy multi-speaker environments such as stations and airports, SeedRealtime can filter out irrelevant interference, continuously understand and retain the key points of a conversation, and recall them accurately when needed.
Outlook
Audio-visual interaction in the real world is complex and continuous: backgrounds are noisy, voices overlap, scenes change, and users may pause, add information, or interrupt at any time. To meet these real-world challenges, we will continue pushing forward in several directions:
Lower latency and more natural timing: Further reducing end-to-end delay from hearing to understanding to responding, so that interruptions, overlaps, backchannels, and pauses in real conversations can be handled more naturally—not just faster, but with better timing.
More proactive perception and decision-making: Continuously understanding what is happening in both audio and video, and deciding when to remind, when to supplement, and when to provide exactly the information the user needs—so the model not only speaks, but also acts when appropriate.
More robust performance in complex multi-party scenarios: Reliably identifying who is speaking, what they are looking at, and who should be addressed in noisy environments, multi-person scenes, and multi-party conversations, bringing these capabilities into more real-life and work settings.
From "conversing" to "acting": Connecting tool use with the real world, so that the model can do more than communicate and observe—it can help users search, book, and complete tasks, turning real-time multimodal understanding into concrete action.
We hope AI interaction will no longer be limited to turn-based question answering, but instead understand context in continuously changing real-world situations, grasp timing sensitively, and offer help at the right moment.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み