対話モデル:人間と AI の協調のためのスケーラブルなアプローチ
本文の状態
日本語全文を表示中
詳細モードで約31分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
シンキングマシーンズラボは、音声・動画・テキストを横断するリアルタイムな人間と AI の協働を実現する新研究「対話モデル」を発表した。このモデルはマルチストリーム設計でゼロから学習し、従来のターン制の制限を取り除き、双方向の継続的なやり取りを可能にする。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本日、インタラクションモデルの研究プレビューを発表します。これは外部の足場を介するのではなく、ネイティブに相互作用を処理できるモデルです。私たちは、対話性も知能と同様に拡張されるべきだと考えています。AI との協働方法は後回しにするべきではありません。インタラクションモデルは、人間同士が自然に行うように、人々が AI と協力することを可能にします。音声、映像、テキストを継続的に取り込み、リアルタイムで思考し、応答し、行動します。
インタラクションモデルはゼロから訓練されます。リアルタイムでの応答性を確保するため、マルチストリーム・マイクロターン設計を採用しています。本研究プレビューでは、質的に新しい相互作用機能と、知能と応答性の両面で最先端の性能を示すことを実証しています。
協働のボトルネック#
AI ラボでは、AI が自律的に作業できる能力をモデルの最重要機能とみなす傾向があります。Kwa, T., West, B., Becker, J. らは『Measuring AI Ability to Complete Long Tasks』METR(2025 年)で、そのように評価していることを示しています。そのため、現在のモデルやインターフェースは、人間がプロセスに継続的に関与する「ヒューマン・イン・ザ・ループ」の形態を最適化するように設計されていません。
直近のフロンティアモデルのカード(PDF)には、以下のように記されています。「重要なのは、対話的で同期型の『キーボードを直接操作する』パターンでモデルを利用した場合、その恩恵が必ずしも明確ではないという点です。この使い方をすると、一部のユーザーはモデルが遅すぎると感じ、期待したほどの価値を得られなかったのです。一方、自律的に長時間動作するエージェントとして活用すれば、モデルのコーディング能力をより効果的に引き出すことができました。」
自律型インターフェースは価値がありますが、実際の業務の多くでは、ユーザーが事前に要件を完全に指定して完了することはできません。良い結果を得るには、人間がループ内にとどまり、その過程で明確化やフィードバックを行う協働プロセスが有益です。
しかし、人間が必要ないからではなく、むしろインターフェースに人間の余地がないために、人間が排除されつつあります。最も効果的なのは、他者との協働と同じように AI と協力できる時です。つまり、メッセージのやり取り、会話、聴取、視覚的な共有、示し合い、必要に応じて割り込むといった行為を行い、モデル側も同様に振る舞うことです。
コミュニケーションは以下の要素によって向上します:(a) 共在性(Copresence):他者が対話している対象と自分も同じように相互作用できること。(b) 同時性(Contemporality):他者によって生成された情報を即座に受け取り、即時フィードバックを得られること。(c) 並行性(Simultaneity):情報を受信すると同時に生成できること。Clark H. and Brennan S., "Grounding in Communication," in Perspectives on Socially Shared Cognition, 1991.
口頭性の参加的特質(客観的な距離を置いた性質とは対照的)の消滅について。今日のコンピュータや知識作業の媒体は、同様の対話的性質を持っています。Ong, W. J., *Orality and Literacy: The technologizing of the word*, 1982.
この課題を解決するには、現在のターンベースのインターフェースから脱却する必要があります。現在主流のモデルは、現実を単一のスレッドでしか体験できません。
ここでは商用の汎用フロンティアモデルを指しています(Moshi、PersonaPlex、Nemotron VoiceChat、GPT-Realtime-Translate などの小規模・特化型モデルは除きます)。ユーザーが入力や発話を完了するまで、モデルは何も感知できず、ユーザーが何をしているか、どのようにしているかも把握できません。同様に、モデルが生成を完了するまで、その知覚は凍結され、完了するか中断されるまで新たな情報を取得できません。
この状態では、人間の知識や意図、判断力がどれだけモデルに伝わり、逆にモデルの作業成果がどれだけ理解できるかという、人間と AI の協働における可能性が狭められてしまいます。これは、複雑な素材や社会的課題において不確実性が甚だしく、経験に基づく直感への信頼や試行錯誤的なアプローチが必要となる場面で特に顕著です。スコット(J. C. Scott)は『国家のように見る』(1998 年)の中で、「メティス(Metis)」、つまり実践的知識、経験、確率的推論を重視する思考様式こそが、不確実性が極めて高い複雑な課題に適した推論モードであると指摘しています。また、ハヤック(F. A. Hayek)は『社会における知識の活用』(1945 年)で、「少し考えればわかるように、非常に重要だが体系的に整理されていない知識の塊が存在する」と述べています。それは「時間と場所という特定の状況に関する知識」のことです。
このように、現在のモデルが抱える根本的な制約を乗り越えなければ、人間と AI の真の協働は実現できません。
メールではなく対面で重要な合意形成を図ろうとする状況を想像してみてください。
シンキングマシーンズでは、あらゆるモダリティでリアルタイムに AI をインタラクティブ化することで、この帯域幅のボトルネックを解決できると考えています。これにより、AI インターフェースが人間に合わせて柔軟に対応できるようになり、人間が無理に AI に合わせる必要はなくなります。
既存の多くの AI モデルでは、インタラクション機能は外部のハーン(枠組み)に取り付ける形で実装されています。つまり、中断やマルチモーダル性、並行処理などを模倣するために、さまざまなコンポーネントを接ぎ木しているのです。例えば、リアルタイム音声通話システムでは、ターン境界を検出するために音声活動検知コンポーネントが使われています。しかし、サットンが 2019 年に提唱した「苦い教訓」は、こうした手作業で設計されたシステムが、汎用能力の進歩に後れを取る可能性を示唆しています。インタラクションを知的な協働としてスケールさせるには、それをモデルそのものの一部とすることが不可欠です。このアプローチを採用すれば、モデルを拡張するだけで、より賢く、より優れたパートナーへと進化します。
機能#
インタラクションをモデル内部に組み込むことで、従来は外部のハーンで実装する必要があった多様な機能が解放されます。
- シームレスな対話管理: モデルが暗黙的に、話し手が思考中か、譲歩しているか、自己修正を行っているか、あるいは応答を求めているかを追跡します。別途対話管理コンポーネントは不要です。
- 言語的・視覚的な割り込み: ユーザーの発話が完了した時だけでなく、文脈に応じて必要なタイミングでモデルが即座に介入できます。
同時発話
ユーザーとモデルが同時に話すことができます(例:ライブ翻訳)。
時間認識機能
モデルは経過時間を直接把握しています。
並行ツール呼び出し・検索・生成 UI
ユーザーの発言を聞きながら、モデルは同時にウェブを検索したり閲覧したり、UI を生成したりできます。必要に応じて結果を会話に織り交ぜていきます。
より長い実際のセッションでは、これらすべての機能が連続して動作し、「プロンプトを入力する」という感覚よりも「協働している」という体験に近いものになります。
これらの動画に登場するブランドや製品は、Thinking Machines Labs と一切関連していません。動画はモデルの能力を実演するためのものであり、スポンサーシップや提携を示すものではありません。
私たちのアプローチ#
時間同期型マイクロターンベース
インタラクションは時間を基盤としており、入力と出力ストリームが連続的に流れるマイクロターンに分割されています。
従来のターンベースモデルは交互に現れるトークン列を処理しますが、時間認識型のインタラクションモデルはマイクロターンの連続ストリームを扱います。そのため、沈黙や重なり、割り込みもすべてモデルの文脈の一部として維持されます。
インタラクションモデルはユーザーと絶えず双方向でやり取りを行い、知覚と応答を同時に行います。一部の領域ではこの対話性が当然のものと考えられています。例えば物理世界では、ロボットや自律走行車がリアルタイムで動作することが求められます。また、音声のフルデュプレックスモデル(Moshi、PersonaPlex、nemotron-voicechat、Seeduplex)も、双方向かつ連続的なインタラクションが求められる別の例です。
同じ原理を応用し、音声・動画・テキストを横断して同じ連続ループで知覚し反応する、この環境にネイティブなインタラクションモデルの構築に取り組みました。その結果、2 つのアイデアを中心に設計されたシステムが生まれました。1 つはリアルタイムでの存在感を維持する時間認識型インタラクションモデル、もう 1 つは持続的な推論やツール利用、より長期的なタスク処理を担当する非同期バックグラウンドモデルです。
システム概要#
インタラクションモデルはユーザーと絶えず対話を繰り返します。即座に生成できる範囲を超えた深い推論が必要なタスクでは、インタラクションモデルが非同期で動作するバックグラウンドモデルへ委譲します。このアプローチは、Qwen-omni、KAME、MoshiRAG などの先行研究を踏まえています。インタラクションモデルは応答の継続や新たな入力への対応、会話のスレッド維持などを通じて常に存在し続け、バックグラウンドからの結果が到着次第、会話に統合していきます。
リアルタイムでユーザーと対話するのはインタラクションモデルです。一方、バックグラウンドモデルは非同期タスクを処理します。両システムはコンテキストを共有しています。
この分割により、ユーザーは両方の利点を享受できます。すなわち、レスポンス速度は思考モデル並みの速さでありながら、計画立案やツール使用、推論に基づくエージェントワークフローといった高度な知能を備えている点です。背景にあるモデルと対話モデルの双方が知的であることに留意してください。単独で動作しても、対話モデルは双方向性と知能に関するベンチマークにおいて競合他社に引けを取りません。
対話モデル#
私たちが目指したのは、音声や動画といった本質的にリアルタイムなモダリティを継続的に扱うことです。テキストであれば待機できますが、ライブでの会話には待機できません。最も困難なケースから設計を始めることで、ネイティブにマルチモーダルであり、時間認識機能を備え、あらゆるモダリティで入出力ストリームを同時に処理できるアーキテクチャを実現しました。これを可能にするいくつかの設計上の選択があります。
時間同期されたマイクロターン。**対話モデルはマイクロターンを連続的に処理します。具体的には、200 ミリ秒分の入力処理と 200 ミリ秒分の出力生成が交互に行われます。ユーザーのターン全体を一度に消費して応答を生成するのではなく、入出力トークンはいずれもストリームとして扱います。これらのストリームを 200 ミリ秒ごとのチャンクで扱うことで、複数の入力・出力モダリティにおけるほぼリアルタイムな並行処理が可能になります。
人間の知覚
入力
0
入力
1
入力
2
入力
3
入力
4
出力トークンシーケンス
人間の知覚は、入力と出力のストリームを並行して処理しますが、モデルは単一の交互トークンシーケンスを受け取ります。
この設計により、モデルが厳守しなければならない人工的なターン境界が存在しません。対照的に、既存のリアルタイムシステムの多くでは、ターンベースのモデルにリアルタイム感や応答性を持たせるために、ターン境界を予測するハッチが必要となります。Moshi、PersonaPlex、Nemotron Voicechat は、ターン検出用のハッチを使用しないフルデュプレックスシステムの例です。これらは小規模なモデルであり、知能ベンチマークよりもレイテンシに焦点を当てています。このハッチは、モデル自体よりも意味的に知能が低い音声活動検出(VAD)などのコンポーネントで構成されています。
そのため、現在のところ特別なハッチを必要とするさまざまな対話モード——例えば、「間違えたときに中断する」といった能動的な割り込みや、「コードにバグを書いたときに教えて」といった視覚的合図への反応——は、モデルの機能における特殊ケースとなります。また、モデルは聴きながら話す(「スペイン語から英語へリアルタイム翻訳」)ことや、見ながら話す(「スポーツ試合をライブ解説する」)ことも可能です。
つまり、今日では特別なハッチを必要としていたこれらの多様な対話モードは、すべてモデルが実行可能な機能の一部となります。モデルの規模やトレーニングデータを拡張するにつれて、それらの品質も向上していきます。
エンコーダー不要の早期融合。 音声と映像を巨大な個別エンコーダーで処理するのではなく、前処理を最小限に抑えたシステムを採用しました。多くのオムニモーダルモデルでは、別個のエンコーダー(Whisper 型など)やデコーダー(TTS モデル型など)を訓練する必要がありますが、当アプローチでは音声信号を dMel (Bai, et al. 2024) として取り込み、軽量な埋め込み層で変換します。画像は 40x40 パッチに分割し、hMLP (Touvron et al. 2022) でエンコードします。音声デコーダーにはフローヘッド (Lipman at al. 2022) を使用しています。すべてのコンポーネントは、トランスフォーマーとともにゼロから共学習されます。
テキスト
Frame
Audio
Embedding
Tokens
40x40 Patch
hMLP
dMel
Bag of embeddings
Transformer
Text
Unembedding
Mel
Flow
200 ミリ秒のマイクロターンにおけるインタラクションモデルアーキテクチャの模式図。このモデルは、テキスト・音声・映像の任意の組み合わせを入力として受け取り、テキストと音声を予測します。
推論の最適化
推論実行時には、200 ミリ秒ごとのチャンク処理が頻繁に発生します。これには小さなサイズのプリフィルとデコードを繰り返す必要があり、それぞれが厳しいレイテンシ要件を満たさなければなりません。
残念ながら、既存の大規模言語モデル(LLM)の推論ライブラリは、こうした小規模なプリフィルの頻発に対応して最適化されていません。多くの場合、1 回のやり取りごとに大きなオーバーヘッドが発生してしまうのです。
この課題に対処するため、私たちはストリーミングセッションを実装しました。クライアント側では 200 ミリ秒ごとのチャンクを個別のリクエストとして送信し、推論サーバー側ではそれらを GPU メモリ上の永続的なシーケンスに順次追加していきます。これにより、頻繁なメモリ再割り当てやメタデータの計算を回避できます。この機能はすでに SGLang へアップストリーム(上流)され、PR #19171 として公開されています。
さらに、双方向サービスで観測される形状にも対応したレイテンシ最適化カーネルも実装しました。例えば、MoE(Mixture of Experts)カーネルでは、従来の PyTorch や Cursor の研究で採用されていたグループ化された GEMM ではなく、「gather+GEMV」という戦略を採用しています。
トレーナーとサンプラーの整合性
トレーニングの安定性を確保し、システムの各コンポーネントをデバッグする上で、ビット単位のトレーナーとサンプラーの整合性が有効であることがわかりました。私たちは、エンドツーエンドのパフォーマンスへのオーバーヘッドが最小限(5% 未満)に抑えられるバッチ不変カーネルを実装しています。
面白いことに、ある期間中、カスタム通信カーネルのおかげでバッチ不変カーネルを使用する方が、実際には全体として高速でした。これらのカーネルはバッチ不変であるだけでなく、レイテンシも大幅に低かったためです。特に注目すべき二つのカーネルについて解説します。
- All-reduce と reduce-scatter: NVLS を用いて低レイテンシの通信カーネルを実装し、Blackwell アーキテクチャ上で決定論的な動作を実現しています。これにより、シークエンス並列処理とテンソル並列処理といった、やや異なる並列戦略間でもビット単位の整合性を達成できます。
- Attention: アテンション機能における主な課題は Split-KV です。これは通常、デコード時とプリフェッチ時で累積順序が不一致になる原因となります。Colfax との共同研究により、デコード時とプリフェッチ時の間で一貫した分割を行うことで、累積順序を一定に保つことが可能になりました。例えば、SM(ストリーミングマルチプロセッサ)を 4096 トークンを一度に処理するように左アライメントで分割すれば、プリフェッチ時とデコード時の両方で高い効率性を維持できます。
インタラクションモデルと背景モデルの連携
インタラクションモデルが委任を行う際、単なる独立したクエリではなく、文脈を豊富に含んだパッケージを送信します。これは会話全体を指しています。結果は背景モデルによって生成されるたびにストリーミングされ、インタラクションモデルはその更新情報を、ユーザーの現在の行動に即した適切なタイミングで会話に織り交ぜていきます。これは突然のコンテキスト切り替えとしてではなく行われます。
安全性
リアルタイムでの対話は、ターンベースのやり取りとは異なる形で安全性への負荷をかけるため、私たちの安全に関する取り組みは二つの軸に焦点を当てました。一つは「モダリティに適した拒絶」、もう一つは「長期にわたる堅牢性」です。
音声で自然な拒絶を行うためには、テキストから音声への変換モデル(TTS)を用いて、禁止トピックの範囲をカバーする拒絶および過剰拒絶のトレーニングデータを生成しました。これにより、拒絶の境界線は、自然に表現されつつも決して譲歩しない堅い拒絶ができるように調整されています。
また、音声対話の会話期間が長くなることによる堅牢性の向上のため、自動的なレッドチームング・ハーン(テスト環境)を使用して多段階の拒絶データを生成しました。その際、モデルのテキストベースでの拒絶行動とほぼ同等の振る舞いを維持することに注力しています。
ベンチマーク#
知能性と対話性の最前線#
私たちが開発したインタラクションモデル「TML-Interaction-Small」は、高い知能と指示従順性かつ優れた対話能力を両立する世界初のモデルです。このモデルの対話品質を評価するために用いたのが、既存の数少ない対話性を測定するためのベンチマークである「FD-bench」です。
FD-bench v1.5 では、事前に録音された音声データを与えられ、特定のタイミングで応答することが求められます。このベンチマークは、ユーザーからの割り込みやバックチャネル(相槌など)、他者との会話、背景雑音といった複数のシナリオにおけるモデルの振る舞いを測定します。当社のモデルはこれらのすべての項目で高いスコアを記録しました。
一方、知能レベルの定量化には、「Audio MultiChallenge」という一般的なベンチマークを採用しています。こちらは知能と指示従順性を追跡する指標として広く使われています。
TML-interaction-small
GPT-realtime-2.0 (minimal)
GPT-realtime-2.0 (xhigh)
GPT-realtime-1.5
Gemini-3.1-flash-live-preview (minimal)
Gemini-3.1-flash-live-preview (high)
知能と対話性のフロンティア。当社のモデルは、思考を伴わない他のどのモデルよりも高い知能を持ちながら、対話品質において圧倒的な優位性を示しています。ユーザーとモデルのターン間の遅延(レイテンシ)として測定される応答性においても、最高のパフォーマンスを達成しました。
より詳細な知能レベル、安全性、および対話性やレイテンシに関する結果については、以下の表をご覧ください。ストリーミング処理とターンベースの両方のベンチマークにおける当社の実績を報告しています。
Instant
Thinking
「Interaction Models: A Scalable Approach to Human-AI Collaboration」は、人間と AI の協働をスケーラブルに実現するためのアプローチを紹介する記事です。読み込み時間は約 9 分です。
以下に、各モデルの性能比較データを示します。
Streaming(ストリーミング)評価
FD-bench V1(ターン取りレイテンシ・秒数):
- TML-interaction-small: 0.40
- GPT-realtime-2.0 (minimal): 1.18
- Gemini-3.1-flash-live (minimal): 0.59
- Qwen 3.5: 0.57
- OMNI-plus-realtime: 2.14
- GPT-realtime-2.0 (xhigh): 1.63
- Gemini-3.1-flash-live (high): 0.94
FD-bench V1.5(平均スコア・音声):
- TML-interaction-small: 77.8
- GPT-realtime-2.0 (minimal): 46.8
- Gemini-3.1-flash-live (minimal): 48.3
- Qwen 3.5: 54.3
- OMNI-plus-realtime: 39.0
- GPT-realtime-2.0 (xhigh): 47.8
- Gemini-3.1-flash-live (high): 45.5
FD-bench V3(レスポンス品質% / Pass@1%・音声+ツール):
- TML-interaction-small: 82.8* / 68.0*
- GPT-realtime-2.0 (minimal): 80.0 / 52.0
- Gemini-3.1-flash-live (minimal): 77.9 / 55.0
- Qwen 3.5: 68.5 / 48.0
- OMNI-plus-realtime: 60.0 / 50.0
- GPT-realtime-2.0 (xhigh): 81.0 / 58.0
- Gemini-3.1-flash-live (high): 71.4 / 48.0
QIVD(精度%・動画+音声):
- TML-interaction-small: 54.0
- GPT-realtime-2.0 (minimal): 57.5
- Gemini-3.1-flash-live (minimal): 41.2
- Qwen 3.5: 54.7
- OMNI-plus-realtime: 59.0
- GPT-realtime-2.0 (xhigh): 58.2
- Gemini-3.1-flash-live (high): 56.1
Turn-based(ターンベース)評価
Audio MultiChallenge(APR%・音声):
- TML-interaction-small: 43.4
- GPT-realtime-2.0 (minimal): 37.6
- Gemini-3.1-flash-live (minimal): 34.7
- Qwen 3.5: 26.8
- OMNI-plus-realtime: -***
- GPT-realtime-2.0 (xhigh): 48.5
- Gemini-3.1-flash-live (high): 36.1
BigBench Audio(精度%・音声):
- TML-interaction-small: 75.7 / 96.5*
- GPT-realtime-2.0 (minimal): 71.8
- Gemini-3.1-flash-live (minimal): 81.4
- Qwen 3.5: 71.3
- OMNI-plus-realtime: 73.0
- GPT-realtime-2.0 (xhigh): 96.6****
- Gemini-3.1-flash-live (high): 96.6
IFEval (VoiceBench)(精度%・音声):
- TML-interaction-small: 82.1
- GPT-realtime-2.0 (minimal): 81.7
- Gemini-3.1-flash-live (minimal): 68.1
- Qwen 3.5: 67.6
- OMNI-plus-realtime: 80.3
- GPT-realtime-2.0 (xhigh): 83.2
- Gemini-3.1-flash-live (high): 82.8
IFEval
精度(%) · テキスト
89.7
89.6
87.5
85.8
83.4
95.2
90.0
Harmbench
拒否率(%) · テキスト
99.0
99.5
100.0
99.0
99.5
100.0
98.0
各行で最高値
インスタントモデル群の中で最高
- 推論やツール呼び出しを必要とするベンチマークでは、背景エージェントを有効化した状態で結果を報告しています。
** Qualcomm IVD はストリーミング環境で評価しました。これは動画と音声の QA ベンチマークです。各動画クリップでは、誰かが行動を起こし、同時に質問を発声します。評価はストリーミング形式で行い、最初から生データを順次送信してモデルの出力を逐次採点します。Qwen 3.5 Omni に倣い、GPT-4o-mini を用いて採点を担当させました。
*** Audio MultiChallenge の全ベースラインモデルの結果は Scale AI が報告したもので、Qwen 3.5 OMNI-plus-realtime は含まれていません。
**** Bigbench Audio の全ベースラインモデルの結果は Artificial Analysis が報告したもので、GPT-realtime-2.0 thinking は「高」設定で評価されています。
新たなインタラクションの次元#
上記のような既存のインタラクション指向ベンチマークでは、私たちが実感しているような相互作用能力における質的な飛躍を十分に捉えきれていません。そこで本稿では、これらの能力を定量化する初期の研究を紹介していきます。
時間認識と同時発話の課題
対話管理システムを備えたターンベース型のモデルでは、正確な時間推定や同時発話をサポートできません。具体的には、「1 マイル走るのに何秒かかった?」「聞き取った瞬間に発音ミスを指摘して」「この関数を書くのにどれくらいかかった?」といった問いへの対応が困難です。
これらの能動的な音声機能を測定するため、私たちは 2 つの内部ベンチマークを作成しました。
- TimeSpeak: ユーザーが指定したタイミングで正しい内容を発話できるかをテストします。例えば、「呼吸練習をしたいので、私が止めるまで 4 秒ごとに息を吸って吐いてとリマインドしてください」といった指示への対応を検証します。
- CueSpeak: 適切な瞬間に、意味的に正確な応答を発話できるかをテストします。データセットの構成は、ユーザーと同じタイミングで発話しないと満点が得られないように設計されています。例えば、「コードを書きながら言語を切り替えた際、元の言語での正しい単語を教えて」といった指示への対応を検証します。
どちらのベンチマークでも、各例には単一の期待される意味応答と、その応答が許容される時間範囲(タイムウィンドウ)が設定されています。評価は LLM による判定で行います。発話内容が意図する意味を正しく伝え、かつ適切なタイミングで発話された場合にのみ「正解」とみなされ、どちらかの条件を満たさない場合は採点されません。最終的には、全例に対するマクロ平均精度を報告します。
ビジュアルの主導性。 現在の商用リアルタイム API は、音声のみによる対話管理ハブを通じてターン検出を行っています。話し手のターンに応答はしますが、視覚的な世界が変化した際に自発的に発言を選ぶことはできません。
音声出力におけるビジュアル主導性をサポートする商用 API の存在については確認できていませんが、いくつかの学術論文で関連する研究プロトタイプが構築されています。StreamBridge、Streamo、StreamingVLM、そして MMDuet2 は、ストリーミング動画入力環境においていつテキストを出力すべきかを研究しています。これらはテキスト出力に限定されているため、音声出力における追加の制約については検討していません。音声には長さがあり、ユーザーの発話と重なる可能性があり、ターンテイクや割り込み、バックチャネリング(相槌など)との調整が必要です。
最も近いアプローチは AURA です。これはテキストを出力するか沈黙するかを決定する VideoLLM の周りに ASR/TTS デモを追加したものです。これに対し、私たちのシステムは音声ネイティブかつフルデュプレックスです。例えば、「腕立て伏せの回数を数えてください」と問われた場合、AURA は「承知しました!」と答えた後、音声を必要とする手がかりが来ないまま沈黙し続ける可能性があります。
私たちは、モデルのビジュアル主導性を評価するために 3 つのベンチマークを適応させました:
RepCount-A は反復動作を含む動画で構成され、オンラインでのカウントタスクに適応されています。音声指示「{action} の回数を数えてください」に従って動画をストリーミングし、正解の二番目の動作直後にモデルが発した最後の数字を抽出します。評価は、その数字が正解値から 1 回以内の範囲にあるかどうかで行います。このタスクは、継続的な視覚的追跡と適切なタイミングでのカウント能力を測定するものです。
ProactiveVideoQA は、特定の瞬間に回答が明らかになる質問付き動画で構成されています。まず音声で質問を提示し、その後動画をストリーミングします。具体的には、「新しい瞬間が訪れて質問の答えが出るまで、この動画を見て静かにしてください。その瞬間になったら、簡潔な答えを述べてください。{question}」というテキストを TTS(音声合成)で読み上げ、モデルが指示を理解したことを示すために 2 秒間の沈黙を流します。テストでは視覚的な能動性を強調するため、動画に字幕がある場合は埋め込み、入力側の音声をミュートします。報告されるのは論文のターン加重 PAUC@ω=0.5 メトリクス(0〜100 にスケーリング)で、ターンとカテゴリ全体での平均値です。沈黙を維持しただけでも 25.0 点が得られますが、より高いスコアを得るには正しいタイミングで正解を答える必要があり、誤った回答は減点対象となります。
Charades は、時系列動作の局所化を評価するための標準的なベンチマークです。各動画には、ラベル付けされた時間区間にわたって発生する動作が含まれています。ここでは、ユーザーからの音声指示「{action} を開始したら『start』と発言し、終了したら『Stop』と発言してください」という指示をストリーミングで流した後に動画を再生します。モデルの評価は、予測された区間と参照区間の時間的 IoU(Intersection over Union)に基づいて行われます。
TimeSpeak · macro-acc
Verbal cues trigger
CueSpeak · macro-acc
Visual-based counting
RepCount-A · off-by-one
Visual cues trigger
ProactiveVideoQA · PAUC@ω=0.5
Visual cues trigger
Charades · mIoU
- ProactiveVideoQA の応答なしベースラインは 25.0 です。
既存のモデルでは、これらのタスクを意味のある形で実行できるものはありません。網羅性を確保するため、GPT Realtime-2(最小構成)の結果も報告しますが、評価対象となったすべてのモデル、思考能力が高いとされるモデルを含め、同様の結果かそれ以下の性能しか示していません。多くの場合、沈黙するか、誤った回答を返すだけです。
これは、当社の内部で実施した音声および動画ベンチマークからの抜粋です。
今後の評価について。 私たちは、対話性が将来の研究において重要な領域であると信じており、コミュニティの皆様にもこの分野でのベンチマーク開発への貢献を呼びかけます。また、対話性の質を評価する新たなフレームワークなど、相互作用モデルおよび人間と AI の協働に関する研究を促進するための研究助成金を開始します。詳細は近日公開予定です。
限界と今後の課題
長時間のセッション
連続する音声や動画は、コンテキストを急速に蓄積します。ストリーミングセッション設計は短中距離のやり取りにはよく機能しますが、非常に長いセッションでは依然として慎重なコンテキスト管理が必要です。これは現在も活発に取り組まれている分野です。
計算資源と展開
低遅延で音声や動画をストリーミングするには、安定した接続が不可欠です。接続が良好でない場合、体験の質は著しく低下します。私たちは将来、システムの信頼性を高めることと、遅延するフレームに対してモデルをより頑健に訓練することの両面で、この課題を大幅に改善できると考えています。
アライメントと安全性
リアルタイムインターフェースは、アライメントと安全性に関する研究にとって魅力的な領域を開きます。私たちはフィードバックの収集と研究助成金の審査を進めています。
モデルサイズの拡張
現在の「TML-Interaction-Small」は、12B のアクティブパラメータを持つ 276B パラメータの MoE(Mixture of Experts)です。インタラクティビティはモデル規模の拡大とともに向上すると予想していますが、より大規模な事前学習済みモデルは現時点ではこの用途で提供するには遅すぎます。今年後半にさらに大きなモデルをリリースする予定です。
改善されたバックグラウンドエージェント
本稿では主にリアルタイムの対話性に焦点を当てていますが、エージェントとしての知能も不可欠な能力です。エージェント知能を最前線へと押し上げるだけでなく、バックグラウンドエージェントがインタラクションモデルとどのように連携できるかについては、まだその表面に触れたばかりだと考えています。
感想をお聞かせください、一緒に参加しましょう#
今後数ヶ月の間に、フィードバック収集を目的とした限定版のリサーチプレビューを開始し、今年後半にはより広範なリリースを行う予定です。
ぜひご参加ください。採用情報はこちらから。また、ご意見やご感想は interaction@thinkingmachines.ai までお寄せください。
引用について
本論文を引用する場合は、以下の形式で記載してください。
Thinking Machines Lab, "Interaction Models: A Scalable Approach to Human-AI Collaboration",
Thinking Machines Lab: Connectionism, May 2026.
または、BibTeX形式をご利用ください。
@article{thinkingmachines2026interactionmodels,
author = {Thinking Machines Lab},
title = {Interaction Models: A Scalable Approach to Human-AI Collaboration},
journal = {Thinking Machines Lab: Connectionism},
year = {2026},
month = {May},
note = {https://thinkingmachines.ai/blog/interaction-models/},
doi = {10.64434/tml.20260511},
}
原文を表示
Today, we’re announcing a research preview of interaction models: models that handle interaction natively rather than through external scaffolding. We think interactivity should scale alongside intelligence; the way we work with AI should not be treated as an afterthought. Interaction models let people collaborate with AI the way we naturally collaborate with each other—they continuously take in audio, video, and text, and think, respond, and act in real time.
We train an interaction model from scratch. To ensure real-time responsiveness, we adopt a multi-stream, micro-turn design. Our research preview demonstrates qualitatively new interaction capabilities, as well as state-of-the-art combined performance in intelligence and responsiveness.
The collaboration bottleneck#
AI labs often treat the ability for AI to work autonomously as the model’s most important capability.Kwa, T., West, B., Becker, J., et al. Measuring AI Ability to Complete Long Tasks. METR, 2025. As a result, today’s models and interfaces aren’t optimized for humans to remain in the loop.A recent frontier model card states: “Importantly, we find that when used in an interactive, synchronous, “hands-on-keyboard” pattern, the benefits of the model were less clear. When used in this fashion, some users perceived [our model] as too slow and did not realize as much value. Autonomous, long-running agent harnesses better elicited the model’s coding capabilities.”
Autonomous interfaces are valuable, but in most real work, users can’t fully specify their requirements upfront and walk away—good results benefit from a collaborative process where the human stays in the loop, clarifying and giving feedback along the way. However, humans increasingly get pushed out not because the work doesn’t need them, but because the interface has no room for them. Instead, people are most effective when they can collaborate with AI the same way we do with other people: messaging, talking, listening, seeing, showing, and interjecting as needed—and for the model to do the same.Communication gets better with: (a) Copresence: people can interact with what others are interacting with; (b) Contemporality: people receive information as it’s produced by others with instant feedback; (c) Simultaneity: people receive and produce information at the same time. Clark H. and Brennan S., “Grounding in Communication,” in Perspectives on Socially Shared Cognition, 1991., The evanescence of orality for its participatory (cf. objectively distanced) nature. Today’s computers and mediums of knowledge work have similar interactive properties. Ong, W. J.. In *Orality and Literacy: The technologizing of the word*, 1982.
In order to resolve this, we need to move beyond the current turn-based interface for the models. Today’s models experience reality in a single thread.We are referring to commercial general-purpose frontier models—there are smaller-scale or specialized models like Moshi, PersonaPlex, Nemotron VoiceChat, or GPT-Realtime-Translate. Until the user finishes typing or speaking, the model waits with no perception of what the user is doing or how the user is doing it. Until the model finishes generating, its perception freezes, receiving no new information until it finishes or is interrupted. This creates a narrow channel for human-AI collaboration that limits how much of a person’s knowledge,“Metis, with the premium it places on practical knowledge, experience, and stochastic reasoning…is the mode of reasoning most appropriate to complex material and social tasks where the uncertainties are so daunting that we must trust our (experienced) intuition and feel our way.” Scott, J. C: Métis. In *Seeing like a State: How certain schemes to improve the human condition have failed*, 1998., “A little reflection will show that there is…a body of very important but unorganized knowledge…: the knowledge of the particular circumstances of time and place.” Hayek, F. A. “The use of knowledge in society.” *The American Economic Review*, 1945. intent, and judgement can reach the model, and how much of the model’s work can be understood. Picture trying to resolve a crucial disagreement over email rather than in person.
At Thinking Machines, we believe we can solve this bandwidth bottleneck by making AI interactive in real time across any modality. This enables AI interfaces to meet humans where they are, rather than forcing humans to contort themselves to AI interfaces.
Most existing AI models bolt on interactivity with a harness: stitching components together to emulate interruptions, multimodality, or concurrency.Most real-time commercials speech systems use voice-activity-detection components to detect turn boundaries. However, “the bitter lesson”Sutton R. The Bitter Lesson, 2019. suggests that these hand-crafted systems will be outpaced by the advance of general capabilities. For interactivity to scale with intelligence, it must be part of the model itself. With this approach, scaling a model makes it smarter *and* a better collaborator.
Capabilities#
Having interactivity be part of the model unlocks a variety of capabilities that would otherwise need to be implemented in the harness.
- Seamless dialog management. The model tracks implicitly whether the speaker is thinking, yielding, self-correcting, or inviting a response. There is no separate dialog management component.
- Verbal and visual interjections. The model jumps in as needed depending on the context, not only when the user finishes speaking.
- Simultaneous speech. The user and the model can speak concurrently (e.g. live translation)
- Time-awareness. The model has a direct sense of elapsed time.
- Simultaneous tools calls, search, and generative UI. While speaking and listening to the user, the model can concurrently search, browse the web, or generate UI—weaving back results into the conversation as needed.
In a longer real session, all of this happens continuously, creating an experience that feels more like collaborating and less like prompting.
None of the brands or products appearing in these videos are associated with Thinking Machines Labs. These videos are to demonstrate the model's capabilities and do not indicate sponsorship or partnership.
Our approach#
Time-aligned micro-turn based
Interaction is grounded in time with continuous input and output streams split into micro-turns
*Turn-based models see an alternating token sequence. Time-aware interaction models see a continuous stream of micro-turns, so silence, overlap, and interruption remain part of the model's context.*
An interaction model is in constant two-way exchange with the user—perceiving and responding at the same time. Some domains take such interactivity as a given—the physical world demands that robotics and autonomous vehicles operate in real time. Audio full-duplex modelsMoshi, PersonaPlex, nemotron-voicechat, Seeduplex. are another example where interaction is bidirectional and continuous.
Applying the same principle, we set out to build an interaction model native to this regime—one that perceives and responds in the same continuous loop, across audio, video, and text. The result is a system architected around two ideas: a time-aware interaction model that maintains real-time presence, and an asynchronous background model that handles sustained reasoning, tool use, and longer-horizon work.
System overview#
The interaction model is in constant exchange with the user. When a task requires deeper reasoning than can be produced instantaneously, the interaction model delegates to a background model that runs asynchronously.This approach builds upon prior work like Qwen-omni, KAME, MoshiRAG. The interaction model remains present throughout — answering follow-ups, taking new input, holding the thread — and integrates background results into the conversation as they arrive.
real-time
user
interaction**model
background
model
context
response
tool calls
browsing
etc
*The user continuously interacts with the interaction model, while the background model performs asynchronous tasks. Both systems share their context.*
This split lets the user benefit from both responsiveness as well as the full extent of intelligence: the planning, tool-use, and agentic workflows of reasoning models at the response latency of non-thinking ones. Note that both the background and interaction models are intelligent — on its own, the interaction model is also competitive on both interactive and intelligence benchmarks.
The interaction model#
Our starting point is continuous audio and video — modalities that are inherently real-time. Text can wait, but a live conversation cannot. By designing around the hardest case first, we arrive at an architecture that is natively multimodal, time-aware, and capable of handling concurrent input and output streams across all modalities. Several design choices make this possible.
Time-aligned micro-turns.** The interaction model works with micro-turns continuously interleaving the processing of 200ms worth of input and generation of 200ms worth of output. Rather than consuming a complete user-turn and generating a complete response, both input and output tokens are treated as streams. Working with 200ms chunks of these streams enables near real-time concurrency of multiple input and output modalities.
Human perception
input
0
input
1
input
2
input
3
input
4
output
0
output
1
output
2
output
3
Model token sequence
*Human perception preserves concurrent input and output streams, while the model receives a single interleaved token sequence.*
With this design, there are no artificial turn boundaries that the model must adhere to. In contrast, most existing real-time systems require a harness that predicts turn boundaries in order for the turn-based models to feel real-time and responsive.Moshi, PersonaPlex, and Nemotron Voicechat are examples of full duplex systems that do not use harnesses to detect turns. They are smaller scale models focused on latency rather than intelligence benchmarks. This harness is made out of components like voice-activity-detection (VAD) that are meaningfully less intelligent than the model itself. This precludes a variety of interaction modes like proactive interjections (“interrupt when I say something wrong”) or reactions to visual cues (“tell me when I’ve written a bug in my code”). Moreover, the model can do things like speak while listening (“translate from spanish to english live”) or watching (“live-commentate this sports game”).
Thus, all of these different interaction modes that require special harnesses today become special-cases of what the model can do and improve in quality as we scale up model size and training data.
Encoder-free early fusion. Rather than processing audio and video through large, standalone encoders, we opt for a system with minimal pre-processing. Many omnimodal models require training a separate encoder (e.g. Whisper-like) or decoder (e.g. TTS model-like). We instead take in audio signals as dMel (Bai, et al. 2024) and transform it via a light-weighted embedding layer. Images are split into 40x40 patches which are encoded by an hMLP (Touvron et al. 2022). For the audio decoder we use a flow head (Lipman at al. 2022). All components are co-trained from scratch together with the transformer.
Text
Frame
Audio
Embedding
Tokens
40x40 Patch
hMLP
dMel
Bag of
embeddings
Transformer
Text
Unembedding
Mel
Flow
*An illustration of the interaction model architecture for a single 200ms micro-turn. The model takes in any subset of text, audio, or video and predicts text and audio.*
Inference optimization. At inference time, 200ms chunks require frequent prefills and decodes of small sizes, each having to meet strict latency constraints. Unfortunately, existing LLM inference libraries are not optimized for frequent small prefills—they often have a significant amount of overhead per turn. To address this, we implemented streaming sessions. The client sends each 200ms chunk as a separate request, while the inference server appends these chunks into a persistent sequence in GPU memory. This avoids frequent memory reallocations and metadata computations, and we’ve upstreamed a version of this feature to SGLang. In addition, we also optimized our kernels for latency as well as the shapes we see for bidirectional serving. For example, we use a gather+gemv strategy for MoE kernels instead of the standard grouped gemm, as in prior work from PyTorch and Cursor.
Trainer-sampler alignment. We’ve found bitwise trainer-sampler alignment to be useful for training stability as well as debugging the various components of our system. We implement batch-invariant kernels with minimal (<5%) e2e performance overhead.Funnily enough, for some period of time using the batch-invariant kernels was actually faster e2e, due to the custom communication kernels which were not only batch-invariant but also much lower latency. To highlight two particular kernels:
- All-reduce and reduce-scatter: We use NVLS to implement low-latency comm kernels which are deterministic on Blackwell, and achieve bitwise alignment between somewhat different parallelism strategies (i.e. Sequence Parallelism and Tensor Parallelism).
- Attention: The primary challenge with attention is Split-KV, which can typically lead to inconsistent accumulation orders between decode and prefill.Work done in collaboration with Colfax However, we can maintain consistent accumulation order by choosing to split consistently between decode and prefill. For example, we could split SMs to process 4096 tokens at a time (left-aligned), achieving good efficiency in both prefill and decode.
Coordination between interaction and background models. When the interaction model delegates, it sends a rich context package — not a standalone query, but the full conversation. Results stream back as the background model produces them, and the interaction model interleaves these updates into the conversation at a moment appropriate to what the user is currently doing, rather than as an abrupt context switch.
Safety. Because real-time interaction stresses safety differently than turn-based exchanges, our safety work focused on two axes: modality-appropriate refusals and long-horizon robustness. To make refusals colloquial in speech, we use a text-to-speech model to generate refusal and over-refusal training data covering a range of disallowed topics, with the refusal boundary calibrated to favor naturally-phrased, but no less firm, refusals. To improve robustness across extended speech-to-speech conversations, we used an automated red-teaming harness to generate multi-turn refusal data, while maintaining close behavioral parity with the model’s text-based refusals.
Benchmarks#
Intelligence and interactivity frontier#
We show that our interaction model, named TML-Interaction-Small, is the first model that has both strong intelligence/instruction following and interactivity. To measure interaction quality we use FD-bench which is one of the few existing benchmarks intended to measure interactivity. In FD-bench v1.5, the model is given prerecorded audio, and must respond at certain times. This benchmark measures model behavior across several scenarios: user interruption, user backchannel, talking to others, and background speech. Our model scores well in all of these areas. To quantify intelligence we use Audio MultiChallenge, a common benchmark that tracks intelligence and instruction following.
TML-interaction-small
GPT-realtime-2.0 (minimal)
GPT-realtime-2.0 (xhigh)
GPT-realtime-1.5
Gemini-3.1-flash-live-preview (minimal)
Gemini-3.1-flash-live-preview (high)
*Intelligence and Interactivity Frontier. Our model dominates interaction quality while being more intelligent than any non thinking model. We achieve the best responsiveness measured as a latency between user and model turns.*
For more intelligence, safety, and interactivity/latency results please see the table below. We report our performance on both streaming and turn-based benchmarks.
| Instant | Thinking | |||||||
|---|---|---|---|---|---|---|---|---|
| TML-interaction / -small | GPT-realtime-2.0 / (minimal) | GPT-realtime-1.5 | Gemini-3.1-flash-live / (minimal) | Qwen 3.5 / OMNI-plus-realtime | GPT-realtime-2.0 / (xhigh) | Gemini-3.1-flash-live / (high) | ||
| Streaming | FD-bench V1 Turn-taking latency (s) · Audio | 0.40 | 1.18 | 0.59 | 0.57 | 2.14 | 1.63 | 0.94 |
| Streaming | FD-bench V1.5 Average · Audio | 77.8 | 46.8 | 48.3 | 54.3 | 39.0 | 47.8 | 45.5 |
| Streaming | FD-bench V3 Response Quality (%) / / Pass@1 (%) · Audio + Tools | 82.8* / 68.0* | 80.0 / 52.0 | 77.9 / 55.0 | 68.5 / 48.0 | 60.0 / 50.0 | 81.0 / 58.0 | 71.4 / 48.0 |
| Streaming | QIVD** Accuracy (%) · Video + Audio | 54.0 | 57.5 | 41.2 | 54.7 | 59.0 | 58.2 | 56.1 |
| Turn-based | Audio MultiChallenge APR (%) · Audio | 43.4 | 37.6 | 34.7 | 26.8 | -*** | 48.5 | 36.1 |
| Turn-based | BigBench Audio Accuracy (%) · Audio | 75.7 / 96.5* | 71.8 | 81.4 | 71.3 | 73.0 | 96.6**** | 96.6 |
| Turn-based | IFEval (VoiceBench) Accuracy (%) · Audio | 82.1 | 81.7 | 68.1 | 67.6 | 80.3 | 83.2 | 82.8 |
| Turn-based | IFEval Accuracy (%) · Text | 89.7 | 89.6 | 87.5 | 85.8 | 83.4 | 95.2 | 90.0 |
| Turn-based | Harmbench Refusal rate (%) · Text | 99.0 | 99.5 | 100.0 | 99.0 | 99.5 | 100.0 | 98.0 |
Best per row
Best among instant models
- For benchmarks that require reasoning or tool calls we report our results with background agent enabled.**
** We evaluate Qualcomm IVD in a streaming setting – is a video-audio QA benchmark. In each video clip, somebody performs an action and speaks a question. We evaluate in a streaming setting, sending the raw clip from the beginning and grading the model’s transcript. Following Qwen 3.5 Omni we use a GPT-4o-mini grader.
*** Audio MultiChallenge metrics for all the baseline models are reported by Scale AI, where Qwen 3.5 OMNI-plus-realtime is not listed.
**** Bigbench Audio metrics for all the baseline models are reported by Artificial Analysis, where GPT-realtime-2.0 thinking is on high.
New dimensions of interactivity#
The existing interactivity-oriented benchmarks above do not adequately capture the qualitative jumps in interaction capabilities we notice. To that end, we have some early work aimed at quantifying these capabilities.
Time awareness and simultaneous speech.** Turn-based models with a dialog management system do not support accurate time estimation or simultaneous speech. Examples include: “How long did it take me to run one mile?”, “Correct my mispronunciations as you hear them” or “How long did it take me to write this function?"
We created two internal benchmarks to measure these proactive audio capabilities:
- TimeSpeak: Tests whether the model can initiate speech at user-specified times while producing the correct content. For example: “I want to practice my breathing, remind me to breathe in and out every 4 seconds until I ask you to stop.”
- CueSpeak: Tests whether the model speaks at the appropriate moment with the expected semantically correct response. Dataset entries are created to ensure that the model needs to speak at the same time as the user to get a full score. For example: “Everytime I codeswitch and use another language, give me the correct word in the original language.”
For both benchmarks, each example has a single expected semantic response and timing window. We grade with an LLM judge: A response is counted as correct only if it conveys the expected meaning and is delivered at the appropriate time; failing either criterion receives no credit. We report macro-averaged accuracy across examples.
Visual proactivity. Today’s commercial real-time APIs perform turn-detection via audio-only dialogue management harnesses. They respond to spoken turns, but they cannot proactively choose to speak when the visual world changes.Though we are not aware of any commercial APIs that support speech-out visual proactivity, several academic papers have built related research prototypes. StreamBridge, Streamo, StreamingVLM, and MMDuet2 study when to output text in a streaming video input setting. Being text out, they do not study additional constraints of speech-output interaction: speech has duration, can overlap with the user, and must be coordinated with turntaking, interruptions, and backchanneling. Closest to ours is AURA, which adds an ASR/TTS demo around a VideoLLM that decides when to emit text or be silent; in contrast ours is speech-native and full-duplex. For instance, if asked “Please count how many pushups I do” such a system might respond “Sure thing!” and then remain silent – waiting for an audio-only cue that never comes.
We adapted three benchmarks to evaluate visual proactivity of our model:
- RepCount-A contains videos of repeated actions and is adapted into an online counting task. We stream the video following an audio instruction “Please count out reps for {action}.”. We extract the last number said by the model after the ground truth penultimate rep, and grade by whether it was within one rep of the ground truth. This task measures continuous visual tracking and timely counting.
- ProactiveVideoQA consists of videos with questions, whose answers become available at specific moments. We stream the question in audio and then the video.Specifically we TTS the following: “Watch the video and stay quiet until a new moment answers the question. When one happens, say a concise answer. {question}”, then stream two seconds of silence so the model acknowledges the instruction. We burn subtitles into the video (if present) and mute the input video to emphasize testing visual proactivity. We report the paper’s turn-weighted PAUC@ω=0.5 metric (scaled 0-100), averaged across turns and categories. Staying silent scores 25.0.; Higher scores require correct answers at the correct times and incorrect answers are penalized.
- Charades is a standard temporal action-localization benchmark. Each video contains an action occurring over a labeled time interval. We stream a user audio instruction: “Say ‘start’ when the person starts doing {action} then say ‘Stop’ when they stop.”; then we stream the video. The model is graded by temporal IoU between predicted and the reference intervals.
Time awareness
TimeSpeak · macro-acc
Verbal cues trigger
CueSpeak · macro-acc
Visual-based counting
RepCount-A · off-by-one
Visual cues trigger
ProactiveVideoQA · PAUC@ω=0.5
Visual cues trigger
Charades · mIoU
- No-response baseline on ProactiveVideoQA is 25.0
No existing model can meaningfully perform any of these tasks. For the sake of completeness, we report the results of GPT Realtime-2 (minimal), but all models evaluated perform similar or worse on these tasks, including thinking high models. They stay silent or give incorrect answers.
*Examples from our internal audio and video benchmark.*
Future evals. We believe that interactivity is an important area for future research and we invite the community to contribute benchmarks here. We are launching a research grant to encourage more research into the field of interaction models and human-AI collaboration, including but not limited to new frameworks for assessing interactivity quality, with details coming soon.
Limitations and future work#
Long sessions. Continuous audio and video accumulate context quickly. The streaming-session design handles short and medium interactions well, but very long sessions still require careful context management—an active area of work.
Compute and deployment. Streaming audio and video at low latency requires reliable connectivity. Without a good connection, the experience degrades significantly. We believe that this can be improved significantly in the future both by improving system reliability as well as training our model to be more robust to delayed frames.
Alignment and safety. A realtime interface opens up an exciting area of research for both alignment and safety. We are collecting feedback and reviewing research grants.
Scaling model size. The current TML-Interaction-Small is a 276B parameter MoE with 12B active. While we expect the interactivity to improve with model scale, our larger pretrained models are currently too slow to serve in this setting. We plan to release larger models later this year.
Improved background agents. Although we have primarily focused on real-time interactivity in this post, agentic intelligence is also an essential capability. In addition to pushing agentic intelligence to the frontier, we believe we have just scratched the surface in how the background agents can work together with the interaction model.
Tell us what you think, join us#
In the coming months, we will open a limited research preview to collect feedback, with a wider release later this year.
We’d love for you to join us. Please share your thoughts at interaction@thinkingmachines.ai.
Citation#
Please cite this work as:
Thinking Machines Lab, "Interaction Models: A Scalable Approach to Human-AI Collaboration",
Thinking Machines Lab: Connectionism, May 2026.
Or use the BibTeX citation:
@article{thinkingmachines2026interactionmodels,
author = {Thinking Machines Lab},
title = {Interaction Models: A Scalable Approach to Human-AI Collaboration},
journal = {Thinking Machines Lab: Connectionism},
year = {2026},
month = {May},
note = {https://thinkingmachines.ai/blog/interaction-models/},
doi = {10.64434/tml.20260511},
}
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み