OpenAI、ChatGPT向け双方向音声モードの展開を準備
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
OpenAIは、アシスタントが同時に話しかけ、聞き取り、応答できる新音声生成モデル「Bidi 1」をChatGPTに導入し、会話の流れを維持しながら中断時に即座にタスクを切り替える機能をロールアウトしている。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
OpenAI は、ChatGPT の音声モードに数ヶ月ぶりの最大規模のアップグレードを施す準備を整えているようです。次世代オーディオモデル「Bidi 1」Bidi 1 が登場しており、これは双方向型(bidirectional)設計を指す略称で、アシスタントが同時に話しかけ、聞き取り、聴取できる機能を実現しています。今週中のリリースの可能性に先駆けて、ChatGPT の Web インターフェースでの言及が始まり、すでにアプリ内の一部のユーザーにも配信され始めています。
機械学習 & 人工知能
初期テストでは、現在の先進的な音声モードとの差は明白です。Bidi 1 は設定のモデル選択画面にあり、「標準」や「高度」オプションの隣に配置されており、選択すると音声バブルが黄色くなります。話者が一時停止したり速度を落としたりした際、割り込まずに「はい」や短い頷きといった自然な応答を示します。また、タスクもその場で切り替えることができます。「10 まで数えて」と指示し、途中で中断して逆数を数えさせると、即座に対応します。
より実用的なのは、会話の文脈を維持できる点です。以前のコンテキストを失うことなく、一連の会話を継続的に保持します。これは長年現在の音声スタックの弱点として指摘されてきた部分であり、また、長い一時停止中に割り込んでくることもなくなりました。
image創造的な振る舞いは、最初の高度な音声ロールアウトから引き継がれており、歌唱やビートボックスも含まれます。ただし、著作権の扱いについてはより厳格化されており、人気のある楽曲は明確に拒否する一方で、指定されたアーティストのスタイルでオリジナル曲を創作しようとする試みは継続されます。
この動きは、OpenAI が、その能力の高いテキストモデルと古くからの音声レイヤーとの距離を縮め、会話を ChatGPT への主要な入口として位置づけていることを示唆しています。同社はこれを正式に発表していません。ウェブおよびモバイルプラットフォームを通じて段階的かつオプトイン形式でリリースされる可能性が高く、欧州経済領域(EEA)ではより長い待機期間が設けられる可能性があります(未確認)。Codex は、このローンチの数週間後に独自の音声アップグレードを予定しており、これは本件とは別に実施されます。API アクセスはさらに後になる見込みですが、具体的なタイムラインは確定していません。
機械学習 & 人工知能
原文を表示
OpenAI looks set to hand ChatGPT's voice mode its biggest upgrade in months, with a next-generation audio model surfacing as Bidi 1, shorthand for the bidirectional design that lets the assistant speak, hear, and listen at once. References to it began appearing in the ChatGPT web interface ahead of a possible release this week, and it has already begun reaching a subset of users in the app.
MachineLearning & Artificial Intelligence
In our early testing, the gap from today's advanced voice mode is plain. Bidi 1 sits in the model selector under settings, beside the standard and advanced options, and turns the voice bubble yellow once picked. It offers small, natural acknowledgments — an "okay" or a brief nod — when you pause or slow down, without cutting across you. It also switches tasks on the fly: ask it to count to ten, interrupt to reverse the count, and it adjusts immediately.
More usefully, it holds the thread of a whole conversation rather than dropping earlier context, the weak point that has long dogged the current voice stack, and it no longer jumps in during longer pauses.

Creative behavior carries over from the first advanced voice rollout, singing and beatboxing included, though copyright handling is tighter; it declines popular songs outright while still attempting an original piece in a chosen artist's style.
The move reads as OpenAI closing the distance between its capable text models and an older voice layer, treating conversation as a core route into ChatGPT. The company has not formally announced it. A gradual, opt-in release across web and mobile looks likely, with the European Economic Area possibly waiting longer (not confirmed). Codex appears set for its own voice upgrade in the weeks after this launch, separate from it, and API access may follow later still (timeline is not confirmed).
MachineLearning & Artificial Intelligence
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み