SpaceXAI、Agent Builder で Grok Voice Think Fast 2.0 を発表
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
xAI は音声対話モデル「Grok Voice Think Fast 2.0」を発表し、ベンチマークで競合を凌駕する性能と低遅延を実現して開発者の活用を促進した。
AI深層分析を開く2026年7月30日 23:12
AI深層分析
キーポイント
ベンチマークでの高性能達成
Artificial Analysis の評価において総合スコア82.9%、エージェント機能56.5%を獲得し、GPT-Realtime-2.1 や Gemini 3.1 Flash を上回る結果となった。
音声認識精度の劇的向上
24言語での評価により、背景ノイズや電話回線の圧縮下で競合比10倍の精度差を記録し、実環境での信頼性を大幅に高めた。
並列推論による低遅延化
音声生成と推論を並行処理する設計により、初回音声までの時間を0.70秒まで短縮し、複雑なクエリへの対応も可能にした。
実社会での効果検証済み
Starlink の電話サービスにおけるA/Bテストで販売転換率とサポート完結率が向上し、実運用での即効性を示した。
旧バージョンへのアクセス方法
以前のモデルを利用したい開発者は、切り替え前にバージョン識別子 1.0 を固定する必要がある。それ以外のユーザーは特にアクションを行う必要はない。
重要な引用
On Artificial Analysis' speech-to-speech benchmark, Think Fast 2.0 scored 82.9% overall, up from 75.7% for version 1.0 and ahead of GPT-Realtime-2.1 at 79.1%
The comparison uses the word error rate, with lower scores preferred.
Think Fast 2.0 reasons in parallel with speech, a design intended to preserve latency while handling more complex queries.
Developers who want the prior model must pin the 1.0 identifier before the switch; everyone else needs to take no action.
編集コメントを表示
編集コメント
xAI は今回のリリースで、単なる性能向上だけでなく実環境でのノイズ耐性と推論速度の両立を成し遂げた。競合他社との明確な差を示す数値は、開発者が音声エージェントを採用する際の判断材料として極めて有用である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
xAI は、知能・文字起こしの精度・会話の振る舞い・ツール利用能力を向上させた次世代音声対話モデル「Grok Voice Think Fast 2.0」を発表しました。このモデルは音声エージェントを開発する開発者向けに設計されており、1 分あたりの料金は 0.08 ドルです。xAI は、既存のプロンプトを変更しなくても、ほぼすべてのユースケースで性能が向上すると期待しています。
Artificial Analysis が実施した音声対話ベンチマークでは、Think Fast 2.0 の総合スコアは 82.9% に達しました。これはバージョン 1.0 の 75.7% を上回り、GPT-Realtime-2.1(79.1%)や Gemini 3.1 Flash(69.5%)を凌駕する結果です。エージェント機能に関するスコアも 56.5% に向上し、前作の 52.1% を上回りました。同様に GPT-Realtime-2.1 の 45.7% や Gemini 3.1 Flash の 37.7% よりも高い数値です。最初の音声出力までの時間(Time to first audio)は、1.25 秒から 0.70 秒に短縮されました。会話能力に関するベンチマークスコアは 95.1% で、GPT-Realtime-2.1 の 95.7% にわずかに及ばないものの、非常に高い水準を維持しています。
知能・文字起こしの精度・会話能力が向上した次世代音声モデル「Grok Voice Think Fast 2.0」を発表しました。https://t.co/XUiX1CouKz pic.twitter.com/Nel3zwzkwY
— SpaceXAI (@SpaceXAI) 2026 年 7 月 29 日
音声認識もまた、重要な焦点の一つです。xAI は 24 カ国語にわたる数千の短いフレーズを評価した結果、Deepgram Nova 3 や ElevenLabs Scribe v2 と比較して精度が 1.5〜2.0 倍向上し、Think Fast 1.0 と比較しても 1.4 倍の改善が見られたと報告しています。特に、大きな背景ノイズや電話回線の圧縮下では、その差は約 10 倍に達すると同社は説明しています。評価には単語誤り率(WER)が用いられ、数値が低いほど優れています。
Think Fast 2.0 は、発話と並行して推論を行う設計となっており、複雑なクエリを処理しつつも遅延を抑えることを目指しています。その結果、推論トークンの使用量は、先行モデルを基準(1.0)とした場合、中位数で 0.4 倍にまで低下しました。これにより、エージェントが最初の文を完了する前に、生産環境でのツール呼び出しが通常実行可能になります。また、強化学習によって、複雑なワークフローを案内する際にも、より短い文章で一度に一つの質問に応じ、余計な言葉を省く方向へモデルが改善されました。
SpaceXAI にとって今回のリリースは、Grok Voice を実際の顧客ワークフローでもっと信頼性の高いものにするための一歩です。同社によると、Starlink の電話サービスで行われた A/B テストでは、売上変換率とサポートの完結率が向上しました。2026 年 8 月 5 日、grok-voice-latest というエイリアスが自動的に grok-voice-think-fast-1.0 から grok-voice-think-fast-2.0 に切り替わります。以前のモデルを利用したい開発者は、切り替え前に 1.0 の識別子を固定する必要がありますが、それ以外のユーザーは特に操作を行う必要はありません。
X 社(旧 Twitter)傘下の AI 企業 x.ai は、エージェント構築プラットフォーム「Agent Builder」において、音声機能を持つ大規模言語モデル「Grok Voice Think Fast 2.0」の公開を開始しました。
原文を表示
xAI has introduced Grok Voice Think Fast 2.0, its next-generation speech-to-speech model, with gains in intelligence, transcription accuracy, conversational behavior, and tool use. The model is aimed at developers building voice agents and costs $0.08 per minute of audio. xAI expects it to raise performance across almost all use cases without changes to existing prompts.
On Artificial Analysis’ speech-to-speech benchmark, Think Fast 2.0 scored 82.9% overall, up from 75.7% for version 1.0 and ahead of GPT-Realtime-2.1 at 79.1% and Gemini 3.1 Flash at 69.5%. Its agentic score reached 56.5%, compared with 52.1% for its predecessor, 45.7% for GPT-Realtime-2.1, and 37.7% for Gemini 3.1 Flash. Time to first audio fell from 1.25 seconds to 0.70 seconds. The model’s 95.1% conversational benchmark score was just below GPT-Realtime-2.1 at 95.7%.
Transcription is another major focus. In xAI’s evaluation of thousands of short phrases across 24 languages, the company reported accuracy improvements of 1.5 to 2.0 times versus Deepgram Nova 3 and ElevenLabs Scribe v2, and 1.4 times versus Think Fast 1.0. xAI says the gap is roughly 10× under substantial background noise and telephony compression. The comparison uses the word error rate, with lower scores preferred.
Think Fast 2.0 reasons in parallel with speech, a design intended to preserve latency while handling more complex queries. Median relative reasoning-token use fell to 0.4 times, using the predecessor’s 1.0 times as a baseline. xAI says this lets production tool calls usually execute before the agent finishes its first sentence. Reinforcement learning also pushed the model toward shorter sentences, one question at a time, and less fluff while guiding users through complex workflows.
For SpaceXAI, this release is a push to make Grok Voice more dependable in real customer workflows. An A/B test on Starlink’s phone service produced higher sales conversion and support containment rates, according to the company. On August 5, 2026, the grok-voice-latest alias will automatically move from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0. Developers who want the prior model must pin the 1.0 identifier before the switch; everyone else needs to take no action.
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み