ボクストラの発表
Mistral AI は、9 か国語の多様な方言に対応し、低遅延かつ高コストパフォーマンスを実現した軽量(4B パラメータ)テキスト音声合成モデル「Voxtral TTS」をリリースした。
キーポイント
高性能な軽量モデルの実装
40 億パラメータというコンパクトなサイズでありながら、多言語・多方言対応とリアルタイム性を両立し、大規模展開におけるコスト効率を劇的に向上させた。
文脈理解に基づく感情表現
単なる文字の読み上げではなく、文脈(中立、喜び、皮肉など)や話者の個性(間隔、リズム、イントネーション)を深く理解し、自然で感情的な発話を生成する。
エンタープライズ向けのカスタマイズ性
新しい声を容易に適応可能であり、企業が自社の音声 AI スタックを完全に支配・カスタマイズできる機能を提供し、信頼性の高いボイスエージェント基盤となる。
自然な音声とカスタマイズ性
ネイティブスピーカーによる評価で、ElevenLabs Flash v2.5 よりも優れた自然性を達成し、3 秒の参考音声からアクセントやイントネーションを含む声を即座に再現できます。
低遅延ストリーミング対応
10 秒・500 文字の入力に対して 70ms のモデル遅延を達成し、リアルタイムファクター(RTF)は約 9.7x で、長文生成もスマートなインターリーブ処理でサポートします。
ゼロショット言語間適応
明示的な訓練がなくても、フランス語の音声プロンプトを使って英語を話すなど、異なる言語間で自然なアクセントを持つ発話を生成するクロスリンガル能力を備えています。
既存システムへの統合とテスト
Voxtral TTS は既存の顧客サポート通話システムに組み込んで自動化された音声応答を実現でき、Mistral Studio のプレイグラウンドで直接モデルを試し、独自の声を録音することも可能です。
重要な引用
A natural voice generation hinges on the model's ability to not only recite but interpret a text accurately.
Audio is the new UX. Create new interactions for collaboration and understanding only found in speech.
What makes speech natural is extremely nuanced and requires a deep understanding of cultural differences and typical speaking patterns.
Voxtral TTS achieves superior naturalness compared to ElevenLabs Flash v2.5 while maintaining similar Time-to-First-Audio (TTFA).
Voxtral TTS is available now via API at $0.016 per 1k characters.
We are building the voice layer for AI, and If this is the kind of problem you want to work on, we'd love to hear from you.
影響分析・編集コメントを表示
影響分析
この発表は、AI エージェントが単なる情報伝達から、人間のような感情と文脈理解を持った対話へ進化するための重要な一歩を示しています。特に軽量モデルで高品質な多言語対応を実現した点は、グローバル展開する企業におけるコスト削減とユーザー体験の向上に直結し、音声 AI の実用化スピードを加速させるでしょう。
編集コメント
4B パラメータという軽量設計で SOTA な性能を達成した点は、エッジデバイスや大規模展開におけるコスト最適化の観点から非常に注目すべき成果です。特に「皮肉」や「感情」のニュアンスを理解する能力は、次世代ボイスエージェントの実現に不可欠な要素と言えます。
本日、私たちは最先端の多言語音声生成性能を備えた最初のテキスト読み上げモデル「Voxtral TTS」を発表します。このモデルは 4B パラメータという軽量設計であり、Voxtral を搭載したエージェントは、大規模展開においても自然で信頼性が高く、コスト効率に優れたものとなります。
ハイライト。
- 9 つの主要言語でリアルかつ感情豊かな発話を実現し、多様な方言にも対応しています。
- 最初の音声出力までの遅延時間が極めて短い。
- 新しい声への適応が容易です。
- Mistral Studio でテスト利用可能です。
- エンタープライズグレードのテキスト読み上げ技術により、重要な音声エージェントワークフローを駆動します。
自然な音声生成には、モデルがテキストを単に朗読するだけでなく、正確に解釈する能力が不可欠です。文脈理解(例えば、無表情、喜び、皮肉など)が、聴衆にとってその生成物が自然なものなのか機械的なものなのかを決定づけます。私たちのモデルは、文脈理解と話者モデリングの両面で卓越しており、特定の人物がどのように自然に話すかを捉えます。当社の音声適応技術は、従来の朗読型アプローチを超え、話者の個性、すなわち自然な間、リズム、イントネーション、そして感情表現の巧みさまでを捉えるものです。コンパクトなサイズ、低コスト、低遅延、そして容易な適応性を備えた Voxtral TTS は、自社の音声 AI スタックを完全に制御・カスタマイズしたい企業にとって、そのための完全なコントロールと柔軟性をもたらします。
オーディオは新たな UX です。会話や理解において音声にしか存在しない新しい相互作用を生み出しましょう。米語、英式英語、フランス語の方言に対応する「Mistral Voices」で AI Studio での利用を今すぐ開始してください。
聞いて判断してください:違いを識別できますか?
私たちのチームは複数の方言を含む数十の言語を話し、文化的なニュアンスの重要性を理解しています。私たちは自分たちを反映したモデルを構築しました。音声生成は、自然なリズム、感情、さらにはユーモアの活用を通じて信頼性を築きます。そのため、音声エミュレーションにおいては、真正性と感情的な表現力に焦点を当てました。
最先端のパフォーマンス
単語誤り率や多言語テキスト読み上げシステムの音声品質スコアといった自動評価指標では、音声の自然さを測定することはできません。何が音声を自然にするのかは極めて微妙で、文化的差異や典型的な発話パターンに対する深い理解を必要とします。そのため、ネイティブスピーカーによる比較的人間評価が不可欠です。
音声エージェントにおいては、レイテンシ(応答遅延)と品質は常に緊張関係にあります。人間評価によると、Voxtral TTS は ElevenLabs Flash v2.5 と比較して同等の Time-to-First-Audio (TTFA) を維持しながら、 superior な自然さを達成しています。また、Voxtral は ElevenLabs v3 の品質と同等のパフォーマンスを発揮し、より生々しい対話を実現するための感情制御(emotion-steering)を成功裏にサポートしています。

Voxtral TTS と ElevenLabs v2.5 Flash のゼロショットカスタム音声コンテキストにおける比較人的評価を実施しました。9 つのサポート言語それぞれについて、それぞれの母語方言で認識可能な 2 つの声を用い、3 人の注釈者が自然さ、アクセントへの準拠度、および元の参照音源との音響的類似性の観点からペアごとの並列選好テストを行いました。このゼロショット多言語カスタム音声設定において、Voxtral TTS は v2.5 Flash との品質格差を拡大し、あらゆる声に対する Voxtral TTS の即時カスタマイズ性を浮き彫りにしています。
母語で話されるように。
大規模な音声データセットでトレーニングされた Voxtral TTS は、グローバルアプリケーション向けに構築されています。英語、フランス語、ドイツ語、スペイン語、オランダ語、ポルトガル語、イタリア語、ヒンディー語、アラビア語の 9 か国語において最先端のパフォーマンスをサポートします。
このモデルは、3 秒というわずかな参照音声でカスタム音声に適応し、単に声質だけでなく、参照音源で表現されるような微妙なアクセント、イントネーション、抑揚、さらには不自然さ(disfluencies)といったニュアンスも捉えるようにトレーニングされています。API ではいくつかのプリセット音声オプションを提供していますが、社内音声ライブラリへの拡張は簡単です。使用ケースに合わせてカスタマイズし、言語やアクセントにローカライズし、中立的またはより感情的な表現、カジュアルまたはフォーマルなトーン、自然で会話的なものからロボットのようなものまで、自由に調整可能です。
このモデルは、明示的にそのために訓練されていないにもかかわらず、ゼロショットの異言語音声適応も示します。例えば、このモデルはフランス語の音声プロンプトと英語テキストを用いて英語の発話を生成できます。結果として得られる発話は自然に聞こえつつ、提供された音声プロンプトのアクセント(本例では、自然なフランス語訛りの英語)を採用しています。これにより、カスケード型の音声対音声翻訳システムの構築においてこのモデルは有用となります。
カスケード型音声対音声翻訳
スピーカーをプロンプトブロックにクリックして接続するか、接続することで、音声からテキストへの翻訳機能を有効化できます。
低遅延ストリーミング向けに設計
遅延は音声エージェントアプリケーションにおいて極めて重要です。Voxtral TTS は、10 秒・500 文字の典型的な入力音声サンプルに対してモデル遅延 70ms を達成し、リアルタイムファクター (RTF: Real-Time Factor) は約 9.7 倍です。このモデルはネイティブで最大 2 分間のオーディオを生成可能であり、当社の API はスマートなインターリーブ処理により任意の長さの生成に対応しています。
Voxtral TTS アーキテクチャ
本モデルは、Ministral 3B を基盤とした、トランスフォーマーベースの自己回帰型フローマッチングモデルです。以下のコンポーネントで構成されています:
- 3.4B パラメータを持つトランスフォーマーデコーダーバックボーン
- 390M のフローマッチング音響トランスフォーマー
- 300M のニューラルオーディオコーデック(対称型エンコーダー・デコーダ)
このモデルは、5〜25 秒の音声プロンプトと、9 言語に対応したテキストプロンプトを受け取ります。各オーディオフレームに対して、トランスフォーマーバックボーンが意味トークンを予測し、その後フローマッチングトランスフォーマーが 16 回の関数評価(NFEs)を実行して音響潜在変数を生成します。
私たちは社内開発のコーデックを開発しました。これは、意味 VQ(語彙数 8192)と音響 FSQ(36 次元、21 レベル)の潜在変数を用いてオーディオを因果的に処理し、12.5Hz のフレームレートで出力します。

エンタープライズ向け音声ワークフローを推進する。
**
**Voxtral TTS はオーディオインテリジェンスのループを完結させ、エンタープライズの音声パイプラインに人間によるテストに合格する出力層を提供します。これは Voxtral Transcribe と併用して完全な音声対音声を実現するか、またはクロスリンガルサポートを備えた既存の音声テキスト変換および LLM スタックへ統合されます。
ワークフロー
カスタマーサポート
自然でブランドに合った音声を用いて、複数のチャネル間で問い合わせをルーティングし解決する音声エージェント。
既存のカスタマーサポート通話システムに Voxtral TTS を組み込み、自動化された音声応答を実現します。出力は既存のワークフローへ統合可能です。
Mistral Studio でモデルを試す
Voxtral TTS を Mistral Studio プレイグラウンド で直接実験できます。Mistral の音声から一つを選択するか、ご自身の音声を録音してください。
Voxtral TTS の利用開始
Voxtral TTS は現在、API を通じて利用可能で、料金は 1,000 文字あたり 0.016 ドルです。
Mistral Studio または Le Chat で今すぐお試しください。
複数のリファレンス音声を持つモデルは、CC BY NC 4.0 ライセンスの下でオープンウェイトとしてHugging Face で利用可能です。
モデルのドキュメントをご覧ください。または、当社の研究論文をお読みください。
詳細については、今後のウェビナーにご登録ください!
採用情報
私たちは AI の音声層を構築しています。もしあなたがこのような課題に取り組みたいとお考えであれば、ご連絡をお待ちしています。
原文を表示
Today we’re releasing Voxtral TTS, our first text-to-speech model with state-of-the-art performance in multilingual voice generation. The model is lightweight at 4B parameters, making Voxtral-powered agents natural, reliable, and cost-effective at scale.
Highlights.
- Realistic, emotionally expressive speech in 9 popular languages with support for diverse dialects.
- Very low latency for time-to-first-audio.
- Easily adaptable to new voices.
- Available to test out in Mistral Studio.
- Enterprise-grade text-to-speech, powering critical voice agent workflows.
A natural voice generation hinges on the model’s ability to not only recite but interpret a text accurately. Contextual understanding - like neutral, happy, sarcastic, etc. - determines whether the listener considers the generation accurate or robotic. Our model excels at both contextual understanding and speaker modeling: capturing how a specific person naturally speaks. Our voice adaptation goes beyond traditional read-speech by capturing a speaker’s personality, including their natural pauses, rhythm, intonation, and emotional dexterity. With its compact size, low cost and latency, and easy adaptability, Voxtral TTS gives full control and customization for enterprises looking to own their voice AI stack.
Audio is the new UX. Create new interactions for collaboration and understanding only found in speech. Begin now in AI Studio with our Mistral Voices in American, British, and French dialects.
Listen and decide: can you tell the difference?
Our team speaks dozens of languages in multiple dialects, we understand the importance of cultural nuance and built a model that is a reflection of us. Speech generation builds trust via natural-like rhythm, emotion, and even the use of humor. That’s why with voice emulation, we focused on authenticity and emotional expressiveness.
State-of-the-art performance.
Automated metrics such as word-error-rate and audio quality scores for multilingual text-to-speech systems are unable to measure naturalness of speech. What makes speech natural is extremely nuanced and requires a deep understanding of cultural differences and typical speaking patterns. Hence, comparative human evaluations performed by native speakers are crucial.
For voice agents, latency and quality are in constant tension. Human evaluations show that Voxtral TTS achieves superior naturalness compared to ElevenLabs Flash v2.5 while maintaining similar Time-to-First-Audio (TTFA). Voxtral also performs at parity with the quality of ElevenLabs v3, successfully supporting emotion-steering for more lifelike interactions.

We conducted a comparative human evaluation of Voxtral TTS and ElevenLabs v2.5 Flash in a zero-shot custom voice context. Using two recognizable voices in their native dialects for each of the 9 supported languages, 3 annotators performed a side-by-side preference test per pair on naturalness, accent adherence, and acoustic similarity to the original reference. Voxtral TTS widens the quality gap to v2.5 Flash in this zero-shot multilingual custom voice setting, highlighting the instant customizability of Voxtral TTS to any voice.
Spoken natively.
Trained on a large speech dataset, Voxtral TTS is built for global application. It supports state-of-the-art performance in 9 languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic.
The model was trained to adapt to a custom voice with a reference as little as 3s and capture not just the voice but also nuances like subtle accent, inflections, intonations and even disfluencies similar to those expressed in the reference. We offer some preset voice options in the API but it is simple to extend to your in-house voice library customizing it to the use-case, localize it to the language and accent, keep it neutral or more emotive, casual or formal, more natural and conversational or robotic.
The model also demonstrates zero-shot cross-lingual voice adaptation even though it’s not explicitly trained for it. For example, the model can generate English speech with a French voice prompt and English text. The resulting speech sounds natural while adopting the accent of the provided voice prompt (in this example, the generated speech has a natural French-accented English). This makes the model useful for building cascaded speech-to-speech translation systems.
Cascaded speech-to-speech translation
Click or connect a speaker to the prompt block to enable cascaded speech-to-text translation.
Built for low-latency streaming.
Latency is critical for voice agent applications. Voxtral TTS achieves a model latency of 70ms for a typical input voice sample of 10 seconds and 500 characters, with a real-time factor (RTF) of ≈9.7x. The model natively generates up to two minutes of audio, and our API handles arbitrarily long generations with smart interleaving.
Voxtral TTS architecture.
The model is a transformer-based, autoregressive, flow-matching model, built on Ministral 3B. It consists of the following components:
- 3.4B parameters transformer decoder backbone
- 390M flow-matching acoustic transformer
- 300M neural audio codec (symmetric encoder-decoder)
The model takes a voice prompt (5 to 25 seconds) and a text prompt in 9 supported languages. For each audio frame, the transformer backbone predicts a semantic token, then the flow-matching transformer runs 16 function evaluations (NFEs) to produce the acoustic latent.
We developed an in-house codec, which processes audio causally using a semantic VQ (8192 vocabulary) and an acoustic FSQ (36 dim and 21 levels) latent and produces them at 12.5Hz frame rate.

Powering enterprise voice workflows.
**
**Voxtral TTS closes the loop on audio intelligence, giving enterprise voice pipelines an output layer that passes the human test. It works alongside Voxtral Transcribe for full speech-to-speech, or integrates into any existing speech-to-text and LLM stack, with cross-lingual support.
Workflows
Customer Support
Voice agents that route and resolve queries across channels with natural, brand-appropriate speech.
Place Voxtral TTS into existing contact support call systems for automated spoken responses, with output that integrates into existing workflows.
Test-run the model in Mistral Studio.
Experiment with Voxtral TTS directly in the Mistral Studio playground. Select one of the Mistral voices or record your own.
Get started with Voxtral TTS.
Voxtral TTS is available now via API at $0.016 per 1k characters.
Try it now in Mistral Studio or in Le Chat.
A model with several reference voices is available as open weights on Hugging Face under CC BY NC 4.0 license.
Explore the model’s documentation or read our research paper.
Sign up for our upcoming webinar to learn more!
We’re hiring!
We are building the voice layer for AI, and If this is the kind of problem you want to work on, we'd love to hear from you.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み