Inworld TTS-1.5 Maxがfalプラットフォームで利用可能に
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
fal.ai Blog
Inworldは、低遅延・高表現力・多言語対応の音声生成モデル「TTS-1.5 Max」をfalプラットフォームに追加した。これにより、開発者はアシスタントやメディア体験など、本番環境でのリアルタイム音声インターフェース構築を強化できる。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
image fal に Inworld TTS-1.5 Max を追加できることを嬉しく思います。これにより、プラットフォーム上の最先端のリアルタイム音声モデルのセットが拡大します。このモデルは、低遅延の音声生成、表現力の向上、および本番環境での利用に向けた多言語サポートに焦点を当てています。
アシスタントからメディア体験まで、アプリケーション全体で音声インターフェースの中核となりつつある現在、開発者は遅延、品質、コストのバランスが取れたモデルを必要としています。TTS-1.5 Max はこれらの制約内で動作するように設計されており、リアルタイムインタラクションをサポートします。
Inworld TTS-1.5 Max とは?
Inworld TTS-1.5 Max は、表現力に富み低遅延の音声合成を実現するためのテキスト読み上げ(TTS: Text-to-Speech)モデルです。これは、高品質版である Max 型と低遅延版である Mini 型の両方を含む TTS-1.5 ファミリーの一部です。
Max モデルは、応答性をほぼリアルタイムに保ちつつ、音声の質と表現範囲を優先し、ほとんどのアプリケーションにおけるデフォルトオプションとして位置づけられています。
主な特徴
リアルタイム遅延
TTS-1.5 Max は、最初の音声までの時間を約 250ms(P90)未満で達成し、応答時間がユーザーエクスペリエンスに影響を与える会話型およびインタラクティブなユースケースを可能にします。
表現力と精度の向上
以前のバージョンと比較して、このモデルはより広い表現範囲と低い単語誤り率を導入しました。これにより、発音ミス、途切れ、不自然なペース配分などのアーティファクトが減少します。
多言語サポート
本モデルは、ローカライゼーションや翻訳などのグローバルアプリケーションおよびユースケースに対応するため、15 か国語をサポートしています。
コストプロファイル
料金は 1 分あたり約 0.01 ドル(100 万文字あたり 10 ドル)で構成されており、多くの同等のリアルタイム TTS(Text-to-Speech:音声合成)システムと比較して低コストオプションとして位置づけられています。
fal でお試しください
fal を介して Inworld TTS-1.5 Max の利用を開始し、表現豊かな音声を生成したり、レイテンシとパフォーマンスのトレードオフをテストしたり、アプリケーションにボイス機能を統合したりできます。
生成メディアや新モデルリリースに関する最新情報は、X(旧 Twitter)、ブログ、または Reddit で随時お知らせいたします!
原文を表示
imageWe’re excited to add Inworld TTS-1.5 Max to fal, expanding our set of cutting-edge real-time voice models on the platform. The model focuses on low-latency speech generation, improved expressiveness, and multilingual support for production use cases.
As voice becomes a core interface across applications, from assistants to media experiences, developers need models that balance latency, quality, and cost. TTS-1.5 Max is designed to operate within these constraints while supporting real-time interactions.
What is Inworld TTS-1.5 Max?
Inworld TTS-1.5 Max is a text-to-speech model built for expressive, low-latency voice synthesis. It is part of the TTS-1.5 family, which includes both Max (higher quality) and Mini (lower latency) variants.
The Max model is positioned as the default option for most applications, prioritizing voice quality and expressive range while maintaining near-realtime responsiveness.
Key characteristics
Realtime latency
TTS-1.5 Max achieves time-to-first-audio under ~250ms (P90), enabling conversational and interactive use cases where response time impacts user experience.
Improved expressiveness and accuracy
Compared to earlier versions, the model introduces higher expressive range and lower word error rates. This reduces artifacts such as mispronunciations, cutoffs, and unnatural pacing.
Multilingual support
The model supports 15 languages, including expanded coverage for global applications and use cases like localization and translation.
Cost profile
Pricing is structured at approximately $0.01 per minute ($10 per million characters), positioning it as a lower-cost option relative to many comparable realtime TTS systems.
Try it on fal
You can start using Inworld TTS-1.5 Max on fal to generate expressive speech, test latency-performance tradeoffs, and integrate voice into your applications.
Stay tuned to our X, blog or Reddit for the latest updates on generative media and new model releases!
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み