Microsoft AI、MAI-Transcribe-1.5 を発表:人工分析で WER2.4%、FLEURS 精度は業界最高水準、長音響変換速度は最大 5 倍向上
マイクロソフト AI は自社開発音声認識モデル「MAI-Transcribe-1.5」を発表し、43 言語・雑音環境に対応し、人工分析で WER2.4%、FLEURS 精度は業界最高水準、長音響変換速度を最大 5 倍向上させた。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
先週、Microsoft AI は MAI-Transcribe-1.5 を発表しました。これは同社が開発した音声認識(Speech-to-Text)ファミリーの2回目のバージョンです。このモデルは、43 の言語、アクセント、およびノイズの多い環境における精度を目標としています。Microsoft チームは、これを本番環境での文字起こしワークロード向けに位置付けています。
MAI-Transcribe-1.5 とは何か
MAI-Transcribe-1.5 は自動音声認識(Automatic Speech Recognition: ASR)モデルです。入力はオーディオで、出力はテキストとなります。Microsoft によって自社開発され、サードパーティの基盤上には構築されていません。このモデルは単一のシステムで43 の言語を処理します。多様なアクセント、方言、および現実世界の音響条件に最適化されています。
Microsoft はこれを Copilot、Teams、GitHub、Dynamics 365 Contact Centre に統合しています。また、同社のモデルプラットフォームである Foundry でも利用可能です。
精度に関するケース
ここでいう精度は、単語誤り率(Word-Error-Rate: WER)によって測定されます。WER が低いほど、文字起こしされた単語あたりのミスが少ないことを意味します。Microsoft は FLEURS において43 の言語で最高クラスの WER を報告しています。FLEURS は標準的な多言語文字起こしベンチマークです。
Artificial Analysis リーダーボードでは、このモデルは2.4%の WER を記録しました。これは競争の激しいオープンなベンチマークにおいて3位に位置することを意味します。つまり、状況は分かれており、Microsoft チームは FLEURS では1位、Artificial Analysis では3位であると主張しています。
言語拡張は、もう一つの精度向上の物語です。対応言語数は 25 から 43 に拡大しました。追加された 18 の新言語において、精度を損なうことなくサポートされています。そのうち 10 は南アジア諸国のもので、ベンガル語、タミル語、テルグ語が含まれます。残りの 8 つはヨーロッパ諸国の言語で、ウクライナ語、ギリシャ語、カタルーニャ語などが該当します。
速度
MAI-Transcribe-1.5 は、Artificial Analysis のリーダーボードにおいて、精度と速度の両面で首位を維持しています。同等の精度を持つ他のモデルと比較して、最大 5 倍高速で動作します。この効果は特に長時間の音声ファイルにおいて顕著です。本モデルであれば、1 時間の音声を 15 秒未満で書き起こすことが可能です。
Microsoft は、長時間音声における処理速度が Gemini 3.1、Scribe v2、GPT-4o-Transcribe を最大 5 倍上回るとしています。先行する MAI-Transcribe-1 と比較した場合、Azure の仕様書には長文推論において最大 5.7 倍の高速化が記載されています。大量のアーカイブを処理するバッチパイプラインにおいては、このレイテンシ(遅延時間)の差はすぐに蓄積・拡大します。
キーワード(エンティティ)バイアス:理解すべき機能
汎用的な書き起こしツールは、ドメイン固有の単語においてしばしば失敗します。これには人名、製品名、医療用語、社内略語などが含まれます。これらの単語は、企業ユーザーにとって最も重要な要素であることが多いです。
MAI-Transcribe-1.5 は、キーワードバイアス(エンティティバイアスとも呼ばれます)機能を追加しました。ユーザーはドメイン固有のキーワードリストを指定します。Azure の仕様書では最大 200 個までのキーワードをサポートしています。本モデルはこのリストに対して予測を偏向させます。ただし、重要な点として、盲目的に一致させるわけではありません。共有された文脈(コンテキスト)を用いて、バイアスを適用すべきタイミングを判断します。Microsoft によると、バイアス機能を使用した場合、FLEURS における WER(単語誤り率)が 30% 削減されたと報告されています。
短い例でその効果が示されます。バイアスを与えない場合、名前は「Sean」「Oif」「Societal」として表示されますが、提供された名前リストを使用すると、「Shaun」「Aoife」「Xochitl」を正しく復元できます。これは、専門用語が多い会議、医療現場、コールセンターにおいて特に重要です。
ユースケース
Azure モデルカードには、具体的な生産環境でのシナリオが記載されています。それぞれが一般的なエンジニアリングの負荷に対応しています:
メディアやコンテンツプラットフォーム向けの動画字幕。
正確な字幕に依存するアクセシビリティツール。
Teams 型コラボレーションツール向けの会議文字起こし。
コンタクトセンターおよびサポート分析のための通話分析。
高速なドラフト文字起こしが必要なコンテンツ作成ワークフロー。
推論前に音声からテキストへ変換するボイスエージェント。
入力言語が不明な場合、自動言語識別機能が役立ちます。このモデルは手動設定なしで発話された言語を検出します。
MAI-Transcribe-1.5 と MAI-Transcribe-1 の比較
以下の表は、公表されている事実のみを用いて 2 つの世代を比較したものです。
属性 | MAI-Transcribe-1 | MAI-Transcribe-1.5
---|---|---
対応言語数 | 25 | 43
キーワード/エンティティバイアス | 記載なし | 最大 200 キーワード
長文推論速度 | ベースライン | 最大 5.7 倍高速化
Artificial Analysis WER | 未指定 | 2.4%(ランク 3 位)
FLEURS 順位(Microsoft 発表) | 最先端技術 | 43 言語で最高クラス性能
自動言語識別 | 未指定 | あり
ライフサイクル | 先行リリース | 一般提供 (GA)
入力/出力 | オーディオ / テキスト | オーディオ / テキスト
強みと制限
強み:
単一モデルによる 43 カ国語対応は、従来の 25 カ国語から拡大されました。
キーワードやエンティティのバイアス適用により、FLEURS 評価において WER(単語誤り率)が最大 30% 削減されます。
1 時間の音声データを 15 秒未満で転写可能です。
現在、Azure AI Foundry を通じて一般利用が可能になりました。
Microsoft によると、ノイズの多い実世界の音声に対しても堅牢です。
制限事項:
ダイアライゼーション(話者分離)機能はまだ提供されていないため、話者ラベルは使用できません。
ネイティブなストリーミング API が存在しないため、リアルタイム利用には制限があります。
精度、速度、コストに関する複数の主張は、第一当事者(Microsoft 自身)によるものです。
Artificial Analysis のランキングでは、2 つの競合他社に次いで第 3 位です。
出典
MAI-Transcribe-1.5 の紹介 — Microsoft AI
MAI-Transcribe-1.5 モデルカード — Azure AI Foundry
MAI-Transcribe-1.5 Foundry API ドキュメント
MAI-Transcribe-1.5 クックブック
MAI プレイグラウンド
本記事「Microsoft AI Introduces MAI-Transcribe-1.5: 2.4% WER on Artificial Analysis, Best-in-Class FLEURS Accuracy, and Up to 5x Faster Long-Audio Transcription」は、もともと MarkTechPost で公開されたものです。
原文を表示
Last week Microsoft AI has announced MAI-Transcribe-1.5. It is the second iteration of the company’s in-house speech-to-text family. The model targets accuracy across 43 languages, accents, and noisy environments. The Microsoft team positions it for production transcription workloads.
What is MAI-Transcribe-1.5
MAI-Transcribe-1.5 is an automatic speech recognition (ASR) model. It takes audio as input and returns text. Microsoft built it in-house, not on a third-party base. The model handles 43 languages with a single system. It is optimized for diverse accents, dialects, and real-world acoustic conditions.
Microsoft is integrating it into Copilot, Teams, GitHub, and Dynamics 365 Contact Centre. It is also available in Foundry, Microsoft’s model platform.
The Accuracy Case
Accuracy here is measured by Word-Error-Rate (WER). Lower WER means fewer mistakes per transcribed word. Microsoft reports best-in-class WER across 43 languages on FLEURS. FLEURS is a standard multilingual transcription benchmark.
On the Artificial Analysis leaderboard, the model posts a WER of 2.4%. That places it third on a competitive open benchmark. So the picture is split. Microsoft team claims first place on FLEURS and third on Artificial Analysis.
The language expansion is the other accuracy story. Coverage grew from 25 languages to 43. The 18 new languages were added without compromising accuracy. Ten of them are South Asian, including Bengali, Tamil, and Telugu. Eight are European, such as Ukrainian, Greek, and Catalan.
Speed
MAI-Transcribe-1.5 leads on accuracy-times-speed on the Artificial Analysis leaderboard. It runs up to 5x faster than models of comparable accuracy. The effect is largest on long audio files. The model can transcribe an hour of audio in under 15 seconds.
Microsoft cites up to 5x speedups over Gemini 3.1, Scribe v2, and GPT-4o-Transcribe on long audio. Against the prior MAI-Transcribe-1, the Azure card lists up to 5.7x faster long-form inference. For batch pipelines processing large archives, that latency gap compounds quickly.
Keyword (Entity) Biasing: The Feature Worth Understanding
Generic transcribers often fail on domain-specific words. These include people, product names, medical terms, and internal acronyms. Those words frequently matter most to enterprise users.
MAI-Transcribe-1.5 adds keyword biasing, also called entity biasing. You supply a list of domain-specific keywords. The Azure card supports up to 200 keywords. The model biases its predictions toward that list. Critically, it does not blindly force matches. It uses shared context to decide when biasing should apply. Microsoft reports a 30% WER reduction on FLEURS when biasing is used.
A short example shows the effect. Without biasing, names render as “Sean,” “Oif,” and “Societal.” With a supplied name list, the model recovers “Shaun,” “Aoife,” and “Xochitl.” This is relevant for meetings, healthcare, and call centers with niche vocabulary.
Use Cases
The Azure model card lists concrete production scenarios. Each maps to a common engineering workload:
Video captions for media and content platforms.
Accessibility tools that depend on accurate captions.
Meeting transcription for Teams-style collaboration tools.
Call analysis for contact centers and support analytics.
Content creation workflows that need fast draft transcripts.
Voice agents that convert speech to text before reasoning.
Automatic language identification helps when the input language is unknown. The model detects the spoken language without a manual setting.
MAI-Transcribe-1.5 vs MAI-Transcribe-1
The table below compares the two generations using stated facts only.
AttributeMAI-Transcribe-1MAI-Transcribe-1.5
Languages covered2543
Keyword/entity biasingNot listedUp to 200 keywords
Long-form inference speedBaselineUp to 5.7x faster
Artificial Analysis WERNot specified2.4% (ranked #3)
FLEURS position (per Microsoft)State-of-the-artBest-in-class across 43 languages
Automatic language identificationNot specifiedYes
LifecyclePrior releaseGenerally available (GA)
Input / OutputAudio / TextAudio / Text
Strengths and Limitations
Strengths:
43-language coverage from a single model, up from 25.
Keyword/entity biasing yields up to 30% WER reduction on FLEURS.
Sub-15-second transcription for an hour of audio.
Generally available now through Azure AI Foundry.
Robust on noisy, real-world audio, per Microsoft.
Limitations:
No diarization yet, so speaker labels are unavailable.
No native streaming API, so real-time use is limited.
Several accuracy, speed, and cost claims are first-party.
Ranked third on Artificial Analysis, behind two competitors.
Sources
Introducing MAI-Transcribe-1.5 — Microsoft AI
MAI-Transcribe-1.5 model card — Azure AI Foundry
MAI-Transcribe-1.5 Foundry API documentation
MAI-Transcribe-1.5 Cookbook
MAI Playground
The post Microsoft AI Introduces MAI-Transcribe-1.5: 2.4% WER on Artificial Analysis, Best-in-Class FLEURS Accuracy, and Up to 5x Faster Long-Audio Transcription appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み