ZDNET が選ぶ音声入力 AI ツール 3 つと無料ツールの比較
本文の状態
日本語全文を表示中
詳細モードで約24分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
ZDNET AI
ZDNET の記者が自身の膨大な音声入力実績に基づき、自己修正機能や信頼性を重視した 3 つの AI 音声入力ツールを比較し、その中から有料版と無料ローカル版の優劣を報告している。
AI深層分析を開く2026年8月6日 22:24
AI深層分析
キーポイント
音声入力の現状分析
記者は 2 月以来で約 120,896 単語、過去 3 ヶ月だけで 50,234 単語の音声を記録し、AI の進歩により精度と速度が劇的に向上したと評価している。
Wispr Flow の優位性
比較対象の中で Wispr Flow が自己修正機能、語彙力、信頼性のすべての項目で最も優れた結果を示し、有料版の優勝者として選定された。
FluidVoice の台頭
無料でローカル環境で動作する FluidVoice が有料版の優勝者にほぼ匹敵する性能を発揮しており、コストを抑えたいユーザーにとって有力な選択肢となる。
音声入力を選択する主な理由
著者は親指の腱鞘炎症状を軽減するために、キーボード入力の代わりに音声を主要な入力モードとして採用している。
Vibe Coding におけるワークフローへの統合
Claude Code や OpenAI の Codex を使用した開発、Slack や Google Chat でのコミュニケーション、メール返信、ノート作成、および AI プロンプト入力に音声を頻繁に活用している。
重要な引用
Fast self-correction matters more than raw accuracy alone.
Wispr Flow led on corrections, vocabulary, and reliability.
Free, local FluidVoice nearly matched the paid winner.
I use voice dictation a lot when working with Claude Code or OpenAI's Codex to vibe code any of the products I'm working on.
編集コメントを表示
編集コメント
記事は特定の製品名を挙げているが、これはあくまで記者個人の体験に基づく評価であり、業界全体での絶対的な優劣を示すものではない。読者は各ツールの最新仕様や自社の環境との適合性を別途確認する必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

*ZDNET をフォロー:* *優先ソースとして追加* on Google.*
ZDNET の注目ポイント
- 正確さそのものよりも、素早い自己修正能力が重要だ。
- 修正機能、語彙力、信頼性の面で Wispr Flow が首位に立った。
- 無料でローカル環境で動作する FluidVoice は、有料の優勝候補とほぼ互角の結果を出した。
2 月以降、私は音声入力を通じて 120,896 単語を記録してきました。直近 3 ヶ月の間にだけで 50,234 単語に達しています。これまでにマイク録音セッションは合計 2,314 回行い、平均して 1 回の入力で約 24 単語を入力しています。
AI の進化に伴い、音声入力の品質や速度、精度も飛躍的に向上しました。ここ数年、私は何度も音声認識ツールを試してきましたが、最近になるまで思うような成果は得られませんでした。
一般的な記事はおよそ1000語程度です。その基準で考えると、2月以降に私が音声入力した文章の量は、約120本分に相当します。
関連記事: 元懐疑派が語る、ChatGPTの音声モードを7つの意外な活用術
もちろん私がタイピングしないわけではありません。実際には多くの文章を入力していますが、音声入力も非常に頻繁に利用しています。私は音声入力を主要な入力手段として採用しました。その主な理由は2つあります。
まず一つ目は、手首の負担を軽減できる点です。私の手首は車掌症候群(カーパル・トンネル症候群)のような症状が出やすい体質です。音声入力に切り替えることで、手首を使う回数を減らすことができます。
本記事では、今年私が実際に検討し、時間を割いて使用した3つの主要なツールをご紹介します。また、それら以外の候補や注目すべきツールについても触れ、市場にある音声入力ツールの種類とそれぞれの違いを把握していただけるように解説します。
私のワークフローにおける音声入力の位置づけ
私はClaude CodeやOpenAIのCodexを使って「バイブコーディング(直感でコードを書く作業)」を行う際、音声入力を頻繁に利用しています。また、編集者やプロジェクトパートナーとのやり取りでは、SlackやGoogle Chatでも同様に活用しています。
また、音声だけで2日間でiOSアプリを構築した記事もご覧ください。その体験は非常に刺激的でした。
私はメールの返信では頻繁に音声入力を利用しています。新規作成時にも使いますが、その頻度はやや低くなります。メモを取る際にも多用します。さらに、Googleでの検索クエリや、ChatGPT、Gemini、Claude Codeへのプロンプト入力としても時々利用しています。
(開示:ZDNETの親会社であるジフ・デイビス社は2025年4月、OpenAIを提訴しました。同社がAIシステムの訓練および運用においてジフ・デイビス社の著作権を侵害したと主張しています。)
また、音声とマウスだけで2つのアプリを構築した記事もご覧ください。IDE(統合開発環境)はもう不要になっているのでしょうか。
音声入力は、キーボードで一字一句入力するよりも精度がやや低く、規律に欠けると感じることもあります。しかし、私の平均的な音声入力の速度は1分間に118語です。一方、タイピングの速度は通常70〜80語/分の範囲にとどまります。
私はこの記事のほとんどを音声入力で作成しています。これは「記事も音声入力で書ける」ということを証明するための実験ですが、私の日常ではありません。音声入力はよく行いますが、記事そのものを音声で書くことは通常ありません。数行だけ声に出すことはあっても、基本的にはタイピングで書いています。一語一句を慎重に選び、構成にも細心の注意を払うためです。タイピングなら、そうした細部へのこだわりが生まれやすいのです。
一方、「バイブコーディング」をしているときは事情が違います。機能の実装方法や挙動の希望、あるいは発見したバグについて話す場合など、音声の方がはるかに速く、手にも優しいからです。
最も重要な機能
私が特に重要だと考える機能が二つあります。一つ目はカスタマイズ可能な辞書です。例えば「ZDNET」と発音しても、製品側が私の望む綴り方や表記を理解できるようにします。二つ目はその場で修正できる機能です。発言した後に自分で訂正して言い直した場合、入力先のテキストにその訂正内容が反映されるようにするものです。
関連記事: AI ツール専門家が現在有料で使っている 4 つのツールと注目している 2 つ
音声入力製品のほとんどは、ホットキーで起動します。私はこのホットキーをマウスのボタンに割り当てており、ボタンをタップするだけで入力が開始・停止できます。これにより、現在どのアプリケーションやウェブページにいるかに関係なく、いつでも音声入力が可能になります。つまり、キーボードが目の前になくても、コンピュータへの入力ができるのです。
そのために、私は以下の 3 つの製品を取り上げます。Wispr Flow、Superwhisper、そして FluidVoice です。
1. Wispr Flow:総合評価最高
年間 144 ドル、あるいは月額 15 ドルと決して安くはありませんが、これは間違いなく投資する価値があるツールだと私は考えます。他の製品をほぼすべて試してきた私にとって、これが唯一の常連であり、最も頻繁に利用しているのが Wispr Flow です。
*Wispr Flow の統計画面(David Gewirtz/ZDNET 撮影)*
Wispr Flow は、紹介する上位 3 つのツールの中で唯一、Mac、Windows、iOS、Android に対応しています。ただし Linux 版はまだ提供されておらず、Linux ユーザー向けには待機リストが設けられています。
私が特に評価している Wispr Flow の最大の特徴は、「使用中の自己修正機能」です。発話中に誤りを指摘すると、その場で書き起こされたテキストが自動的に更新されます。この機能に慣れてしまうと、もう元に戻りたくなりません。私のように大量の音声入力を行う人にとっては、特におすすめです。
つまり、発話ミスがあっても話しながら簡単に修正できるため、結果として実用的なテキストが得られるケースが多々あるのです。
私が試した他のモデルは、これほどスムーズにはできませんでした。中には全くできないものもありました。
FluidVoice は近いところまで来ましたが、Wispr Flow は 10 回中 8 回の頻度で正確な修正を行っていました。一方、FluidVoice は 10 回中 4 回程度でした。飛行中の自己修正機能は、多くの作業を行う際にその差が明確に測れるものです。
関連記事: 音声からテキストへ変換する AI モデル 3 つをテストして最良のものを探しました - その結果
また、Wispr Flow の辞書機能も信頼性が高く効果的だと分かりました。具体的には、誤ったスペルや誤認識された単語を学習させてから一度登録すれば、その後ほとんど再修正する必要がなくなります。
例えば「ZDNET」という単語を学習させると、Wispr Flow はほぼ 100% の確率で正しく認識します。私が日常的に修正している約 90 個の単語群についても同様です。
Wispr Flow には 2 つの辞書登録オプションがあります。一つは「Gewirtz」のような個別の単語を直接入力する方法、もう一つは誤記や誤認識された語句と、その正しい表記をセットで登録する方法です。
例えば、Claude という単語について頻繁に「call it」と誤認識される問題がありました。そこで「call it code」という定義を追加し、「Claude Code」に変換するルールを設定したところ、それ以来全く問題がなくなりました。
稀に、Wispr Flow は指定した場所にテキストの断片を挿入し忘れることがあります。例えば、Notes アプリに追加したい段落を発話しても、その場所へ反映されないケースです。
ただし Wispr Flow には発話履歴がアプリ内に保存されるため、挿入に失敗した場合でも、アプリを開いて履歴からコピーして貼り付ければ大丈夫です。最初の試行で目標の場所に届かないことがあっても、発話内容自体を失うことはありません(これは非常に稀なケースですが)。
*直近の発話履歴が記録されたクリップボード画面 (David Gewirtz/ZDNET 撮影)*
価格面以外で私が最も懸念しているのは、Wispr Flow がクラウドベースのモデルである点です。つまり、すべての音声スニペットは転送されてクラウド上で文字起こしが行われます。
名前が似ているからといって、Wispr Flow が OpenAI のオープンソース音声認識技術「Whisper」を基にしているわけではありません。Wispr Flow は独自モデルか、あるいは複数のモデルプロバイダーを組み合わせたスタックに基づいているようです(詳細はこちら)。同社は使用している具体的なモデルについては明言していません。
関連記事: ChatGPT のライブ音声アップグレードを試した感想 - 人間そっくりで驚き、使い方も解説
Wispr Flow の設定には、プライバシーモードやプライベートクラウド同期のオン/オフ切り替え、ローカルデータストレージなど、さまざまなデータおよびプライバシーオプションが用意されています。ただし、UI 上に表示される項目名と、実際にそれぞれの機能がどう動作するかは一致していません。
*プライバシーオプションのスリーンショット(David Gewirtz/ZDNET 撮影)*
「プライバシーモード」という名称から想像するほど、すべての入力フレーズを監視しないわけではありません。このモードを有効にすると、AI の学習データとして利用するためにユーザーのデータを外部へ送信しなくなります。
「プライベートクラウド同期」をオフにしても、データがクラウドに転送されないという意味ではありません。これは、他のデバイスと同期するためにクラウド上にデータを保存しないことを意味します。音声認識処理のためにデータはクラウドへ転送されますが、Wispr Flow は認識完了直後にそのデータを即座に削除します。
「ローカルデータストレージ」オプションも、データがローカルに保存されるのかクラウドに保存されるのかを制御するものではありません。これは、例えば Wispr Flow がローカルデータを 24 時間ごとに自動削除するか、あるいは一切ローカルに保存しないかなどを制御するものです。後者の場合、音声入力履歴が利用できなくなります。
関連記事: Gmail の AI ツールで 10 分間に数時間の作業を完了させた体験 - プロンプトは 3 つだけ
データ管理ポリシーへの懸念、機密保持の必要性、開示制限、あるいはクラウド上にデータを置いたくないその他の法的理由がある場合は、Wispr Flow の利用を避けたほうがよいでしょう。
基本的な生産性向上の観点から、私が最も頻繁に活用している音声入力ツールは Wispr Flow です。これまで数多くのツールを試してきましたが、日常使いとして長く続けられるものを探し求めていました。その結果、Wispr Flow が私のトップ推奨品となりました。
Wispr からは評価用に 1 年間の Pro アカウントを提供してもらいましたが、この期間が終了した際には自費で更新する可能性が高いです。これだけで、その価値のほどがお分かりいただけるでしょう。また、Wispr Flow には最大 2,000 単語まで利用可能なトライアル版も用意されています。
2. Superwhisper: 音声入力のカスタマイズに最適
2 つ目のツールである Superwhisper は、機能の宝庫です。しかし残念ながら、私にはしっくり来ませんでした。紹介しているのは、あなたにとって役立つ特殊な機能が多数搭載されているからです。
オンザフライでの修正機能がないことが、私の使用を断念させる決定的な理由になりました。正直に言うと、この点については驚きました。自分がこれまでどれほど積極的にその機能を活用していたかには気づいていなかったのですが、それが欠如していることに気づいた瞬間、その重要性を痛感したのです。
Superwhisper の利用料は月額 8.49 ドル、年額 84.99 ドル、あるいは一度払いで 249 ドルの永久ライセンスがあります。長期間使用する予定なら、この永久ライセンスは非常に魅力的な選択肢です。一方で、AI ベースのソリューションは急速に進化しており、永久ライセンスを完全に活用する前に、さらに優れた、あるいはより安価な、あるいは無料の代替案が登場している可能性も十分にあります。音声認識モデルが小さいバージョンであれば、無料で利用することも可能です。
関連記事: 実際に支払う価値がある AI ツールはどれか?私が 2026 年も継続するサブスクリプションとその理由
念のため付け加えておくと、上位 2 つの製品の間で認識精度に大きな差があるわけではありません。どちらも全体的な認識タスクにおいて成功するモデルを実行可能です。Superwhisper は多くのモデル選択肢を提供していますが、その中には音声認識の精度は劣るものの、他の点で優れているものも含まれています。
*Superwhisper のモデル選択画面(David Gewirtz/ZDNET 撮影)*
Superwhisper の Wispr Flow に対する最大の強みは、Mac 専用でオフライン動作し、クラウドへのデータ送信も不要なバージョンを構築できる点です。一度課金して永続ライセンスを購入すれば、追加請求は一切なく、所有権と完全なコントロールを得られます。これにより、すべての音声入力を完全にローカル環境で行うことが可能になります。Superwhisper はその実現を約束しますが、Wispr Flow はできません。
Superwhisper の際立った特徴は、そのモードシステムです(詳細はこちら)。モードとは保存された処理パイプラインのことで、現在アクティブなモードに応じて、Superwhisper の動作や機能が根本的に変化します。
*Superwhisper でモードを作成する様子(David Gewirtz/ZDNET 撮影)*
モードは以下の 4 つの要素で構成されます。
- 音声モデル: 発話内容をテキストに変換する音声認識エンジンです。
- 言語モデル: 生きた文字起こしデータを整形・修正する部分です。
- 処理指示セット: そのモードに紐づくレシピやスキルのような、具体的な処理手順の集合体です。
- 自動起動ルール: 特定のアプリやウェブサイトを使用している際にモードを自動的に有効にする設定です。例えば、Gmail を使用中は A モードが、Notion では B モードが、Apple Notes では C モードがそれぞれ自動的に作動するような設定が可能です。
音声認識のモードシステムを使えば、単なる文字起こしを超えた驚くべきことが可能です。例えば、「ドキュメントのディレクトリを表示して」と口頭で指示するだけで、ターミナルに直接実行可能なコマンドに変換してくれる機能があります。
また、複数の項目を口頭で指定すると、その内容をそのまま JSON や YAML 形式で出力することもできます。
関連記事: デスクトップ向け AI ツールは数多く試しましたが、Ollama を使った Hermes が私の新お気に入りです。その理由はこちら
さらに「悪魔の弁護人モード」という機能もあります。文や段落、あるいは概念を口頭で入力すると、Superwhisper はそのまま貼り付けるのではなく、内部で処理して、あなたが主張した内容に対する反論を即座に生成します。学生や、執筆中に概念的な挑戦を動的に行いたい研究者にとって有用でしょう。強力でありながら非常に特化した機能ですが、万人向けではないものの、確かに面白い技術です。
私は Superwhisper の無料トライアル期間を超えて利用しませんでした。その理由は、特別な機能や設定オプション、モードが豊富にあるにもかかわらず、肝心の音声入力自体の精度が私の期待に届かなかったからです。音声入力製品を選ぶ際、最も重要な要件はやはり正確な文字起こし能力です。
音声入力システムを構築し、発話された言葉を他の形式に変換したいのであれば、Superwhisper がおすすめです。しかし、単に話しかけて使用しているアプリケーション内できれいな文字起こしを行いたいだけなら、他の 2 つのツールよりも Superwhisper を推すことはできません。
Superwhisper は Mac、Windows、iOS で利用可能です。Android 版はありません。
3. FluidVoice:最良の無料オプション
他の 2 つの製品と比較すると、FluidVoice は圧倒的なコスパです。無料であり、オープンソースでもあります。また、ほとんどの用途では Wispr Flow とほぼ同等の性能を発揮します。
*Fluid-1 エンジンの選択画面(David Gewirtz/ZDNET 撮影)*
FluidVoice は、その場での修正や辞書登録された単語の処理にはやや苦戦します。例えば、「ZDNET」という単語に対しては非常に苦手です。私はすでに辞書に「ZDNET」を追加し、何度も訂正を行い、音声トレーニングも実施しましたが、それでも認識に失敗してしまいます。「ZDNET」と発話すると、毎回手動で入力し直す必要があります。
*FluidVoice の修正テスト(David Gewirtz/ZDNET 撮影)*
FluidVoice は完璧ではありませんが、その場で修正する機能は時としてうまく機能します。残念ながら、この機能が有効になるのは半分以上の時間ではないものの、全く使えないよりはマシです。
*音声認識ポップアップ(David Gewirtz/ZDNET 撮影)*
FluidVoice の優れた点の一つに、発話中にテキストをプレビューできる小さなポップアップウィンドウがあることが挙げられます。私は基本的な Mac の音声処理機能を使用していた際にこの機能に慣れました。書き込んでいるテキストに貼り付ける前に、自分が何と言ったかを確認し、それがどのように解釈されているかを確認するのは役立ちます。
FluidVoice には独自ネイティブの「Fluid-1」言語処理モデルを搭載していますが、OpenAI の Whisper とも連携可能です。Intel Mac で FluidVoice を利用したい場合は Whisper に接続する必要があります。私がテストしている Fluid モデルを利用するには Apple Silicon 搭載の Mac が必要です。現時点では Windows や iOS 向けのバージョンはありませんが、同社によると両方のプラットフォーム向け開発は進行中です。
関連記事: リアルで信頼性の高い製品を迅速にリリースするために私が活用する AI コーディングテクニック 7 つ
無料かつオープンソースであることに加え、もう一つの大きな利点は、FluidVoice を完全オフラインで使用できることです。つまり、すべての音声入力が端末内に留まり、データ管理について懸念する必要がありません。
もちろん、OpenAI の Whisper などのクラウドモデルを選択して、処理能力が低い Intel Mac で実行する場合は、クラウド側での処理が必要になります。
現時点では、Wispr Flow を使い続ける予定です。これは私が行っている作業において最も信頼性の高いソリューションであり、特にその場で修正できる機能と辞書機能が優れているからです。ただし、音声認識のために Netflix の月額料金にほぼ匹敵する金額を支払い続けるのが面倒な場合は、更新時期には FluidVoice に切り替えるかもしれません。
Lightning round(早押しラウンド)
もし音声入力製品について少しでも時間を費やして調べているなら、数多くの候補に出会うことになるでしょう。私の推奨は上記の 3 つの中から選ぶことです。ただし、ここでは追加の候補をいくつか紹介する「早押しラウンド」です。
Also-rans and honorable mentions(惜しくも敗れた候補と特別賞)
これらはシステム全体で使える音声入力ツールなので、検討の価値があります。
「MacWhisper」(https://cc.zdnet.com/v1/otc/00hQi47eqnEWQ6T9d4QLBUc?element=BODY&element_label=MacWhisper&group_uuid=d840ecc2-6364-4e9c-b415-23a5036c699b&module=LINK&object_type=commerce-link&object_uuid=34b35fbf-5058-43ab-81d4-5477568dda49&object_version=25c968a7-532d-4210-8662-de16f890a870&position=1&template=article&track_code=__COM_CLICK_ID__&view_instance_uuid=dd0b45fe-144b-4b67-8169-137b08c9b2d1&url=https%3A%2F%2Fwww.macwhisper.com%2F): デバイス上で動作する文字起こし機能と、オプションの音声入力機能を備えています。*強み:* 精度が高く、プライバシーに配慮しており、字幕生成や話者識別も可能です。*弱み:* 音声入力が主目的ではなく、Mac専用です。*価格:* 無料プランあり、Pro版は約69米ドルの買い切り。
「Paraspeech」(https://cc.zdnet.com/v1/otc/00hQi47eqnEWQ6T9d4QLBUc?element=BODY&element_label=Paraspeech&group_uuid=6054ccea-743e-43fd-8d43-6a6d937afdcf&module=LINK&object_type=commerce-link&object_uuid=34b35fbf-5058-43ab-81d4-5477568dda49&object_version=25c968a7-532d-4210-8662-de16f890a870&position=1&template=article&track_code=__COM_CLICK_ID__&view_instance_uuid=dd0b45fe-144b-4b67-8169-137b08c9b2d1&url=https%3A%2F%2Fparaspeech.com%2F): Macで動作する低価格な音声入力ツールです。*強み:* 安価、オフライン対応、買い切りオプションあり。*弱み:* 使用モデルが非公開、レビュー数が極めて少ない。*価格:* 月額8.99米ドル、年額89米ドル、または買い切り(ローカル利用のみ)。
「Vibe コーディング」中に音声入力に役立つ、私が最もお気に入りの AI ツール 3 つ(そのうち 1 つは無料)
- VoiceInk: Mac 向けのオープンソース音声入力ツール。ローカルで動作する Parakeet AI モデルを採用しています。
*強み:* ローカル環境、プライバシー保護、カスタマイズ性、オープンソース
*弱み:* Mac のみに限定される、セットアップに手間がかかる
*価格:* コンパイルして自作すれば無料、バイナリ版を購入すると 25〜49 ドル
- Handy: 無料でオープンソースのオフライン音声入力ツール。
*強み:* 無料、クロスプラットフォーム対応、完全なプライバシー保護
*弱み:* 機能は最小限、使い心地に荒削りな部分がある
*価格:* 無料(オープンソース)
「Talon Voice」: 音声による完全な制御とコーディングが可能。*強み:* パワフルでハンズフリーのコーディング、スクリプト対応。*弱み:* 習得に時間がかかる。*価格:* 無料(ベータ版は有料 Patreon)。
OS ネイティブアプリ
MacOS や Windows を使用している場合、これらのツールはオペレーティングシステムに標準搭載されています。
「バイブコーディング」中に音声入力に使う AI ツールで私が気に入っている 3 つと、1 つは無料です。
MacOS の内蔵音声入力: Mac に標準搭載されたリアルタイム音声入力機能。*強み:* 無料で、端末内で処理され、リアルタイム文字変換に対応。*弱み:* 辞書機能が貧弱で、誤字訂正が苦手。*価格:* 無料(標準搭載)
MacOS の音声コントロール: アクセシビリティ向けの音声による編集・コマンド機能。*強み:* 音声での編集が可能で、カスタム語彙に対応し、オフライン動作も可能。システム操作には非常に強力。*弱み:* 習得に時間がかかる。コード入力には向いておらず、オン/オフの切り替えが面倒。*価格:* 無料(標準搭載)
Windows の音声入力 (Win + H): Windows に標準搭載されたクラウド型音声入力機能。*強み:* 無料で使いやすく、句読点の自動挿入に対応。*弱み:* クラウド依存で、誤字訂正機能が弱い。*価格:* 無料(標準搭載)
Windows の音声アクセス: アクセシビリティ向けの音声による編集・コマンド機能。*強み:* オフライン動作が可能で、音声での修正やカスタム語彙に対応。*弱み:* 習得に時間がかかる。コード入力には向いていない。*価格:* 無料(標準搭載)
Claude、ChatGPT、Gemini の音声入力機能について
Anthropic と OpenAI は最近、それぞれ音声入力のオプションを発表しました。どちらも実際にかなり優秀ですが、利用可能なのはそれぞれのサービス内のインターフェースに限られています。
関連記事: 2026 年最高の AI チャットボット:専門家によるテストとレビュー
「Claude の音声入力とDictation」:アシスタントとの会話に限定。*強み*:正確な文字起こし、自然な会話形式。*弱み*:アプリ内のみでシステム全体のDictation には対応せず(Cowork や Claude Code では利用不可)。*価格*:無料(すべてのClaude プランに含まれる)。
「ChatGPT の音声入力とDictation」:アシスタントとの会話に限定。*強み*:編集可能なDictation、高度な音声モード。*弱み*:AI との対話に主眼が置かれている。*価格*:無料(すべてのChatGPT プランに含まれる)。
「Gemini の音声入力とDictation」:アシスタントとの会話に限定。*強み*:Gemini Live による連続的な会話フロー、途中での自然な割り込み対応。*弱み*:アプリ内のみでシステム全体のDictation には対応せず、Web インターフェースの音声入力は基本的な機能にとどまる。*価格*:無料(すべてのGemini プランに含まれる)。
精度がかなり向上している
ここ数年、私は音声入力ツールを断続的に試してきましたが、成功は限定的でした。まだ完璧ではありませんが、現在は実用レベルに達しています。Wispr Flow の利用統計によると、2月以降で10万語以上の文字起こしを行ったことがわかりました。Wispr Flow が日々の生産性フローの重要な一部になっていることは知っていましたが、これほどまでに日常業務の中心となっているとは思いもよりませんでした。
関連記事: ChatGPT、Gemini、Copilot、Claudeとの会話をいかにプライバシー保護するか
さまざまなツールの利用経験について、ぜひお聞かせください。例えば、Wispr Flow のより確実な校正機能のために有料版を利用するか、無料で使える FluidVoice を使い続けるか、ご自身の判断をお聞かせいただければ幸いです。
日々のプロジェクトの進捗状況は、ソーシャルメディアで随時更新しています。週刊ニュースレター my weekly update newsletter の購読や、Twitter/X @DavidGewirtz、Facebook Facebook.com/DavidGewirtz、Instagram Instagram.com/DavidGewirtz、Bluesky @DavidGewirtz.com、YouTube YouTube.com/DavidGewirtzTV でのフォローもぜひご検討ください。
原文を表示

*Follow ZDNET: *Add us as a preferred source* on Google.*
ZDNET's key takeaways
- Fast self-correction matters more than raw accuracy alone.
- Wispr Flow led on corrections, vocabulary, and reliability.
- Free, local FluidVoice nearly matched the paid winner.
Since February, I have dictated 120,896 words. In the last three months alone, I have dictated 50,234 words. I've done this across 2,314 individual microphone recording sessions, averaging about 24 words per dictation sequence.
With improvements in AI, dictation quality, speed, and accuracy have come a long way. Over the years, I've tried to work with speech recognition many times, but it hasn't been very successful until quite recently.
A typical article is roughly a thousand words. So if you look at it that way, I have dictated the equivalent of roughly 120 articles since February.
Also: 7 surprisingly useful ways to use ChatGPT's voice mode, from a former skeptic
That's not to say I don't type. I type a lot, but it's clear that I also use dictation *a lot*. I've adopted voice dictation as a primary input modality for two key reasons. First, it helps to protect my wrist, which tends to have carpal tunnel symptoms. By dictating, I'm using my wrist a little bit less.
In this article, I'll show you the three primary contenders that I've looked at and spent time with this year. I'm also going to go over a number of also-rans and honorable mentions, just to give you an idea of the dictation tools that are out there and how they differ.
Where voice dictation fits into my workflow
I use voice dictation a lot when working with Claude Code or OpenAI's Codex to vibe code any of the products I'm working on. I use it a lot in Slack and Google Chat when talking to my editors and some of my project partners.
Also: I built an iOS app in just two days with just my voice - and it was electrifying
I use voice dictation quite a lot when replying to email messages. I use it a little bit less when composing new messages, but I use it there sometimes as well. I use it a lot when taking notes. I also use it sometimes for search phrases in Google or prompts that I give to ChatGPT, Gemini, or Claude Code.
*(Disclosure: Ziff Davis, ZDNET's parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)*
Also: I built two apps with just my voice and a mouse - are IDEs already obsolete?
I find that voice dictation is sometimes a little bit less precise and a little bit less disciplined than writing each word with the keyboard. But I can dictate at an average of 118 words per minute, while my typing is usually in the 70- to 80-word-per-minute range.
I am dictating most of this article just as a proof of concept to make the point that an article can be dictated. But that's not the norm for me. Despite all the dictation I do, I don't usually dictate my articles. I sometimes dictate a few lines of an article, but mostly I type them out. I sweat each sentence, choosing the words and structure with great care. Typing lends itself to that kind of attention to detail.
But when I'm vibe coding, for example, and I'm discussing how I want a feature to be instantiated, how I want something to behave, or a bug that I've noticed, speaking is considerably faster and also considerably more gentle on my hands than typing it into the computer.
The features that matter most
I find two features to be mission-critical. The first is a customizable dictionary, so that when I say something like ZDNET, the dictation product understands how I want it spelled and presented. The second is an on-the-fly correction capability, so that when I say something and then correct myself and re-say it, the version that lands in whatever I'm dictating into contains those corrections.
Also: I'm an AI tools expert, and these are the 4 I pay for now (plus 2 I'm eyeing)
Almost all the dictation products are initiated by a hotkey. I bind the dictation hotkey to a button on my mouse so that when I tap the button, the dictation starts or stops. This allows me to dictate regardless of what application or web page I'm in at the moment. It means I can do computer input even if my keyboard isn't in front of me.
To that end, I will be spotlighting three products: Wispr Flow, Superwhisper, and FluidVoice.
1. Wispr Flow: Best overall
At $144 a year, or $15 a month, Wispr Flow is certainly not cheap, but I would argue it's actually worth it. Despite trying almost all the other products, this is the one I keep coming back to and have used more than any other.
*Wispr Flow stats Screenshot by David Gewirtz/ZDNET*
Wispr Flow is the only one of our top three available for Mac, Windows, iOS, and Android. It is not, however, available on Linux, although the company has a waitlist for Linux users.
Wispr Flow's standout feature, at least in terms of my usage, is its in-flight self-correction. As you're dictating, you can correct yourself, and it updates what's being transcribed. Once you get used to this feature, you really don't want to go back, especially if you're doing a large amount of dictation like I do.
It means that the text you produce is, more often than not, usable because if you misspeak, you can fairly easily correct it as you're speaking and end up with a decent result.
None of the other models that I tested were able to do this as smoothly. Some couldn't do it at all. FluidVoice has come close, but I would say that Wispr Flow made accurate corrections eight out of 10 times, and FluidVoice made accurate corrections maybe four out of 10 times. For in-flight self-correction, that's measurable when you're doing a lot of work.
Also: I tested 3 text-to-speech AI models to see which is best - hear my results
I also found that Wispr Flow's dictionary is reliable and effective. What I mean by that is that once I've trained it on an incorrectly spelled word or incorrectly interpreted word, I almost never have to go back and correct it again. Once I trained it on the word ZDNET, for example, Wispr Flow reliably gets it correct just about 100% of the time. That's also the case with my library of 90 or so other words that I regularly correct.
Wispr Flow has two dictionary options: It allows you to feed it individual words like Gewirtz, and it allows you to feed it misspellings or misinterpretations and then the corrected word. For example, it regularly had trouble with the word Claude, which it would represent as "call it." I set up a dictionary definition for "call it code" that converted to Claude Code, and I've never had a problem since.
Once in a while, Wispr Flow misses the insertion of a chunk of text into the destination location. For example, I might dictate a paragraph that I want to go into Notes, and it never winds up there. Wispr Flow keeps a history of dictation in its app. If it misses insertion, I can open it up in the app, copy from the history, and paste it in. I don't ever actually lose any of my dictation, even if it doesn't always arrive on target the first time out (which is a fairly rare occurrence).
*Clipboard with recent dictation Screenshot by David Gewirtz/ZDNET*
Beyond price, my biggest concern about Wispr Flow is that it's a cloud-only model, meaning that all of your voice snippets are sent to the cloud for transcription. Despite the similarity in names, Wispr Flow is not based on OpenAI's open-source Whisper speech recognition technology. Wispr Flow appears to be its own model or based on a stack of a variety of model providers. The company does not disclose the exact model used.
Also: I tested ChatGPT's Live Voice upgrade, and it almost felt human - how to try it
Wispr Flow offers a number of data and privacy options in its settings, including a privacy mode, the option to turn private cloud sync on and off, and local data storage. However, what the options are called in the UI and what the options actually do are different.
*Privacy options Screenshot by David Gewirtz/ZDNET*
Privacy mode isn't really what you would think. It's not that it doesn't look at any of your phrases. It's that when turned on, it will not send any of your data to be used for training the AI.
Private Cloud Sync, when turned off, does not mean that the data is not sent up to the cloud. It means that it's not stored in the cloud to sync to other devices. It is still sent up to the cloud for transcription, but Wispr Flow then deletes the data immediately after transcription.
The local data storage option does not control whether data is stored locally or in the cloud, but instead controls factors like whether or not Wispr Flow will auto-delete local data every 24 hours or never store any data locally, meaning, for example, that the dictation history would not be available to you.
Also: I used Gmail's AI tool to do hours of work for me in 10 minutes - with 3 prompts
If you have data control policy concerns, confidentiality concerns, disclosure restrictions, or any other legal reason you don't want your data up in the cloud, you might want to avoid Wispr Flow.
I have found, for basic productivity, that Wispr Flow has become my most actively used voice dictation product. I have been cycling through a bunch of them to try to find one that I could live with as a daily driver. So far, that's Wispr Flow, and that's why it's my top recommendation.
Wispr provided me with a Pro account to use for a year for evaluation, but there's a very good chance that when that year runs out, I will renew it with my own money. That should tell you something. There is a trial version of Wispr Flow that allows you to use it for up to 2,000 words.
2. Superwhisper: Best for voice dictation customization
Our second tool, Superwhisper, is just chock-full of features. Even so, I just really haven't been able to mesh with it. I am presenting it here because it does have so many specialized capabilities that you may find helpful.
I found the lack of on-the-fly correction to be a deal killer. To be honest, that surprised me because I didn't even realize that I had been using the on-the-fly correction as actively as I was until it became apparent that when it was missing, I missed it greatly.
Superwhisper is available for $8.49 a month, $84.99 a year, or a one-time purchase of $249 that provides unlimited lifetime use. If you think you're going to be using it for a number of years, that's a good deal. On the other hand, AI-based solutions are changing so rapidly that there may be a far better solution, a far cheaper solution, or a free solution available before you fully utilize the unlimited lifetime use. There is a free version that you can use with smaller voice recognition models.
Also: Which AI tools are actually worth paying for? I'm keeping these subscriptions in 2026 - here's why
I should point out that recognition accuracy is not really that much of an issue between the two top products. Both are able to run models that are successful for overall recognition. Although Superwhisper provides a lot of model choices, some of them do not perform voice recognition as accurately but have other advantages.
*Superwhisper's model selection Screenshot by David Gewirtz/ZDNET*
Superwhisper's biggest advantage over Wispr Flow is that you can create a Mac-only, offline-only, no-data-in-the-cloud version. If you pay for the one-time lifetime use with no additional billing ever, you can own it and control it all and have all your dictation happen entirely on your computer. Superwhisper will do that for you. Wispr Flow will not.
Superwhisper's standout fiddly feature is its mode system. Modes are saved processing pipelines. Basically, what this means is that depending on what mode you're in, Superwhisper can behave or function completely differently.
*Creating a mode in Superwhisper Screenshot by David Gewirtz/ZDNET*
A mode consists of four elements:
- The voice model, which is a speech-to-text engine that transcribes the spoken word for you.
- The language model, which is the part that cleans up and reshapes the raw transcript or modifies it in some way.
- A set of processing instructions, essentially a recipe or a skill attached to that mode.
- An auto-activation rule, which says that the mode becomes active on a given app or website. For example, you could have a mode that becomes active when you are using Gmail, another mode that becomes active when you are using Notion, and still a third that becomes active when you are using something like Apple Notes.
Beyond processing basic speech, the mode system lets you do some crazy stuff. Take, for example, the idea of dictating a description of a command-line operation that you want to be typed into the terminal, but you describe it like "Give me a directory of my documents." The engine then converts that into an actual command line directly from your dictation.
Another example might be dictating a series of items and having the output be JSON or YAML directly from your dictation.
Also: I've tested so many desktop AI tools, but Hermes with Ollama is my new favorite - here's why
Another one might be a devil's advocate mode where you dictate a sentence, paragraph, or concept. Instead of pasting in the dictation, Superwhisper takes that dictation in, processes it, and spits back out an argument against whatever it is that you asserted. This might be useful for students or people researching concepts that they'd like to have dynamically challenged as writing continues. It's powerful. It's very specialized. It might not be usable by everyone, but it is cool.
I never went beyond the free trial for Superwhisper. That's because, despite all of the special features, configuration options, and modes, actual dictation was not as effective as I wanted it to be. When I am looking at an overall voice dictation product, dictation is at the core of my requirements.
If you want to build a dictation system that takes spoken words and turns them into other forms, then Superwhisper is for you. If you simply want to speak and get clean transcription in whatever application you're using, I would not recommend Superwhisper above the other two.
Superwhisper is available for Mac, Windows, and iOS. There is no Android version.
3. FluidVoice: Best free option
Compared to the other two offerings, FluidVoice is a total bargain. It's free. It's also open source. And it works nearly as well as Wispr Flow in most uses. Almost.
*Choosing the Fluid-1 engine Screenshot by David Gewirtz/ZDNET*
FluidVoice does struggle with on-the-fly correction and dictionary words. For example, FluidVoice has a very hard time with the word ZDNET, even though I've put ZDNET into the dictionary, corrected it multiple times, and done a voice training version of it. It still fails. Whenever I say "ZDNET," I need to retype it by hand.
*Testing FluidVoice correction Screenshot by David Gewirtz/ZDNET*
But even though FluidVoice does struggle, on-the-fly correction does sometimes work fine. Unfortunately, that sometimes happens less than half of the time, but it's better than nothing.
*Speech recognition pop-up Screenshot by David Gewirtz/ZDNET*
One nice feature of FluidVoice is that it has a little pop-up window that lets you preview the text as you're speaking it. I got used to this when I was using the basic Mac voice processing. It can be helpful to remember what you just said and see how it is being interpreted before you have it pasted into the text you're writing.
FluidVoice has its own native Fluid-1 language processing model, but it also works with OpenAI's Whisper. If you want to use FluidVoice on an Intel Mac, you'll want to connect it to Whisper. If you want to use the Fluid model, which is what I've been testing, you'll need an Apple Silicon Mac. There is no version at this time for Windows or iOS, although the company says versions for both are under development.
Also: 7 AI coding techniques I use to ship real, reliable products - fast
In addition to being free and open source, another key advantage is that you can use FluidVoice entirely offline. That means all of your dictation stays on your machine, and you don't have to worry about how it's being managed.
Of course, if you choose to use one of the cloud models, like OpenAI's Whisper, and run it on a lower-powered Intel Mac, then you will have some cloud processing.
For now, I plan to stay with Wispr Flow because it is the overall most reliable solution for the work I'm doing, especially because of the on-the-fly correction and its dictionary capability. But if I don't feel like spending almost as much as a Netflix subscription for speech recognition, I may move to FluidVoice when it comes time to renew.
Lightning round
If you spend any time at all looking at voice dictation products, you'll run into a whole bunch of contenders. My recommendation is that you choose from the above three. But here's a quick lightning round of additional contenders.
Also-rans and honorable mentions
These are systemwide dictation tools you might want to consider.
- MacWhisper: On-device transcription plus add-on dictation; Strengths: Accurate, private, subtitles, and diarization; Weaknesses: Dictation not its main service, Mac-only; Price: Free tier; about $69 one-time for Pro
- Paraspeech: Cheap local Mac dictation; Strengths: Inexpensive, on-device, lifetime option; Weaknesses: Undisclosed models, almost no reviews; Price: $8.99 a month, $89 a year, or lifetime (local only)
- VoiceInk: Open-source Mac dictation, local Parakeet AI model; Strengths: Local, private, customizable, open; Weaknesses: Mac-only, some setup; Price: Free if you compile it yourself; $25 to $49 if you buy a binary
- Handy: Free, open-source, offline dictation; Strengths: Free, cross-platform, fully private; Weaknesses: Bare-bones, rough edges; Price: Free (open source)
- Talon Voice: Full voice control and coding; Strengths: Powerful, hands-free coding, scriptable; Weaknesses: Steep learning curve; Price: Free; paid Patreon for beta
OS-native apps
If you're running MacOS or Windows, these come with the operating system.
- MacOS Dictation: Built-in live Mac dictation; Strengths: Free, on-device, live text; Weaknesses: Weak dictionary, no correction; Price: Free (built in)
- MacOS Voice Control: Accessibility voice editing and commands; Strengths: Spoken editing, custom vocabulary, offline, quite powerful for system manipulation; Weaknesses: Learning curve, not code-friendly, not easy to toggle on or off; Price: Free (built in)
- Windows Voice Typing (Win + H): Built-in Windows cloud dictation; Strengths: Free, easy, auto-punctuation; Weaknesses: Cloud-based, weak correction; Price: Free (built in)
- Windows Voice Access: Accessibility voice editing and commands; Strengths: Offline, spoken correction, custom vocabulary; Weaknesses: Learning curve, not code-friendly; Price: Free (built in)
Claude, ChatGPT, and Gemini dictation
Both Anthropic and OpenAI have made recent announcements about their voice input options. They're actually pretty good, but they're limited to use in their own interfaces.
Also: The best AI chatbots of 2026: Expert tested and reviewed
- Claude voice and dictation: Speak to the assistant only; Strengths: Clean transcription, spoken conversation; Weaknesses: In-app only, not systemwide (and off in Cowork and Claude Code); Price: Free, included on all Claude plans.
- ChatGPT voice and dictation: Speak to the assistant only; Strengths: Editable dictation, Advanced Voice mode; Weaknesses: Primarily aimed at interaction with the AI; Price: Free, included in all ChatGPT plans.
- Gemini voice and dictation: Speak to the assistant only; Strengths: Gemini Live for continuous conversational flow, natural mid-sentence interruptions; Weaknesses: In-app only, not systemwide dictation, web interface voice input is basic; Price: Free, included on all Gemini plans.
They're getting quite good
I've tried voice dictation tools on and off over the years, with limited success. They're still not perfect, but voice dictation has reached the point where it's perfectly usable. Wispr Flow tracks usage stats. I was quite taken aback to realize I had dictated more than 100,000 words since the beginning of February. I knew that Wispr Flow had become a key part of my daily productivity flow, but I didn't know it had become quite that central to my daily work.
Also: How to keep your conversations with ChatGPT, Gemini, Copilot or Claude as private as possible
I'm particularly interested in hearing about your experiences with the various tools, so please feel free to share. Would you pay for Wispr Flow's more reliable corrections instead of using FluidVoice for free? Let us know in the comments below.
*You can follow my day-to-day project updates on social media. Be sure to subscribe to my weekly update newsletter, and follow me on Twitter/X at @DavidGewirtz, on Facebook at Facebook.com/DavidGewirtz, on Instagram at Instagram.com/DavidGewirtz, on Bluesky at @DavidGewirtz.com, and on YouTube at YouTube.com/DavidGewirtzTV.*
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み