Fish Audio、クリエイター向け AI 音声モデル構築に 5000 万ドルシード調達
Fish Audio は、クリエイターおよび企業向けの AI 音声モデル開発を目的として、5000 万ドルのシード資金調達を実施したと発表した。
AI深層分析を開く2026年7月28日 23:58
AI深層分析
キーポイント
大規模資金調達の完了
Fish Audio は Coreline Ventures と Capital Today が主導し、359 Capital や Parable など複数の投資家が参加する形で 5000 万ドルのシードラウンドを完了した。
市場での確固たる地位
同社は昨年の立ち上げ以降、オープンソース版とホスト版合わせて 800 万人以上のユーザーを獲得し、年間経常収益(ARR)が 2100 万ドルに達している。
多様なユースケースへの対応
Fish Audio は 15,000 以上の自然言語コントロールを備えたライブラリにより、クリエイター向けの表現豊かな音声と、企業向けのカスタマーサポートや営業自動化のための制御可能な音声の両方を提供している。
技術的基盤とモデル展開
元 NVIDIA 研究者が単一 GPU でトレーニングしたオープンソースプロジェクトを起源とし、過去 1 年間で 4 つの音声生成モデルと 1 つの音声テキスト変換モデルを発表し、そのうち最新モデルは有料 API のみで利用可能となっている。
音声の自動削除プロセスと課題
Fish Audio は作成者が音声サンプルや契約書を提出すれば3分以内にプラットフォームから音声を削除する自動化プロセスを備えている。しかし、許可なくアップロードされた場合、発見されるまでその音声は使い続けられるという課題が残る。
重要な引用
"Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls," Cao said.
"A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially."
"Fine-grained controls for developers and cost-efficient model training will help Fish Audio compete better with big AI labs."
"I think what they've been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices," Mallozzi told TechCrunch over a call.
編集コメントを表示
編集コメント
Fish Audio の成長は、AI 音声技術が単なる実験段階から実用的なビジネスインフラへと成熟したことを示している。特に、オープンソースコミュニティからの信頼を基盤としつつ、企業向けの高品質・高制御性モデルで収益化に成功した事例として注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI 生成音声モデルの市場は巨大です。クリエイティブな用途では、より表現豊かな音声が必要とされますが、顧客サポートや営業業務の自動化を目指す企業にとっては、制御しやすいモデルが求められます。
パロアルトに拠点を置く Fish Audio は、15,000 以上の自然言語によるコントロール機能を備えたライブラリを通じて、これらの多様なニーズに応えようとしています。昨年の立ち上げ以来、同社のオープンソース版やホスト版を利用するユーザーは 800 万人を超え、年間経常収益(ARR)は 2,100 万ドルに達しています。
この勢いをさらに高めるため、Fish Audio は火曜日に Coreline Ventures と Capital Today が主導し、359 Capital、Parable、Play Time、Alphalist Partners、Bayhouse Ventures、Carya Venture Partners、HF0 が参加するシードラウンドで 5,000 万ドルを調達したと発表しました。
Fish Audio は元 NVIDIA 研究者の廖詩佳(Shijia Liao)氏による小規模プロジェクトから始まりました。市場に存在する表現力に欠ける合成音声に不満を抱いた廖氏は、単一の GPU で音声生成モデルを訓練し、それをオープンソース化しました。現在、GitHub 上の Fish Speech リポジトリはスター数 31,000 を超え、インディーズ開発者やゲームデザイナー、クリエイターたちに利用されています。
同社は過去 1 年間で 5 つのモデルをリリースしました。音声生成モデルが 4 つ、音声からテキストへの変換モデルが 1 つです。そのうち 3 つの音声生成モデルはオープンソース化されていますが、最新の S2.1 Pro モデルは有料 API のみを通じて利用可能です。
Fish Audio は、クリエイターやチーム向けに月額課金プランを提供しており、生成できる音声の分量とボイスクローン機能を解放します。また同社は企業向けの API やプラットフォームも展開しており、HeyGen、Sanas、Plaud といった組織がすでに利用していることを明かしています。
「企業ごとに用途や好みが異なります」と曹氏は語ります。「例えば、AI アバターの音声に Fish Audio の声を活用する HeyGen のような企業は、声のリアリティを重視します。ゲームスタジオならキャラクターに表現豊かな声を求めます。一方、LiveKit などのボイスエージェント企業は、通話に適した自然で低遅延かつ表現力のある声を必要としています」。
同スタートアップが音声ライブラリを構築する一つの方法として、ユーザーに自らの声をトレーニングデータとして提供してもらう仕組みがあります。その声が利用された場合、報酬を支払うというものです。しかし数ヶ月前、この手法が問題を引き起こしました。一部のクリエイターが、同意なく自分の声が Fish Audio にアップロードされたと主張したのです。
同社はこうした懸念に対応するため DMCA によるコンテンツ削除手続きを設けていましたが、実際の削除処理には時間がかかっていました。
Fish Audio のCEO兼共同創設者であるRissa Cao氏は、同社がすでに音声の削除プロセスを自動化したとTechCrunchに語った。クリエイターは、アップロードされた音声の所有権を証明する短い音声サンプルや契約書を簡単に提出でき、その結果、プラットフォームから対象となる音声が3分以内に削除されるという。
ただし、これではアーティストの許可なく他人がその声をアップロードすることを完全に防ぐことはできない。アーティストがそれに気づき、削除を申請するまでの間、その声はプラットフォーム上で使い続けられることになる。
Coreline VenturesのパートナーであるOskue Honda氏は、クリエイターがプラットフォームを信頼していない限り、コミュニティ主導のモデルは機能しないと指摘した。「クリエイターがプラットフォームを信頼することが前提となる場合のみ、コミュニティ中心のアプローチが持続的な強みとなり得ます。つまり、同意、透明性、帰属表示は、後付けの対応ではなく、製品に組み込まれるべきものです。業界は、検証された音声所有権、明確なライセンス条項、簡単な報告・削除プロセスへと移行し、最終的にはクリエイターが音声のライセンスや商業利用によって収益を得られるモデルへと進む必要があります」と彼は述べた。
Cao氏は、同社が当初はクリエイター向けのオープンソースプロジェクトとして製品を提供していた頃は効率的に運営できており資金が必要なかったと振り返る。しかし、より高度なモデルの開発を進めると同時に、投資家の関心が高まる中で企業顧客の受け入れも視野に入れる必要が生じ、資金調達を行うに至ったという。
今後は、Fish Audio は今年中に音声理解モデルの公開を予定しています。また、音声から音声を生成するモデルの開発も進めています。
音声生成市場は競争が激しく、ElevenLabs や WellSaid、Cartesia、Speechify、Async(旧 Podcastle)、Krisp などが、クリエイターや企業の予算を獲得しようと競い合っています。
359 Capital のパートナーであるリコ・マロッツィ氏によれば、開発者向けのきめ細かい制御機能と、コスト効率の高いモデル学習が、Fish Audio が大手 AI ラボとの競争で優位に立つ鍵になるとのことです。
「彼らが現在のチームで実現した最先端のモデルは、資金力のある他の AI ラボや企業と比較しても驚異的です。これは、人工的な声と人間のような声の差を埋める技術的洞察力を示しています」と、マロッツィ氏は TechCrunch のインタビューで語りました。
*当記事内のリンクを通じて購入された場合、私たちは少額のコミッションを受け取る可能性があります。ただし、これは当社の編集の独立性には影響しません。*
原文を表示
The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable.
Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup today has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million.
To continue building on that traction, the startup on Tuesday said it has raised $50 million in a seed round that was led by Coreline Ventures and Capital Today. The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
Fish Audio started as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice generation model on a single GPU, which he open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars, and is used by indie developers, video game designers, and creators.
The company has launched five models in the last year: four speech generation models and one speech-to-text model. It has open-sourced three of its speech generation models, but its latest S2.1 Pro model is available only through its paid API.
Fish Audio offers paid monthly plans suited for creators and teams that unlock a set number of minutes of generation, plus voice cloning features. The company also offers an enterprise version of its APIs and platform, and says organizations like HeyGen, Sanas and Plaud are already using it.
“Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls,” Cao said.
One way the startup has built its library of voices is by simply asking users to submit their own voices for training its models, and compensating them if their voices are used. That resulted in some trouble a few months ago, however, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA content take-down process in place to address such concerns, but the take-downs themselves took a long time.
Fish Audio’s CEO and co-founder Rissa Cao told TechCrunch that the company has now automated the take-down process. Creators can easily submit a short voice sample or a contract to prove that an uploaded voice belongs to them, and their voice will be taken off the startup’s platform in less than 3 minutes, she said.
Still, that doesn’t prevent anyone from uploading an artist’s voice without their knowledge. And until the artist finds out, their voice will continue to be used on the platform until they file for it to be taken down.
Oskue Honda, a partner at Coreline Ventures, said a community-driven model only works when creators trust the platform.
“A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially,” he said.
Cao said when the startup was only offering its product as an open-source project with plans for creators, it was running efficiently and didn’t need money. But it wanted to develop more advanced models, and also wanted to accommodate enterprises as investor interest was ramping up, which led it to seek capital.
Looking ahead, Fish Audio plans to release an audio understanding model this year. It’s also building a speech-to-speech model.
The speech generation market is crowded, with companies like ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp competing for creators and enterprises’ wallets.
According to Rico Mallozzi, a partner at 359 Capital, fine-grained controls for developers and cost-efficient model training will help Fish Audio compete better with big AI labs.
“I think what they’ve been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices,” Mallozzi told TechCrunch over a call.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
AI算出
資金調達・M&Aainew評価高い
記事は Fish Audio の 5000 万ドル調達と、その背景にある技術的特徴(S2.1 Pro モデルやオープンソース戦略)および倫理的課題への対応を詳細に伝えている。新規性については、単なる資金調達の発表だけでなく、具体的なモデル名や削除プロセスの自動化など独自情報が含まれているため高評価としたが、日本固有の情報は限定的である。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 75
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み