Stability AIとArm、スマートフォン向けオンデバイス生成オーディオを実現
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Stability AI
Stability AIはArmと提携し、モバイルデバイス向け生成オーディオ技術「Stable Audio Open」を最適化した。Arm KleidiAIライブラリを活用し、生成速度を30倍高速化。インターネット接続不要で、スマートフォン上で数秒以内に高品質な音響効果やサンプルを生成可能となる。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
主なポイント
私たちは Arm と提携し、モバイルデバイスに生成型オーディオをもたらすことで、インターネット接続を必要とせず、オンデバイスで高品質な音響効果やオーディオサンプルの生成を可能にしました。
Arm の KleidiAI ライブラリと Stability AI の最先端技術である Stable Audio Open を活用することで、Arm CPU 搭載のスマートフォンデバイス上で実行速度が 30 倍向上し、生成時間が数分から数秒へと短縮されました。
この画期的な成果は、2025 年 3 月 3 日(月)にバルセロナで開催される MWC で披露され、エッジにおける前例のない AI 駆動型コンテンツ作成の実証が行われます。本提携の詳細については、Built on Arm ページにてご覧いただけます。
本日、私たちは Arm とのパートナーシップを通じて、最先端の生成 AI モデルをより多くの人々が利用できるようにします。Arm の技術は、世界中のスマートフォンの 99% に搭載されています。Together, we have achieved what was once thought impossible by running Stable Audio Open, our industry-leading text-to-audio model, entirely on Arm CPUs without requiring an internet connection for the first time.
生成 AI が企業とプロのクリエイターの両方にとってますます不可欠なものとなる中、ビルダーが構築し、クリエイターが創作するあらゆる場所でモデルやワークフローを容易に利用可能にし、視覚メディア制作パイプラインへのシームレスな統合を提供することが極めて重要です。
この需要の高まりに伴い、エッジ(端末側)でモデルが効率的に動作することの確保が不可欠です。今回の協力により、音響効果、オーディオサンプル、制作要素を数秒間でオンデバイスかつオフラインで生成することが可能になります。
MWC バルセロナでは、エッジにおける生成メディアの実世界での応用例を紹介し、オンデバイスのテキストからオーディオへのモデルがどのようにして迅速かつ高品質なオーディオ生成を実現するかを実演します。
技術的進展
モバイルデバイス向けに Stable Audio Open を最適化する取り組みは当初大きな課題であり、Arm CPU 上での初期の音声生成には 240 秒を要していました。モデルの蒸留(distillation)と Arm のソフトウェアスタック、特に XNNPack を介した ExecuTorch 内の KleidiAI 由来の int8 matmul カーネル(行列乗算カーネル)を活用することで、Stability AI と Arm は Armv9 CPU 上で 11 秒分のクリップ生成時間を 8 秒未満に短縮し、応答速度を 30 倍向上させることに成功しました。
Stable Audio Open は完全に Arm CPU 上で動作するため、重厚なハードウェア要件なしで利用可能となり、互換性のあるモバイルデバイスを持つ誰でもアクセスできるようになりました。
今後の展望
音声生成は始まりに過ぎません。私たちは画像、動画、3D を含むすべての最先端モデルをエッジ(端末側)へ展開することを目指しています。Arm とのこのパートナーシップは、あらゆる視覚メディアモダリティにおいて高品質なメディア生成をモバイルデバイス上で直接可能にするための重要な一歩であり、視覚メディアの制作方法を変革するものです。
パートナーシップの詳細やデモについては、Built on Arm のウェブページ(here)でご覧いただけます。また、Arm パートナーカタログ内の Stability AI パートナーページ(here)もご参照ください。
今後の進捗状況については、X、LinkedIn、Instagram でフォローいただくか、Discord コミュニティにご参加ください。
原文を表示
Key Takeaways
We’ve partnered with Arm to bring generative audio to mobile devices, enabling high-quality sound effects and audio sample generation directly on-device with no internet connection required.
Leveraging Arm KleidiAI libraries and Stability AI’s cutting-edge technology, Stable Audio Open, can now run 30x faster on smartphone devices on Arm CPUs, reducing generation time from minutes to seconds.
This breakthrough will be showcased at MWC Barcelona on Monday, March 3rd, 2025, demonstrating unprecedented AI-powered content creation at the edge. You can learn about the partnership on the Built on Arm page here.
image
Today, we are making our cutting-edge generative AI models more accessible through our partnership with Arm, whose technology powers 99% of smartphones globally. Together, we have achieved what was once thought impossible by running Stable Audio Open, our industry-leading text-to-audio model, entirely on Arm CPUs without requiring an internet connection for the first time.
As generative AI becomes increasingly integral to both enterprises and professional creators alike, it's crucial that our models and workflows are easily accessible everywhere builders build and creators create, providing seamless integration into their visual media production pipelines.
With this rising demand, ensuring our models run efficiently at the edge is crucial. This collaboration enables generation of sound effects, audio samples, and production elements in seconds all on-device and offline.
At MWC Barcelona, we’ll showcase real-world applications of generative media at the edge, demonstrating how our on-device text-to-audio model enables rapid, high-quality audio generation.
Technical Advancements
Optimizing Stable Audio Open for mobile devices began as a significant challenge, with initial audio generation on an Arm CPU taking 240 seconds. By distilling the model and using Arm’s software stack, including the int8 matmul kernels from KleidiAI in ExecuTorch via XNNPack, Stability AI and Arm reduced the generation time for an 11-second clip to under 8 seconds on Armv9 CPUs, representing a 30x faster response time.
By running entirely on Arm CPUs, Stable Audio Open is now accessible without heavy hardware requirements, making it available to anyone with a compatible mobile device.
What’s Next
Audio is just the beginning. We aim to bring all of our cutting-edge models across image, video, and 3D to the edge. This partnership with Arm is a key step toward enabling high-quality media generation directly on mobile devices across all visual media modalities, transforming how visual media is created.
You can learn more about the partnership and view a demo on the Built on Arm webpage here and visit the Stability AI partner page here in the Arm partner catalog.
To stay updated on our progress follow us on X, LinkedIn, Instagram, and join our Discord Community.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み