MiniMax、歌詞と構造化キャプションから 5 分曲を生成する音楽モデル公開
本文の状態
日本語全文を表示中
詳細モードで約6分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
MiniMax は歌詞と構造化キャプションを入力として受け取り、最大 5 分間の完全な楽曲を生成するオープンウェイトモデル「MiniMax-Music3」を発表し、商用利用可能なライセンスと共に即時デプロイ可能であることを示した。
AI深層分析を開く2026年8月18日 03:46
AI深層分析
キーポイント
即時デプロイ可能なオープンウェイトモデルの公開
MiniMax は重み、推論コード、および 3 つのドキュメント化されたサービングパスを同日で公開し、研究段階のプレビューではなく即座に実装可能な状態であることを示した。
ハイブリッド LLM とフローマッチングを組み合わせた新アーキテクチャ
8B のグローバル LLM と 0.6B のローカル LLM を組み合わせ、連続合成スタックにフローマッチングと Flow-VAE を採用することで、構造的な長距離依存性と詳細な音響特性を同時に処理する。
商用利用を可能にする明確なライセンス条件
コミュニティライセンスにより商用利用が許可される一方、製品 UI への明記義務と年間収益 2,000 万ドルを超える組織による事前承認要件が設定されている。
多様な産業分野での応用可能性
ゲーム開発、広告、短編動画、e ラーニング、ポッドキャストなど、背景音楽や適応型サウンドが必要な幅広い業界での即座の適用が期待される。
ハイブリッドアーキテクチャと合成方式
8BのグローバルLLMと0.6BのローカルLLMを組み合わせ、離散トークナイザーデコーダーをスキップして連続隠れ状態上で合成を行う。
重要な引用
Yes, MiniMax published usable weights, inference code and three documented serving paths on day one, so this is deployable now rather than a research preview.
The MiniMax-Music3 Community License permits commercial use, but it requires you to display 'MiniMax-Music3' prominently in the product UI
Rather than decoding from discrete RVQ tokens, MiniMax fuses the final hidden states of both LLMs and conditions a 2.4B flow-matching module on them
Synthesis runs on fused continuous hidden states, skipping the discrete tokenizer decoder entirely.
編集コメントを表示
編集コメント
MiniMax は、単なるモデルの公開に留まらず、ライセンス条件と実装コードを同時に提供することで、業界における「研究から実用へ」の壁を低くする明確な意図を示した。特に収益規模に応じたライセンス区分は、大企業向けと小規模クリエイター向けのバランスを取る戦略的アプローチと言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
MiniMax は、テキストから音楽を生成するオープンウェイトモデル「MiniMax-Music3」を発表しました。このモデルは、セクションタグを含む歌詞と、詳細な楽曲説明という 2 つの別入力を受け取り、最大 5 分間の完全な楽曲を 1 回の生成で出力します。音声フォーマットはサンプリングレート 32kHz、16 ビットのステレオ WAV です。
アーキテクチャでは、8B のグローバル LLM と 0.6B のローカル LLM を組み合わせた Hybrid-LM に、フローマッチングと Flow-VAE(変分オートエンコーダー)を基盤とした連続合成スタックを組み合わせました。重みデータ、推論コード、および 3 つのドキュメント化された運用パスは、すべて同日に公開されています。
実用性は確保されているのでしょうか?
はい、MiniMax は公開初日に使用可能な重みデータ、推論コード、そして 3 つの運用ガイドを公開しました。つまり、これは研究段階のプレビューではなく、すぐに導入して使える状態です。
企業レベルでは、ソロクリエイター、インディースタジオ、中堅チームが直接このモデルを活用して製品をリリースできます。MiniMax-Music3 コミュニティライセンスは商用利用を許可していますが、製品の UI において「MiniMax-Music3」の表示を明確に行う必要があります。また、これらの製品からの年間総収益が 2,000 万ドルを超える組織は、別途 MiniMax から事前の書面による承認を得る必要があります。第三者向けの生成サービスを提供する場合は、侵害される可能性のある出力に対する対策の実装と維持も義務付けられています。
活用可能な業界としては、ゲーム開発、広告・ブランディング代理店、ショート動画やクリエイター向けツールの提供、e ラーニング、ポッドキャスト制作、フィットネス・ウェルネスアプリ、小売店舗内のオーディオ、そして音楽テック SaaS などが挙げられます。
ユースケースとしては、UGC(ユーザー生成コンテンツ)動画の背景音楽、ゲームやレベルごとの適応型音楽、ローカライズされた広告用サウンドブランディング、作曲家向けのデモ曲作成、ムードに応じたプレイリスト生成、そして楽曲ごとの API コストが制約となるオフラインバッチ処理などが挙げられます。
アーキテクチャ
MiniMax-Music3 は、階層的な自己回帰スタックと連続的な合成パスを組み合わせたモデルです。
トレーニングに用いるトークナイザーは、8 層の残差ベクトル量子化(RVQ)で構成されています。最初のセマンティックコードブックには 16,384 エントリがあり、音楽のコアとなる意味や構造を担います。残りの 7 つの音響コードブックはそれぞれ 1,024 エントリを持ち、詳細な残留情報を符号化します。トレーニングではまずセマンティック層を最適化し、その後 8 層全体を同時に最適化します。
Hybrid-LM はモデル化問題を分割して扱います。8B のグローバル LLM がフレームごとに最初の RVQ コードブックを予測し、長期的な構造を保持します。一方、0.6B のローカル LLM は各フレーム内で残りのコードブックを予測します。モデルカードとライセンスではグローバル LLM が Qwen3-8B から初期化されていると明記されていますが、MiniMax Research の投稿では Qwen3.5-8B と記載されており、正確なベースチェックポイントについてはまだ確定していません。
合成段階こそが、より興味深い設計選択です。離散化した RVQ トークンからデコードするのではなく、MiniMax は両方の LLM の最終的な隠れ状態を融合させ、それらを条件として 2.4B のフローマッチングモジュールに与えます。このモジュールは、MiniMax Speech から継承された 123M の Flow-VAE によってデコードされる潜在空間へとマッピングされます。推論時には、離散トークナイザーのデコーダーは一切読み込まれません。
2 つの入力による制御
歌詞は、それぞれの行に単語とセクションタグ([Intro]、[Verse]、[Pre-Chorus]、[Chorus]、[Post-Chorus]、[Bridge]、[Instrumental]、[Solo]、[Outro])が記載されます。一方、構造化キャプションにはグローバルメタデータ、ボーカル詳細、編曲情報が別枠で含まれます。MiniMax はまた、短い説明をこの 3 つの部分形式に拡張する「music-caption-rewriter」エージェントスキルも提供しており、これはオフラインでも利用可能です。
インタラクティブな解説ツールがあります。
実行方法として、文書化された 3 つのパスが用意されています。リファレンスサーバーは SGLang-Omni です。GitHub ページによると、2 枚の CUDA GPU を使用し、GPU 0 で Qwen3 と RVQ の自己回帰生成を、GPU 1 でフローマッチングと DAV デコーディングを実行します。
diffusers モジュラーパイプラインは、フル精度で 24 GB 未満の VRAM で動作し、自動 CPU オフロードを使用すれば約 22 GB、リーフレベルでのグループオフロードを利用すれば 8 GB まで圧縮可能です。ComfyUI では、Comfy-Org から再パッケージされた FP16/INT8 重みを使用したネイティブの Text to Music テンプレートが利用できます。
主なポイント
- 2026 年 8 月 13 日に公開されたオープンウェイトモデル。32 kHz、16 ビットステレオで、5 分間の完全な楽曲を生成可能。
- ハイブリッド LM デザイン:8B のグローバル LLM と 0.6B のローカル LLM を組み合わせ、2.4B のフローマッチングと 123M の Flow-VAE にデータを送信。
- 合成処理は融合された連続隠れ状態で行われ、離散トークナイザーのデコーダーを完全にスキップします。
- SGLang-Omni を介して 2 枚の GPU で実行可能。diffusers では 24 GB 未満で動作し、グループオフロードを利用すれば 8 GB でも稼働します。
- 商用利用は表示付きでの許諾が可能ですが、収益が 2000 万ドルを超える場合は書面による承認が必要です。
GitHub リポジトリや Hugging Face ページの紹介、製品発表、ウェビナーなどのプロモーションをご希望の場合は、お気軽にお問い合わせください。
本記事は MarkTechPost にて公開された「MiniMax Releases MiniMax-Music3: An Open-Weights Music Model Generating Complete Five-Minute Songs From Lyrics and a Structured Caption」を翻訳したものです。
原文を表示
MiniMax released MiniMax-Music3, an open-weights text-to-music model. The model takes two separate inputs: lyrics carrying section tags, and a detailed music description. It returns a complete song of up to five minutes in a single generation, as 32 kHz, 16-bit stereo WAV. The architecture pairs a Hybrid-LM, an 8B Global LLM with a 0.6B Local LLM, with a continuous synthesis stack built on flow matching and a Flow-VAE. Weights, inference code and three documented serving paths shipped the same day.
Is it deployable?
Yes, MiniMax published usable weights, inference code and three documented serving paths on day one, so this is deployable now rather than a research preview.
Company level: Solo creators, indie studios and mid-market teams can ship on it directly. The MiniMax-Music3 Community License permits commercial use, but it requires you to display ‘MiniMax-Music3’ prominently in the product UI, and any organization whose aggregate yearly revenue from those products exceeds US$ 20 million must obtain separate prior written authorization from MiniMax. Anyone hosting third-party generation must also implement and maintain safeguards against infringing outputs.
Industries: Game development, advertising and brand agencies, short-form video and creator tools, e-learning, podcasting, fitness and wellness apps, retail in-store audio, and music-tech SaaS.
Applications: Background scoring for UGC video, adaptive game and level music, localized ad beds and sonic branding, scratch and demo tracks for songwriters, mood-conditioned playlist generation, and offline batch generation where per-song API cost is the constraint.
The Architecture
MiniMax-Music3 combines a hierarchical autoregressive stack with a continuous synthesis path.
The training tokenizer uses eight layers of residual vector quantization (RVQ). The first, semantic codebook has 16,384 entries and carries core musical semantics and structure. The remaining seven acoustic codebooks have 1,024 entries each and encode residual detail. Training optimizes the semantic layer first, then all eight jointly.
The Hybrid-LM splits the modeling problem. An 8B Global LLM predicts the first RVQ codebook frame by frame and holds long-range structure; a 0.6B Local LLM predicts the remaining codebooks within each frame. The model card and license state the Global LLM is initialized from Qwen3-8B; the MiniMax Research post says Qwen3.5-8B, so treat the exact base checkpoint as unsettled.
The synthesis stage is the more interesting design choice. Rather than decoding from discrete RVQ tokens, MiniMax fuses the final hidden states of both LLMs and conditions a 2.4B flow-matching module on them, which maps into a latent space decoded by a 123M Flow-VAE inherited from MiniMax Speech. At inference the discrete tokenizer decoder is not loaded at all.
Two-input control
Lyrics carry the words and section tags on their own lines: [Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo], [Outro]. A separate Structured Caption carries Global Metadata, Vocal Details and Arrangement. MiniMax also ships a music-caption-rewriter agent skill that expands a short description into that three-part format offline.
Interactive explainer
Running it
Three documented paths. SGLang-Omni is the reference server; the GitHub page specifies two CUDA GPUs, with GPU 0 running Qwen3 and RVQ autoregressive generation and GPU 1 running flow matching and DAV decoding. The diffusers modular pipeline fits under 24 GB VRAM at full precision, about 22 GB with automatic CPU offload, and down to 8 GB with leaf-level group offloading. ComfyUI has a native Text to Music template using repacked FP16/INT8 weights from Comfy-Org.
Key Takeaways
Open-weights model generating complete five-minute songs at 32 kHz, 16-bit stereo, released August 13, 2026.
Hybrid-LM design: 8B Global LLM plus 0.6B Local LLM, feeding 2.4B flow matching and a 123M Flow-VAE.
Synthesis runs on fused continuous hidden states, skipping the discrete tokenizer decoder entirely.
Runs on two GPUs via SGLang-Omni, under 24 GB via diffusers, or 8 GB with group offloading.
Commercial use allowed with visible attribution; above USD 20M revenue needs written authorization.
Check out the Model on HF and GitHub Repo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post MiniMax Releases MiniMax-Music3: An Open-Weights Music Model Generating Complete Five-Minute Songs From Lyrics and a Structured Caption appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み