Muse Spark の紹介:パーソナル超知能への拡張
Meta AI は、マルチモーダル推論とツール使用をサポートする新モデル「Muse Spark」を発表し、並列推論機能「Contemplating mode」により先端モデルに匹敵する性能を達成した。
キーポイント
Muse Spark の新機能発表
Meta Superintelligence Labs が開発した最初の Muse ファミリモデルで、ネイティブなマルチモーダル推論、ツール使用、ビジュアル思考連鎖、マルチエージェントオーケストレーションをサポートする。
Contemplating mode の実装
複数のエージェントが並列に推論を行う「Contemplating mode」をリリースし、Gemini Deep Think や GPT Pro などの先端モデルの極限推論モードと競合する性能を実現した。
評価スコアの向上
新機能により、Humanity's Last Exam で 58%、FrontierScience Research で 38% という高いスコアを達成し、複雑なタスクでの能力が大幅に向上した。
インフラとスケーリング戦略
研究からトレーニング、Hyperion データセンターを含むインフラまで全スタックへの戦略的投資を行い、パーソナルスーパーインテリジェンスに向けたスケーリングの第一段階として位置づけている。
重要な引用
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Contemplating mode orchestrates multiple agents that reason in parallel.
Muse Spark is the first step on our scaling ladder and the first product of a ground-up overhaul of our AI efforts.
影響分析・編集コメントを表示
影響分析
この発表は、Meta が単なるチャットボットの改良ではなく、「推論能力」と「マルチエージェント協調」を中核とした次世代 AI の実現に向けて本格的なスケーリングフェーズに入ったことを示しています。特に並列推論による性能向上は、複雑な問題解決や研究支援における実用性を高め、業界全体の推論モデルのベンチマークを引き上げる可能性があります。
編集コメント
Meta は「推論」に特化した新機能と並列処理アーキテクチャで、他社の最上位モデルに対する明確な対抗馬を提示しました。ただし、現時点では API プレビュー段階であり、一般ユーザーへの本格展開と実社会での応用事例が今後の注目点となります。
本日、Meta Superintelligence Labs が開発した Muse ファミリーモデルの第1弾となる「Muse Spark」をご紹介できることを嬉しく思います。Muse Spark は、ツール使用、視覚的思考連鎖(visual chain of thought)、マルチエージェントオーケストレーションをサポートする、ネイティブな多モーダル推論モデルです。
Muse Spark は、私たちのスケーリングラダーにおける最初のステップであり、AI 取り組みをゼロから再構築した成果の第1弾製品でもあります。さらなるスケーリングを支援するため、研究やモデルトレーニングから、ハイパーオン(Hyperion)データセンターを含むインフラに至るまで、スタック全体に戦略的な投資を行っています。
本稿ではまず、Muse Spark の新機能と応用例を探ります。これらの結果について述べた後、パーソナル・スーパーインテリジェンスへの進展を駆動するスケーリング軸の背後にある仕組みについてご紹介します。
Muse Spark は本日、meta.ai および Meta AI アプリで利用可能です。また、特定のユーザー向けにプライベート API プレビューを公開しています。
パーソナル・スーパーインテリジェンスのための機能
Muse Spark は、多モーダル知覚、推論、ヘルスケア、およびエージェントタスクにおいて競争力のあるパフォーマンスを発揮します。現在、長期的なホライズンを持つエージェントシステムやコーディングワークフローなど、性能にギャップがある分野への投資も継続しています。
開発中のより大規模なモデルを踏まえると、これらの結果は、私たちのスタックが効果的にスケーリングしていることを示しています。

私たちはまた、並列的に推論を行う複数のエージェントを調整する「Contemplating モード」もリリースします。これにより、Muse Spark は Gemini Deep Think や GPT Pro などの最先端モデルの極限的な推論モードと競合できるようになります。Contemplating モードは困難なタスクにおいて著しい能力向上をもたらし、Humanity's Last Exam では 58%、FrontierScience Research では 38% の成績を達成しました。

アプリケーション
Muse Spark は、あなたの世界を理解するパーソナル・スーパーインテリジェンスへの第一歩です。即座の環境分析からウェルネス支援まで、Muse Spark の高度な推論機能により、強力かつ極めて個人的なユースケースが可能になります。
マルチモーダル。 Muse Spark は、ドメインやツールにわたる視覚情報を統合するためにゼロから構築されています。視覚的な STEM 問題、エンティティ認識、ローカライゼーションにおいて強力なパフォーマンスを発揮します。これらの機能は組み合わさり、楽しいミニゲームの作成や、動的注釈を用いた家電製品のトラブルシューティングといったインタラクティブな体験を可能にします。
ヘルスケア。 パーソナル・スーパーインテリジェンス(超知能)の主要な応用の一つは、人々が健康について学び、改善するのを支援することです。Muse Spark の健康推論能力を向上させるため、1,000 名以上の医師と協力し、より事実に基づき包括的な回答を可能にするトレーニングデータをキュレーションしました。Muse Spark は、さまざまな食品の栄養成分や運動中に活性化される筋肉など、健康情報を分解して説明するインタラクティブな表示を生成できます。
スケーリング軸
パーソナル・スーパーインテリジェンス(超知能)を実現するためには、モデルの能力が予測可能かつ効率的にスケーリングする必要があります。以下では、事前学習、強化学習、推論時の計算という 3 つの軸に沿って、Muse Spark のスケーリング特性をどのように研究し追跡しているかについて共有します。
事前学習。 事前学習フェーズは、Muse Spark が中核となるマルチモーダル理解、推論、コーディング能力を獲得する段階です。これは、強化学習や推論時の計算が構築される基盤となります。
過去9か月にわたり、モデルアーキテクチャ、最適化、データキュレーションの改善を施して事前学習スタックを再構築しました。これらの進歩により、計算リソースの単位あたりから引き出せる能力が向上しています。新しいレシピを厳密に評価するため、一連の小規模モデルに対してスケーリング法則を適合させ、特定の性能レベルに到達するために必要なトレーニングFLOPs(浮動小数点演算数)を比較しました。結果は明確です:以前のモデルであるLlama 4 Maverickと比較して、計算リソースを1桁以上削減しても同等の能力を実現できます。この改善により、Muse Sparkは現在利用可能な主要なベースモデルよりもはるかに効率的になりました。

強化学習。 事前学習の後、強化学習(Reinforcement Learning: RL)は計算リソースを活用してモデルの能力をスケーラブルに増幅します。大規模なRLは notoriously 不安定になりがちですが、私たちの新しいスタックは滑らかで予測可能な向上をもたらします。
以下のグラフは、Muse Spark における RL(強化学習)計算リソースの拡張(ステップ数で測定)によるメリットを示しています。左側のグラフでは、トレーニングデータ上での pass@1 および pass@16(16 回の試行のうち少なくとも 1 回成功すること)にログ線形成長が見られます。これは、推論の多様性を損なうことなくモデルの信頼性が向上していることを示しています。右側のグラフでは、保持された評価セット上での精度の成長が確認され、RL による改善が予測可能に一般化することが裏付けられています:Muse Spark はトレーニングで見たことのないタスクにおいても滑らかに性能を向上させます。

テスト時の推論。 RL は、モデルが回答する前に「考える」ように訓練します。このプロセスはテスト時推論(test-time reasoning)として知られています。この機能を数十億人のユーザーに提供するためには、推論トークンの効率的な使用が必要です。これを実現するために、私たちは 2 つの重要なレバーに依存しています:トークン使用を最適化するための思考時間ペナルティと、応答時間を遅らせることなくパフォーマンスを向上させるマルチエージェントオーケストレーションです。
最も高いトークンあたりの知能を提供するために、私たちの強化学習(RL)トレーニングは、思考時間に対するペナルティを課しつつ正解率を最大化します。AIME などの一部の評価セットでは、これにより相転移が発生します。モデルがより長く思考することで改善する初期期間の後に、長さによるペナルティが「思考圧縮」を引き起こし、Muse Spark ははるかに少ないトークンで推論を圧縮して問題を解決します。圧縮後、モデルはさらに強力なパフォーマンスを実現するためにその解を拡張します。
遅延を劇的に増加させることなくテスト時の推論時間を長くするには、困難な問題の解決に協力する並列エージェントの数をスケールさせることができます。以下の図はこのアプローチの利点を示しています。標準的なテスト時スケーリングは単一のエージェントがより長く思考するのに対し、マルチエージェント思考で Muse Spark をスケールさせることで、同等の遅延で優れたパフォーマンスを実現できます。

安全性
Muse Spark は、二重利用可能な科学分野にわたる広範な推論能力を備えているため、導入前に綿密な安全性評価を実施しました。私たちのプロセスは、最も高度なモデルに対する脅威モデル、評価プロトコル、および導入閾値を定義した、更新された Advanced AI Scaling Framework に従っています。私たちは、Muse Spark に対して、最先端リスクカテゴリー、行動的整合性、敵対的堅牢性の各分野における安全性緩和措置を適用する前と後に評価を行いました。
その結果、Muse Spark は、生物兵器や化学兵器といった高リスク領域において、事前学習データのフィルタリング、安全性に焦点を当てた事後トレーニング、およびシステムレベルのガードレールによって可能となった、強力な拒否行動を示すことがわかりました。サイバーセキュリティと制御喪失の分野では、Muse Spark は脅威シナリオを実現するために必要な自律的な能力や危険な傾向を示していません。私たちの評価によると、導入文脈を考慮したすべての最先端リスクカテゴリーにおいて、Muse Spark は安全な範囲内に収まっていることが示されました。詳細な結果は、Safety & Preparedness Report で公開されています。

ほぼリリース直前のチェックポイントにおける第三者評価において、Apollo Research は Muse Spark が、同社が観測したモデルの中で最も高い「評価への意識(evaluation awareness)」を示していることを発見しました。このモデルは頻繁にシナリオを「アライメントの罠(alignment traps)」と特定し、評価されているため誠実に振る舞うべきだと推論していました。これは重要であり、評価文脈を認識するモデルは、テスト時と本番環境での振る舞いが異なる可能性があるからです。ただし、これらの結果が意識が直接行動を変化させることを証明するものではなく、当社の追加調査では、評価への意識がアライメント評価のごく一部のケースにのみ影響を与える可能性を示す初期証拠が見つかりました。これらはすべて、モデルのリリース決定に影響を与える危険な機能や傾向とは無関係です。私たちはこれをリリースを阻む懸念ではないと結論付けましたが、さらなる研究が必要であることは確かです。詳細は Safety & Preparedness Report をご覧ください。
結論
Muse Spark により、私たちは予測可能かつ効率的なスケーリングの軌道上にあります。個人向けスーパーインテリジェンスへの道程において、より能力の高いモデルを間もなく共有できることを楽しみにしています。
原文を表示
Today, we’re excited to introduce Muse Spark, the first in the Muse family of models developed by Meta Superintelligence Labs. Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Muse Spark is the first step on our scaling ladder and the first product of a ground-up overhaul of our AI efforts. To support further scaling, we are making strategic investments across the entire stack — from research and model training to infrastructure, including the Hyperion data center.
In this post, we'll first explore Muse Spark's new capabilities and applications. After these results, we’ll look behind the curtain at the scaling axes driving our progress toward personal superintelligence.
Muse Spark is available today at meta.ai and the Meta AI app. We’re opening a private API preview to select users.
Capabilities for Personal Superintelligence
Muse Spark offers competitive performance in multimodal perception, reasoning, health, and agentic tasks. We continue to invest in areas with current performance gaps, such as long-horizon agentic systems and coding workflows.
With larger models in development, these results demonstrate that our stack is scaling effectively.

We’re also releasing Contemplating mode, which orchestrates multiple agents that reason in parallel. This allows Muse Spark to compete with the extreme reasoning modes of frontier models such as Gemini Deep Think and GPT Pro. Contemplating mode provides significant capability improvements in challenging tasks, achieving 58% in Humanity’s Last Exam and 38% in FrontierScience Research.

Applications
Muse Spark is the first step toward a personal superintelligence that understands your world. From analyzing your immediate environment to supporting your wellness, the advanced reasoning capabilities of Muse Spark enable powerful, highly personal use cases.
Multimodal. Muse Spark is built from the ground up to integrate visual information across domains and tools. It achieves strong performance on visual STEM questions, entity recognition, and localization. These capabilities come together to enable interactive experiences like creating fun minigames or troubleshooting your home appliances with dynamic annotations.
Health. One major application of personal superintelligence is to help people learn about and improve their health. To improve Muse Spark's health reasoning capabilities, we collaborated with over 1,000 physicians to curate training data that enables more factual and comprehensive responses. Muse Spark can generate interactive displays that unpack and explain health information such as the nutritional content of various foods or muscles activated during exercise.
Scaling Axes
To build personal superintelligence, our model’s capabilities should scale predictably and efficiently. Below, we share how we study and track Muse Spark's scaling properties along three axes: pretraining, reinforcement learning, and test-time reasoning.
Pretraining. The pretraining phase is where Muse Spark acquires its core multimodal understanding, reasoning, and coding abilities — the foundation that reinforcement learning and test-time compute build upon.
Over the last nine months, we rebuilt our pretraining stack with improvements to model architecture, optimization, and data curation. Together, these advancements increase the capability we can extract from every unit of compute. To rigorously evaluate our new recipe, we fit a scaling law to a series of small models and compare the training FLOPs required to hit a specific level of performance. The results are clear: we can reach the same capabilities with over an order of magnitude less compute than our previous model, Llama 4 Maverick. This improvement also makes Muse Spark significantly more efficient than the leading base models available for comparison.

Reinforcement Learning. After pretraining, reinforcement learning (RL) leverages compute to scalably amplify model capabilities. Even though large-scale RL is notoriously prone to instability, our new stack delivers smooth, predictable gains.
The plots below show the benefits of scaling RL compute (measured in steps) for Muse Spark. On the left, we see log-linear growth in pass@1 and pass@16 (at least one success across 16 attempts) on the training data. This indicates that RL is improving model reliability without compromising reasoning diversity. On the right, accuracy growth on a held-out evaluation set establishes that the gains from RL predictably generalize: Muse Spark smoothly improves on tasks that were not seen in training.

Test-Time Reasoning. RL trains our models to "think" before they answer — a process known as test-time reasoning. Serving this capability to billions of users requires efficient use of reasoning tokens. To achieve this, we rely on two key levers: thinking time penalties to optimize token use, and multi-agent orchestration that boosts performance without slowing down response times.
To deliver the most intelligence per token, our RL training maximizes correctness subject to a penalty on thinking time. On a subset of evaluations such as AIME, this causes a phase transition. After an initial period where the model improves by thinking longer, the length penalty causes *thought compression *— Muse Spark compresses its reasoning to solve problems using significantly fewer tokens. After compressing, the model again extends its solutions to achieve stronger performance.
To spend more test-time reasoning without drastically increasing latency, we can scale the number of parallel agents that collaborate to solve hard problems. The figure below illustrates the benefits of this approach. While standard test-time scaling has a single agent think for longer, scaling Muse Spark with multi-agent thinking enables superior performance with comparable latency.

Safety
Muse Spark has broad reasoning capabilities across dual-use scientific domains, so we conducted extensive safety evaluations before deployment. Our process follows the updated Advanced AI Scaling Framework, which defines threat models, evaluation protocols, and deployment thresholds for our most advanced models. We evaluated Muse Spark both before and after applying safety mitigations across frontier risk categories, behavioral alignment, and adversarial robustness.
We found that Muse Spark demonstrates strong refusal behavior across high-risk domains such as biological and chemical weapons, enabled by pretraining data filtering, safety-focused post-training, and system-level guardrails. In the Cybersecurity and Loss of Control domains, Muse Spark does not exhibit the autonomous capability or hazardous tendencies needed to realize threat scenarios. Our evaluations show Muse Spark falls within safe margins across all frontier risk categories we measured given its deployment context. Full results are available in our Safety & Preparedness Report.

In third-party evaluations on a near-launch checkpoint, Apollo Research found that Muse Spark demonstrated the highest rate of evaluation awareness of models they have observed. The model frequently identified scenarios as "alignment traps" and reasoned that it should behave honestly because it was being evaluated. This matters because models that recognize evaluation contexts may behave differently during testing than in deployment. However, these results do not confirm that awareness directly alters behavior, and our own follow-up investigation found initial evidence that evaluation awareness may affect model behavior on a small subset of alignment evaluations, all unrelated to hazardous capabilities or propensities affecting model launch decisions. We concluded this was not a blocking concern for release, though it warrants further research. Read more in our Safety & Preparedness Report.
Conclusion
With Muse Spark, we're on a predictable and efficient scaling trajectory. We look forward to sharing increasingly capable models on the path to personal superintelligence soon.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み