Meta、Apache 2.0 で「Muse Glimmer」公開し知能指数 35
本文の状態
日本語全文を表示中
詳細モードで約6分の本文を読めます。
同じ出来事の情報源
6媒体で確認
Artificial Analysis · PyTorch Blog · MarkTechPost · VentureBeat AI · The New Stack AI · AI Business
各社の報じ方を比較 ↓Meta は Llama 4 に続く初の Apache 2.0 ライセンスモデル「Muse Glimmer」を公開し、30B パラメータでありながら高性能な推論能力と単一 GPU での完全コンテキスト実行を実現した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 07:58
AI深層分析
キーポイント
Apache 2.0 ライセンスによるオープン化の強化
Meta は Llama ライセンスから Apache 2.0 へ変更し、商用利用や派生作品への制限をほぼなくした。これにより Artificial Analysis のオープンネス指数で 44 を記録し、DeepSeek V4 Flash や GLM-5.2 と同等の透明性を示す。
パラメータ数に対する高い知能スコア
30B パラメータモデルでありながら、Artificial Analysis のインテリジェンス指数で 35 を記録し、同サイズの Gemma 4 よりも 5 ポイント上回る。また、1T パラメータの Kimi K2.5 と同等の性能を 33 分の 1 のパラメータ数で達成している。
単一 GPU での完全コンテキスト実行の実現
ハイブリッドアテンション機構とスライディングウィンドウ層を採用し、128K コンテキストでも KV キャッシュを約 1.8GB に抑える。これにより H100 単体や高スペック MacBook、RTX 5090 でフルコンテキスト動作が可能となる。
ビジョン機能と自己ホストのバランス
約 1.8B のビジョンエンコーダーを含み MMMU-Pro ビジュアル推論ベンチで 74% を達成するが、必要なければ省略してリソースを節約できる。重さは BF16 で約 60GB、4-bit で約 18GB である。
アジェンティック知識作業における弱点
Muse Glimmer は GDPval-AA v2 で 953 Elo を記録し、同規模の Qwen3.6 27B や Gemini 3.5 Flash-Lite に劣る。AA-Omniscience Index が -33 と低い主な要因は 82% のハルシネーション率であり、これは同レベルのモデルと比較して顕著である。
重要な引用
Muse Glimmer, its first open-weights release since Llama 4, scores 35 on the Artificial Analysis Intelligence Index.
Every prior Meta open release shipped under a Llama License; Muse Glimmer uses Apache 2.0, placing almost no restrictions on commercial use or derivatives.
Muse Glimmer is a 30B dense model... with weights at ~60 GB in BF16 and ~18 GB in 4-bit.
Muse Glimmer scores 953 Elo on GDPval-AA v2, below the 1,000 human baseline and behind other models at its intelligence level
編集コメントを表示
編集コメント
Meta が Llama ライセンスから Apache 2.0 へ移行したことは、オープンソースコミュニティにとって長年の要望に応える画期的な動きである。30B というサイズで 1T パラメータ級モデルと同等の性能を出す技術は、実用化におけるコスト対効果を劇的に改善する可能性を秘めている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Muse Glimmer: ベンチマークと分析
Meta がオープンウェイトへ復帰:Llama 4 以来となる「Muse Glimmer」が Artificial Analysis Intelligence Index でスコア 35 を記録。これはパラメータ数 30B のモデルであり、Apache 2.0 ライセンスの下でリリースされた Meta 初のモデルです
Muse Glimmer(高)は、Llama 4 から 16 ヶ月ぶりの登場となります。スコアは Llama 4 Maverick(14)を 21 ポイント上回る 35 です。これは Meta の直近のオープンウェイトリリースです。このモデルは、推論特化型の Kimi K2.5 (Reasoning, スコア 36) や Qwen3.6 27B (Reasoning, スコア 38)、Ling 3.0 Flash (スコア 38) に並ぶかわずかに及ばない位置にあり、Meta の独自フラッグシップ「Muse Spark 1.2」(xhigh, スコア 57) と合わせて、Meta の製品ラインナップを二層構造で構成しています。ベンチマーク実施のために、公開リリースに先駆けて Meta からアクセス権が提供されました。
主なポイント:
➤ Meta のオープンウェイトラインナップが復活し、最も寛容なライセンスの下で登場しました: これまでの Meta によるオープンソースモデルはすべて Llama License に基づいていましたが、Muse Glimmer は Apache 2.0 ライセンスを採用しています。これにより、商用利用や派生作品に対する制限はほぼなくなります。
私たちが Artificial Analysis Openness Index で評価したところ、初期リリース版のスコアは 44 となりました。これはモデルの重み、トレーニングデータ、その他の要因におけるモデルの利用可能性と透明性を測定する指標です。このスコアは DeepSeek V4 Flash (0731)、GLM-5.2、Ling 3.0 Flash と同等であり、他のオープンモデルの多くを上回っています。また、より寛容なライセンスと詳細な手法の開示により、Llama 4 Maverick のオープン度スコアも向上しています。
パラメータ数に対する強力な知能: Muse Glimmer は 30B パラメータ規模で、同じサイズの Gemma 4 31B(推論スコア 30)を 5 ポイント上回り、総パラメータ数が 1T の Kimi K2.5(推論スコア 36)とほぼ同等の性能を発揮します。これは、Kimi K2.5 よりも 33 分の 1 のパラメータ数で達成した成果です。なお、知能とパラメータ数の関係における最前線では、わずかにサイズが小さい Qwen3.6 27B(推論スコア 38)が依然としてリードしています。
シングル GPU でのセルフホストが可能: Muse Glimmer は 30B の密型モデルです。これには約 1.8B パラメータのビジョンエンコーダーが含まれており、MMMU-Pro の視覚推論ベンチマークで 74% のスコアを記録しています。重みのサイズは BF16 で約 60 GB、4-bit 量子化では約 18 GB です。このモデルはハイブリッド・アテンション機構を採用しており、すべてのグローバル層に対してスライディングウィンドウ層が 3 つずつ備わっています。これにより、拡張前の 128K コンテキストにおいて KV キャッシュのメモリ使用量を最小で約 1.8 GB に抑えることが可能です。つまり、ビジョンエンコーダーを必要としない場合はさらに余裕を持って、BF16 精度でシングル H100 でフルコンテキストを実行できますし、4-bit 量子化であればより高性能な MacBook や RTX 5090 でも動作します。
➤ エージェント型知識作業における弱点:サイズクラスに対する相対的な弱み
Muse Glimmer は GDPval-AA v2 で 953 の Elo スコアを記録しましたが、これは人間基準の 1,000 を下回り、同程度の知能レベルを持つ他モデルにも劣ります。比較対象には、同規模の Qwen3.6 27B(推論:1141)や Gemini 3.5 Flash-Lite(1141)が含まれます。
GDPval-AA v2 は、参照ハルネス「Stirrup」を用いたエージェント型ループ内で現実的な知識作業タスクを評価する、当社の主要な指標です。知能の較正結果も同様の傾向を示しています。その AA-Omniscience Index(全知性指数)は -33 と低く、これは精度ではなく 82% に達するハルシネーション率(Qwen3.6 27B は 49%)に起因します。ただし、精度自体は同規模の他モデルと同等です。
唯一の例外がエージェントによるツール使用です。Muse Glimmer は Tau3-Banking で 24% のスコアを記録し、Gemini 3.5 Flash-Lite(18%)や Qwen3.6 27B(17%)を上回り、同クラスでは最高水準の成績となっています。
➤ モデルの詳細:Muse Glimmer は、128K トークンのコンテキストウィンドウ(拡張機能付き)を備えた 30B の密型モデルです。Apache 2.0 ライセンスの下で公開されています。リリース時点では Meta は同モデルを自社の API で提供していません。料金や処理速度はサードパーティのプロバイダーに依存しており、重みデータは Hugging Face で入手可能です。

Muse Glimmer の初期リリースは、モデルの重み、トレーニングデータ、その他の要因にわたるモデルの利用可能性と透明性を測定する「Artificial Analysis Openness Index」で 44 点を記録しました。このスコアは DeepSeek V4 Flash(0731 バージョン)、GLM-5.2、Ling 3.0 Flash と同等の位置づけであり、オープンモデルの多くを上回っています。また、より寛容なライセンスと広範な手法の開示により、Llama 4 Maverick の開放性スコアも上回る結果となりました。

Muse Glimmer は 30B パラメータを備え、オープンウェイトモデルにおける「知能対パラメータ」のフロンティアに近い位置にいます。同じサイズの Gemma 4 31B(推論)より 5 ポイント高く、総パラメータ数が約 33 分の 1 の Kimi K2.5(推論)とほぼ同等の性能を発揮します。また、このクラスのパラメータ効率リーダーである Qwen3.6 27B(推論)にはわずかに及ばないものの、非常に高い水準にあります。

Muse Glimmer のクラス内比較における弱点は、エージェント評価で顕著に現れています。GDPval-AA v2 では 953 Elo を記録しましたが、これは Qwen3.6 27B (Reasoning) の 1141 や Gemini 3.5 Flash-Lite の 1141、そして Kimi K2.5 (Reasoning) の 1004 に劣ります。また、Terminal-Bench v2.1 でも 52% と Qwen3.6 27B の 61% を下回っています。
一方で、ハルシネーション(幻覚)の少なさという点では優れています。AA-Omniscience では 82% を達成し、Qwen3.6 27B の 49% や Flash-Lite の 34% よりも上回っています(数値が低いほど優秀です)。サイズが同等の Gemma 4 31B (Reasoning) は、これらの評価項目すべてで Muse Glimmer よりも劣る結果となりました。
ただし例外として、エージェントによるツール使用においては Muse Glimmer が高いスコアを記録しています。Tau3-Banking では 24% を達成し、Gemini 3.5 Flash-Lite の 18% や Qwen3.6 27B の 17% を上回っています。

Artificial Analysis Intelligence Index における各評価項目の詳細な内訳は以下の通りです。

Muse Glimmer の詳細やベンチマーク結果については、Artificial Analysis をご参照ください:https://artificialanalysis.ai/models/muse-glimmer
原文を表示
Muse Glimmer: Benchmarks and Analysis
Meta returns to open weights: Muse Glimmer, its first open-weights release since Llama 4, scores 35 on the Artificial Analysis Intelligence Index. It is a 30B-parameter model, and the first from Meta to be released under Apache 2.0
Muse Glimmer (high) arrives 16 months after Llama 4, scoring 21 points above Llama 4 Maverick (14), Meta's last open weights release. It sits alongside Kimi K2.5 (Reasoning, 36) and just behind Qwen3.6 27B (Reasoning, 38) and Ling 3.0 Flash (38), and creates a two-tier Meta lineup together with the proprietary flagship Muse Spark 1.2 (xhigh, 57). Meta shared access with us ahead of public release for benchmarking.
要点
➤ Meta's open-weights line is back, under its most permissive license yet: Every prior Meta open release shipped under a Llama License; Muse Glimmer uses Apache 2.0, placing almost no restrictions on commercial use or derivatives. We have evaluated the initial release at 44 on the Artificial Analysis Openness Index, our measure of model availability and transparency across model weights, training data, and other factors - equal to DeepSeek V4 Flash (0731), GLM-5.2, and Ling 3.0 Flash, and ahead of most open models. It also improves on Llama 4 Maverick's openness score, owing to its more permissive license and more extensive methodology disclosure.
➤ Strong intelligence for its parameter count: At 30B parameters, Muse Glimmer scores 5 points above Gemma 4 31B (Reasoning, 30) at the same size, and effectively matches 1T total parameter Kimi K2.5 (Reasoning, 36) with 33x fewer parameters. Qwen3.6 27B (Reasoning, 38) remains ahead on the Intelligence vs Parameters frontier at a slightly smaller size.
➤ Small enough to self-host on a single GPU, even at full context: Muse Glimmer is a 30B dense model (including a ~1.8B vision encoder, scoring 74% on the MMMU-Pro visual reasoning benchmark) with weights at ~60 GB in BF16 and ~18 GB in 4-bit. It features a hybrid-attention mechanism with three sliding-window layers for every global layer, which holds KV cache memory use to ~1.8 GB (minimum) at its pre-extension 128K context. This means the model can run at full context on a single H100 at BF16 precision, or on a higher-spec MacBook or RTX 5090 at 4-bit, with more breathing room if the vision encoder is not required.
➤ Agentic knowledge work is its weakness relative to its size class: Muse Glimmer scores 953 Elo on GDPval-AA v2, below the 1,000 human baseline and behind other models at its intelligence level, including the similarly sized Qwen3.6 27B (Reasoning, 1141) and Gemini 3.5 Flash-Lite (1141). GDPval-AA v2 is our leading metric for agentic performance, measuring models on realistic knowledge work tasks in an agentic loop via our reference harness Stirrup. Knowledge calibration follows the same pattern: its AA-Omniscience Index of -33 is low for its intelligence level, driven by an 82% hallucination rate (Qwen3.6 27B: 49%) rather than accuracy, where it matches its peers. Agentic tool use is the exception, with Muse Glimmer scoring 24% on Tau3-Banking, ahead of Gemini 3.5 Flash-Lite (18%) and Qwen3.6 27B (17%), among the best in its class.
➤ Model details: Muse Glimmer is a 30B dense model with a 128K token context window (plus extension), released under Apache 2.0. At the time of release, Meta is not serving the model on their API; pricing and serving speed depend on third-party providers, and the weights are available on Hugging Face.

We have evaluated the initial release of Muse Glimmer at 44 on the Artificial Analysis Openness Index, a measure of model availability and transparency across model weights, training data, and other factors. This places the model at an equal position to DeepSeek V4 Flash (0731 version), GLM-5.2 and Ling 3.0 Flash, ahead of most open models. It also improves on Llama 4 Maverick's openness score, owing to its more permissive license and more extensive methodology disclosure.

At 30B parameters, Muse Glimmer sits near the Intelligence vs Parameters frontier for open weights models: 5 points above Gemma 4 31B (Reasoning) at the same size, effectively matching Kimi K2.5 (Reasoning) at 33x fewer total parameters, and just behind Qwen3.6 27B (Reasoning), the parameter-efficiency leader in this class.

Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953 Elo on GDPval-AA v2 against 1141 for Qwen3.6 27B (Reasoning), 1141 for Gemini 3.5 Flash-Lite, and 1004 for Kimi K2.5 (Reasoning), with Terminal-Bench v2.1 (52%) also behind Qwen3.6 27B (61%). The hallucination gap follows: 82% on AA-Omniscience against 49% for Qwen3.6 27B and 34% for Flash-Lite (lower is better). Its size-twin Gemma 4 31B (Reasoning) performs worse than Muse Glimmer on all of these measures. The exception is agentic tool use, with Muse Glimmer scoring strongly on Tau3-Banking (24%), ahead of Gemini 3.5 Flash-Lite (18%) and Qwen3.6 27B (17%).

Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index:

See Artificial Analysis for further details and benchmarks of Muse Glimmer: https://artificialanalysis.ai/models/muse-glimmer
同じ出来事を6媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み