SpaceXAI、Grok 4.6 で知能フロンティア復帰しコスト効率で首位
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
SpaceXAI が発表した大規模言語モデル「Grok 4.6」は、人工知能指数で 61 を記録し GPT-5.6 Sol と並ぶ最前線に位置し、特に低コストでのエージェント性能が際立っている。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 02:07
AI深層分析
キーポイント
最高峰の知能指数達成
SpaceXAI の「Grok 4.6」は Artificial Analysis Intelligence Index で 61 を記録し、OpenAI の GPT-5.6 Sol と同等の最前線レベルに到達した。
卓越したコスト効率性
同モデルは Claude Opus 5 や GPT-5.6 Sol に比べ、1M トークンあたり約 60% 低い価格で提供され、タスクあたりのコストも最前線クラスに位置する。
強力なエージェント能力
Grok 4.6 は GDPval-AA ベンチマークで Claude Opus 5 に次ぐスコアを獲得し、複雑なタスクをより少ないターン数とトークン量で完了する効率性を示した。
エージェント性能での優位性
Grok 4.6 は静的推論よりもエージェント作業において最も優れた結果を示し、GDPval-AA v2 で Elo 1753 を記録した。
多様なタスクでの競争力
顧客サービスやターミナルベースのソフトウェアタスクなど複数の領域でトップクラスに位置し、コスト対性能のパレートフロンティアを達成している。
重要な引用
Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release
It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier
Grok 4.6's strongest results are on agentic work rather than static reasoning.
Combined with its pricing, this places Grok 4.6 on the cost vs. performance Pareto frontier for every agentic evaluation in the Intelligence Index.
編集コメントを表示
編集コメント
SpaceXAI は短期間のアップデートで知能指数を大幅に向上させ、価格競争力においても他社を凌駕する成果を示した。これは AI モデルの性能とコストのトレードオフ関係における新たな基準となる可能性が高い。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
SpaceXAI の Grok 4.6 が人工知能分析指数で 61 点を記録、GPT-5.6 Sol と並ぶ最前線へ。低コストながら突出したエージェント性能を発揮
リリースからわずか 1 ヶ月足らずで Intelligence Index でのスコアが 5 ポイント上昇し、Grok 4.3 と比較すると 23 ポイントもの差がつきました。これにより、SpaceXAI は OpenAI に次ぐ最前線の知能レベルを確立し、Anthropic の後を追う形となりました。
主なポイント:
➤ Grok 4.6 が人工知能分析指数の最前線に参入:スコアは 61 で、GPT-5.6 Sol(最大値)と同等です。Claude Opus 5(最大値 63)や Claude Fable 5(フォールバックありで最大値 62)には及びませんが、Kimi K3 よりわずかに上回っています。
➤ 強力なエージェント性能:GDPval-AA v2 Elo は 1753 を記録し、Claude Opus 5 に次ぐ順位です。Claude Fable 5 や Qwen3.8 Max とは信頼区間が重なる結果となりました。また、𝜏³-Banking では 50.7% のスコアを達成し、Qwen3.8 Max(51.3%)と並んでトップ 2 位に入りました。Terminal-Bench v2.1 では 88.4% を記録し、主要モデル群と同等の性能を示しています。
➤ 低コストでの最前線レベル知能:基本料金は Grok 4.5 と同じく、入力・出力トークン 100 万あたりそれぞれ$2/$6 で据え置かれています。これは Claude Opus 5($5/$25)や GPT-5.6 Sol($5/$30)を 60% 以上下回る価格設定です。タスクあたりのコストは$0.84 で、Kimi K3 と同等ながら知能レベルがわずかに高いことから、「知能対タスクあたりコスト」のパレート最適フロンティアに位置しています。
Grok 4.6 は、私たちが長期的なエージェント型知識作業タスクを対象に設定した独自ベンチマーク「AA-Briefcase」の Fable 5 ティアにおいて、Elo 1577 を記録しました。これは Claude Opus 5 シリーズには及ばないものの、特にターン効率に優れています。具体的には、平均で約 53 ターン、入力トークン数は約 0.5B でタスクを完了しています。一方、Claude Opus 5(最大値)では約 103 ターン、約 2.0B の入力トークンを要します。
その他のモデル詳細:
➤ コンテキストウィンドウは 500k トークンで、Grok 4.5 と同じです。
➤ 価格設定は、入出力とも 1M トークンあたり 2 ドル/6 ドル。キャッシュヒット時は 1M トークンあたり 0.5 ドルに割引されます(Grok 4.5 のキャッシュヒット料率 0.3 ドルから値上げされています)。


エージェント性能
Grok 4.6 の最も優れた結果は、静的な推論よりもエージェント型作業において示されています。現実世界のエージェント型知識作業を測る主要指標である「GDPval-AA v2」では、Elo 1753 を記録しました。これは Claude Opus 5 に次ぐ成績であり、信頼区間が重なるため統計的に有意差はないものの、Claude Fable 5 や Qwen3.8 Max と同等の性能を示しています。
この傾向はタスクの種類全体にわたって一貫しています。𝜏³-Banking(50.7%)では、ツールを活用した複数回のカスタマーサービス対応をテストし、Grok 4.6 は上位 2 位にランクインしました。また Terminal-Bench v2.1(88.4%)では、ターミナルベースのソフトウェアタスクにおいてリーダー陣と互角の結果を残しています。知識労働、カスタマーサービス、ターミナル操作のすべてで同時に競争力を持つモデルはほとんどなく、その価格設定も相まって、Grok 4.6 はインテリジェンス・インデックスにおけるすべてのエージェント評価で、コスト対性能のパレートフロンティアをリードする存在となっています。

コストパフォーマンス
最先端モデルにおいて、性能向上に伴って価格が引き上げられるのが一般的である中で、世代を超えて価格を据え置くことは異例のことです。Grok 4.6 は、$2/$6 の価格設定を変えずにインテリジェンス・インデックスで 5 ポイントの向上を実現しました。また、タスクあたりの測定コストが $0.84 であることも、この価格設定と合理的なトークン効率を反映した結果です。
購入者が最も重視すべき比較対象は、Grok 4.6 とスコア差 2 ポイント以内のモデルたちです。具体的には、$5/$25 の Claude Opus 5 や $5/$30 の GPT-5.6 Sol です。Grok 4.6 は、推論を要するワークロードにおいてコストに大きく影響する出力トークン価格が大幅に抑えられたにもかかわらず、GPT-5.6 Sol と同等のインテリジェンス・インデックススコアを提供します。

長期にわたる知識労働
Grok 4.6 は、長期にわたるエージェント型の知識作業タスクを対象とした私的ベンチマーク「AA-Briefcase」で初登場し、Elo レーティングは 1577 を記録しました。これは Fable の 5 つのティアに分類され、Claude Opus 5 シリーズには及びませんが、評価基準の適合性、プレゼンテーションの質、分析力というすべての項目で一貫して高いパフォーマンスを発揮しています。特定の分野での突出した強さが他の分野の弱さを相殺するといった偏りのないバランスが特徴です。
スコア以上に注目すべきは、その効率性です。Grok 4.6 はタスク解決に平均で約 53 ターン、入力トークン数は約 0.5B で済みます。一方、Claude Opus 5(最大値)では約 103 ターン、約 2.0B の入力トークンを必要とします。長期にわたるエージェント作業ではコンテキストが急速に蓄積するため、半分のターン数と四分の一のトークン数で同等の結果を達成できるモデルは、トークン単価の問題を超えて、明確なコスト優位性を持っています。

詳細結果
Artificial Analysis Intelligence Index における個別評価の完全な内訳は以下の通りです。

Grok 4.6 の詳細やベンチマークデータについては、Artificial Analysis をご参照ください:https://artificialanalysis.ai/models/grok-4-6
原文を表示
SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost
Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic.
要点
➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3
➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on 𝜏³-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models
➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier
➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max)
Other model details:
➤ Context window of 500k tokens (unchanged from Grok 4.5)
➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5’s $0.3 per 1M tokens for cache hits


Agentic performance
Grok 4.6's strongest results are on agentic work rather than static reasoning. On GDPval-AA v2, our leading measure of real-world agentic knowledge work, it scores an Elo of 1753 - behind only Claude Opus 5, and statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max given overlapping confidence intervals.
The pattern holds across task types. 𝜏³-Banking (50.7%) tests multi-turn customer service with tool use and places Grok 4.6 in the top two, while Terminal-Bench v2.1 (88.4%) puts it level with the leaders on terminal-based software tasks. Few models are simultaneously competitive across knowledge work, customer service and terminal use; combined with its pricing, this places Grok 4.6 on the cost vs. performance Pareto frontier for every agentic evaluation in the Intelligence Index.

Cost
Holding headline pricing flat across a generation is unusual at the frontier, where intelligence gains have typically been accompanied by price increases. Grok 4.6 delivers a 5-point Intelligence Index gain at unchanged $2/$6 pricing, and our measured cost per task of $0.84 reflects both that pricing and reasonable token efficiency.
The comparison that matters for buyers is against the models scoring within two points of it: Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30. Grok 4.6 offers effectively the same Intelligence Index score as GPT-5.6 Sol at a fraction of the output token price, which is the dimension that dominates cost in reasoning-heavy workloads.

Long-horizon knowledge work
Grok 4.6 debuts on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577. This places it at Fable 5-tier, behind the Claude Opus 5 family, with consistently strong performance across rubric grading, presentation quality and analytical quality rather than strength in one dimension offsetting weakness in another.
The efficiency profile is as notable as the score. Grok 4.6 resolves tasks in ~53 turns and ~0.5B input tokens on average, against ~103 turns and ~2.0B input tokens for Claude Opus 5 (max). Long-horizon agentic work accumulates context rapidly, so a model that reaches a comparable answer in half the turns and a quarter of the input tokens has a cost advantage well beyond its per-token pricing.

Full results
Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index:

See Artificial Analysis for further details and benchmarks of Grok 4.6: https://artificialanalysis.ai/models/grok-4-6
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み