SpaceXAI、最新モデル「Grok 4.6」発表し世界第3位に
本文の状態
日本語全文を表示中
詳細モードで約18分の本文を読めます。
SpaceXAI は新モデル「Grok 4.6」を発表し、人工知能分析指数で Kimi K3 を上回り GPT-5.6 Sol と同点となる世界第3位の性能を達成した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 03:06
AI深層分析
キーポイント
性能評価での躍進
Grok 4.6 は Artificial Analysis の指数で 61 を記録し、中国の Moonshot 社製 Kimi K3 を上回り、OpenAI の GPT-5.6 Sol と同点となる世界第3位の座を確立した。
コスト競争力の強化
入力 token 100 万あたり2ドル、出力 token 100 万あたり6ドルという価格設定により、同等性能の他社モデルと比較して半額以下のコストで利用可能なミッドレンジモデルとして位置づけられる。
エージェント機能への注力
本モデルは長時間稼働するエージェント、コーディング、および知識作業に特化しており、前作 Grok 4.5 High と比較してこれらのベンチマークで大幅な性能向上を遂げた。
エコシステムへの統合
Grok Build や Cursor(SpaceXAI が買収した企業)での利用が可能となり、OpenRouter や Vercel などのパートナーを通じて広く展開される予定である。
エージェント機能の強化と学習手法
Grok 4.6 はより長い作業シーケンスに集中し、自己検証を行うように設計された。これは curated な推論データやエンジニアリングデータを用いた追加トレーニングと強化学習によるものだと SpaceXAI は説明している。
重要な引用
The model scores 61 on the third-party Artificial Analysis Intelligence Index, surpassing the popular open weights Chinese model from Moonshot, Kimi K3, and tying rival OpenAI's GPT-5.6 Sol Max.
Grok 4.6 posts sizable gains over its predecessor across coding, terminal, knowledge-work and agent benchmarks while retaining an application programming interface (API) price starting at $2 per million input tokens and $6 per million output tokens.
Still, that's less than half of what GPT-5.6 Sol costs over OpenAI's API in standard mode.
SpaceXAI describes Grok 4.6 as being built specifically to stay on task across longer sequences of work, including researching unfamiliar topics, analyzing information, navigating codebases and converting product ideas into working applications.
編集コメントを表示
編集コメント
SpaceXAI は新モデルの性能と価格設定において、既存の大手競合他社に明確な対抗馬としての地位を築いた。特にエージェント機能への最適化とコスト削減は、実務での採用拡大を促す重要な要因となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Elon Musk 率いるスペースX AI(旧名:xAI)は、最新のアプローチ型AIモデル「Grok 4.6」を公開しました。このモデルは、長時間稼働するエージェント、コーディング、知識労働に焦点を当てており、これらのワークロードを実行するコストを抑えるための価格戦略も用意されています。
同モデルは、第三者の評価サイト「Artificial Analysis Intelligence Index」で61点を記録し、中国のMoonshot社が公開している人気オープンウェイトモデル「Kimi K3」を上回りました。また、競合であるOpenAIの「GPT-5.6 Sol Max」と同等の性能を達成し、自社の前世代モデル「Grok 4.5 High」よりも5ポイント向上しています。なお、1位と2位はAnthropic社の「Claude Opus 5」と「Fable 5」が占めています。
AIエージェントの評価を検討する企業にとってより重要なのは、Grok 4.6 がコーディング、ターミナル操作、知識労働、およびエージェント関連のベンチマークにおいて、前世代モデルから大幅な性能向上を達成している点です。さらに、API価格は入力100万トークンあたり2ドル、出力100万トークンあたり6ドルから開始されており、これは世界中で提供されている独自開発型とオープンソース型の主要オプションと比較しても、中価格帯のフロンティアモデルとして位置づけられます(VentureBeatによる分析)。
- モデル:入力 ($/1M) / 出力 ($/1M) / 合計 ($/1M) / ソース
- Muse Spark 1.2 Contributor:$0.10 / $0.20 / $0.30 / Meta
- MiMo-V2.5 Flash:$0.10 / $0.30 / $0.40 / Xiaomi
- deepseek-v4-flash:$0.14 / $0.28 / $0.42 / DeepSeek
- deepseek-v4-pro:$0.435 / $0.87 / $1.305 / DeepSeek
- GPT-5.6 Luna:$0.20 / $1.20 / $1.40 / OpenAI
- MiniMax-M3:$0.30 / $1.20 / $1.50 / MiniMax
- LongCat-2.0 — 期間限定プロモーション:$0.30 / $1.20 / $1.50 / LongCat
- MiMo-V2.5:$0.40 / $2.00 / $2.40 / Xiaomi
LongCat-2.0 — standard
$0.75
$2.95
$3.70
LongCat
MiMo-V2.5 Pro (≤256K)
$1.00
$3.00
$4.00
Xiaomi
Muse Spark 1.1 / 1.2
$1.25
$4.25
$5.50
Meta
GLM-5.2
$1.40
$4.40
$5.80
Z.ai
Grok 4.6 — <200K prompt tokens
$2.00
$6.00
$8.00
xAI
MiMo-V2.5 Pro (>256K)
$2.00
$6.00
$8.00
Xiaomi
Qwen3.8-Max
$2.00
$6.00
$8.00
QwenCloud
Gemini 3.6 Flash
$1.50
$7.50
$9.00
GPT-5.6 Terra
$2.00
$12.00
$14.00
OpenAI
Grok 4.6 — ≥200K prompt tokens
$4.00
$12.00
$16.00
xAI
GPT-5.4
$2.50
$15.00
$17.50
OpenAI
Kimi K3
$3.00
$15.00
$18.00
Moonshot AI
Claude Opus 5
$5.00
$25.00
$30.00
Anthropic
Sakana Fugu Ultra (≤272K)
$5.00
$30.00
$35.00
Sakana AI
GPT-5.6 Sol — Standard mode
$5.00
$30.00
$35.00
OpenAI
Claude Fable 5 / Claude Mythos 5
$10.00
$50.00
$60.00
Anthropic
GPT-5.6 Sol — Fast mode
$10.00
$60.00
$70.00
OpenAI
なお、これは OpenAI の API を通じて標準モードで GPT-5.6 Sol を利用する際の料金と比較すると、まだ半分にも満たない金額です。
SpaceXAI は、Grok 4.6 が本日より Grok Build で利用可能になったと発表しました。これは Anthropic の Claude Code や OpenAI の Codex に相当するもので、月額$30 の SuperGrok プランから利用開始できます。
また、SpaceX が最近買収した AI コーディングスタートアップの Cursor からも利用可能です。OpenRouter、Vercel、Cloudflare といったパートナー企業経由でもアクセスできます。
SpaceXAI は、Cursor と Grok Build における Grok 4.6 の利用について、初週は通常の利用枠の倍を提供しています。
SpaceXAI は今月、コード作成やエージェントタスク、知識作業を目的としたモデル「Grok 4.5」のリリースからわずか数週間後となる7月に、さらにその翌日には AI エージェントに指定されたタスクを完了させるための新たなシステム「Grok Bot」を発表した直後に、Grok 4.6 を発表しました。
今回の大きな変化は単なるベンチマークスコアの向上ではなく、エージェントの振る舞いそのものの進化にあります。
SpaceXAI は Grok 4.6 を、見知らぬトピックのリサーチや情報分析、コードベースの探索、そして製品アイデアを実働するアプリケーションへの変換など、より長い作業シーケンスにわたってタスクに集中できるよう特別に設計されたモデルと説明しています。
同社によると、Grok 4.5 よりも長期間にわたる追加トレーニングを実施しました。これには、厳選されたモデル生成による推論データや技術データを、エンジニアリングデータやオプティマイザーの改良、トレーニングレシピの変更と組み合わせて使用しています。その後、Grok 4.5 を用いて、推論レベル、エージェントハネス、STEM(科学・技術・工学・数学)、ソフトウェアエンジニアリング、知識作業にわたる教師あり微調整の軌道を再生成し、モデルベースのチェックで問題のある軌道を取り除きました。
強化学習もまた、汎用的なコーディング、知識作業、カーネル最適化、ウェブ開発、コンピュータ支援設計(CAD)など、広範なエージェント環境を対象としています。
これは重要です。なぜなら、企業の AI 導入は単なるプロンプトとレスポンスのやり取りから脱却し、状態を維持し、ツールを操作し、コードを変更し、より長い実行パス全体で問題から回復することが求められるエージェントへと移行しているからです。
SpaceXAI は、テスト期間中に Grok 4.6 がより長いタスク軌道において自己検証を強化し、処理を進める前に自身の作業を確認するようになったと発表しました。また、対話型や視覚的なプロジェクトにおける最初の試行成功率も、Grok 4.5 より向上していると報告しています。これらはあくまで同社が観察した結果であり、製品としての動作を保証する独立したデータではありませんが、SpaceXAI がモデルの学習後処理に注力した方向性を示すものです。
Grok 4.6 は最前線に到達したが、それを独占したわけではない
現在の AI モデル競争の局面において、Grok 4.5 から Grok 4.6 への進化はあまりにも重要です。
Artificial Analysis のデータによると、Grok 4.6 は GDPVal-AA v2 ベンチマークで 1,753 という Elo スコアを記録しました。これはスケジュール管理や図面作成といった実務タスクでの性能を測る指標です。対照的に、Grok 4.5 は 1,526、GPT-5.6 Sol Max は 1,728、Fable 5 Max は 1,741 です。
SpaceXAI が発表したコーディング関連の結果も、世代間の大きな進化を示していますが、最前線での競争は激化しています。
Grok 4.6 の CursorBench v3.2 スコアは 69.9% に達し、Grok 4.5 の 66.7% から向上しました。一方、Fable 5 Max は 70.5% を記録しています。DeepSWE v1.1 では、Grok が 54% から 65.9% と大幅に躍進しましたが、GPT-5.6 Sol Max が 73% で首位を維持しています。FrontierCode v1.1 Extended では、Grok が 56.6% から 61.3% に向上しましたが、GPT-5.6 Sol Max は 60.6%、Fable 5 Max は 63.6% と依然として高いスコアを示しています。
エージェントベンチマークの結果も概ね同じ傾向を示しています。Grok 4.6 は APEX-Agents で 57.5% を達成し、前世代の Grok 4.5(47.1%)から 10.4 ポイント向上しました。このスコアは GPT-5.6 Sol Max の 56.7% をわずかに上回りますが、Fable 5 Max の 59.2% には及びません。また APEX-SWE では 53.6% から 56.4% に上昇しましたが、Fable 5 Max は 58.8% を記録しています。
Terminal-Bench v3.0 では、まだ明確な差が残っていることが浮き彫りになりました。Grok 4.6 は 15.7% から 26% に改善しましたが、GPT-5.6 Sol Max と Fable 5 Max はそれぞれ 34.6%、34.1% をマークしています。
Grok 4.6 の最も目覚ましい成果は、長期にわたる専門的なタスクにおけるものです。AA-Briefcase では Elo 1,577 を記録し、Fable 5 Max(1,574)をわずかに上回り、GPT-5.6 Sol Max(1,502)を大きく引き離しました。Harvey LAB では 15.8% に達し、これは Grok 4.5 の 12.9% や Fable 5 Max の 11.3% を上回るだけでなく、GPT-5.6 Sol Max の 2.5% と比較しても圧倒的な差をつけています。
SpaceXAI は重要な注意点も示しています。同社が提示した表の他社スコアは、ベストな自己申告値または公に利用可能な結果を基にしているため、これは厳密に統制された 4 モデル間の直接比較として解釈すべきではありません。
つまり、Grok 4.6 が Grok 4.5 よりも大幅に進化したことは明確なエビデンスが示していますが、競合する最先端モデル全体に対して一貫して優位であるという証拠は乏しいと言えます。Grok 4.6 は提示された評価項目のいくつかで勝利しましたが、GPT-5.6 Sol Max と Fable 5 Max は他の分野で依然として意味のあるリードを維持しています。
コストが企業にとってより重要なベンチマークとなる可能性も
Artificial Analysis が提供した評価はもう一つの側面、つまり「投入した資金に対してモデルがどれだけの作業をこなせるか」という視点も加えています。
テスト結果によると、Grok 4.6 の「タスクあたりの知能対コスト」のパレート最適フロンティアは、1 タスクあたり約 0.84 ドルに位置しています。これは前世代の Grok 4.5 よりも割安ではなく、OpenAI の GPT-5.6 Luna、z.ai の GLM-5.2、Meta の新モデル Muse Spark 1.2 などと比較しても経済性に劣ります。
Artificial Analysis のレポートによれば、Grok 4.6 は AA-Briefcase ワークロードを平均で約 53 ターン、入力トークン数にして約 5 億回で完了しました。一方、Claude Opus 5 Max は同じタスクに約 103 ターン、20 億回の入力トークンを要しています。
これらの測定値が、すべての実運用エージェントでトークン使用量が減り、処理速度が倍になることを保証するものではありません。エージェントのコストは、ハネス設計やプロンプト、ツール呼び出し、キャッシュの有無、リトライ動作、そしてタスクそのものによって大きく左右されます。しかし、これらは「100 万トークンの生成コスト」ではなく、「ワークフロー完了にかかるコスト」という、ますます重要になる企業指標を示唆しています。この区別こそが、SpaceXAI のポジショニングの核心です。
標準的な Grok 4.6 API は、入力トークン 100 万あたり 2 ドル、出力トークン 100 万あたり 6 ドルからスタートします。また、SpaceXAI はさらに高速なバリアントも提供しており、その料金は標準版の倍額です。
提供された API ドキュメントには、長文コンテキストの展開における重要な注意点が記載されています。Grok 4.6 は 50 万トークンのコンテキストウィンドウをサポートしていますが、20 万トークン未満のプロンプトでは、入力トークンあたり 100 万トークンにつき 2 ドル、キャッシュされた入力トークンあたり同様に 0.50 ドル、出力トークンあたり 6 ドルで課金されます。一方、プロンプトが 20 万トークンに達すると、これらの料金はそれぞれ 4 ドル、1 ドル、12 ドルへと引き上げられ、そのリクエストに含まれるすべてのトークンに対して高い価格が適用されます。
つまり、総所有コスト(TCO)を試算する際、企業は 2 ドル/6 ドルという headline の料金をモデルのコンテキストウィンドウ全体に単純に当てはめてはいけません。
Artificial Analysis によると、標準的な headline の料金は、Claude Opus 5 や GPT-5.6 Sol といった競合する最先端モデルの価格と比較して依然として 60% 以上も低く設定されています。実際の節約効果は、同じワークロードを完了させる際に各モデルが消費するトークン数に依存します。
Grok という名称には、一般的な AI への懐疑論とは別に、多くの負のイメージや論争という重荷が伴っています。
性能や価格が、SpaceX AI が Grok 4.6 のベンチマークでの成果を企業導入に結びつける上で直面する唯一の障壁ではない可能性があります。Grok ブランドには、安全性やガバナンスをめぐる問題が非常に目立つ歴史があります。具体的には、過激派や反ユダヤ主義的な出力、政治的に偏った回答、Elon Musk への過度な賛美、そして最近では画像生成機能を使って同意のない性的なイメージを生成した事例などです。コンプライアンスが厳格な企業や、ブランドセーフティ、責任ある AI の要件を満たす必要がある組織にとっては、この歴史が技術的な能力とは別に調達上の懸念材料となる可能性があります。
最も悪名高いテキスト生成の出来事は 2025 年 7 月に発生しました。Grok が反ユダヤ主義的な投稿を生成し、アドルフ・ヒトラーを賛美したのです。一部の回答では、自身を「メカヒトラー」と呼称することさえありました。これを受け、SpaceX AI の前身である xAI は不適切な投稿の削除と、Grok によるヘイトスピーチの公開防止に向けた措置を講じると発表しました。
2025 年夏、Grok は無関係な質問への回答に、南アフリカにおける「白人虐殺」の存在を根拠のない主張として挿入し始めた。xAI は、この問題行動の原因が、政治的なトピックで特定の回答を生成するようシステムを誘導し、通常のレビュープロセスを迂回させる unauthorized な改変によるものだと説明した。同社は今回の変更がポリシー違反であると認め、Grok のシステムプロンプトの公開と、問題のある回答に対する 24 時間体制の監視体制の構築を約束した。南アフリカ政府は、白人市民に対する虐殺が行われているという主張を強く否定している。
同年 11 月、チャットボットがムスク氏に対して不合理に称賛的な評価を繰り返し出力したことで、Grok の客観性への疑念が再燃した。当時報じられた事例には、ムスク氏を一流のアスリートや歴史的な知的人物よりも上位に位置づける主張が含まれていた。ムスク氏は、敵対的なプロンプト(adversarial prompting)を通じて Grok が操作され、「あり得ないほど肯定的」な発言を引き出されたのだと語った。根本的な原因が何であれ、この事件は、開発者である最高経営責任者のパブリック・ペルソナとモデルの出力が混同されるリスクという、エンタープライズモデル固有の評判上の課題を浮き彫りにした。
最も深刻な論争は画像生成に関わるものです。2026年1月、英国の規制当局Ofcomは、Grokアカウントが人々の裸体画像や児童を性的に描写した画像の作成・配布に利用されているとの報告を受け、Xに対して正式調査を開始しました。Ofcomによると、調査対象となっている素材には、同意のない親密な画像の悪用(非合意型インティメートイメージの乱用)、ポルノグラフィ、児童性的虐待材料が含まれる可能性があります。これに対しXは、Grokアカウントが人々の親密な画像を作成するために利用されないよう対策を講じたと発表しましたが、Ofcomは調査を継続中と述べています。
この監視の目はOfcomだけにとどまりません。英国の情報コミッショナー庁(ICO)も、Grokの開発および導入プロセスにおいて個人データの法的扱いが適切だったか、有害な改変画像を防ぐための十分な safeguards が存在したかなどについて、XとxAIを調査しています。一方、欧州委員会もデジタルサービス法に基づき別個の正式調査を開始し、Grokに関連するシステムリスク、特に性的に露骨な改変素材の拡散に対するXの管理態勢を検査しています。
これらの調査は、新たにリリースされたGrok 4.6 APIが法律に違反しているという事実認定を目的としたものではなく、あくまでXおよびその前身であるxAI組織への関与を問うものです。しかしながら、こうした状況はSpaceXAIが企業向けにGrokを販売する上で大きな足かせとなることは間違いありません。
SpaceX は 2026 年 2 月に xAI を買収し、AI 事業部門は現在「SpaceXAI」としてブランド化されています。これにより、Grok の最新モデルは企業構造上は新しい傘下に入りますが、消費者向けのブランド名はそのまま維持されています。
今回検討した資料には、Grok 4.6 が過去の Grok デプロイメントで問題となった「メヒトラー」発言や「白人絶滅」論、性的な画像の生成、あるいはマスク氏への過度な賛美といった具体的な事例を繰り返しているという証拠は見当たりません。しかし、企業の調達チームがモデルをベンダーや製品の歴史から切り離して評価することは稀です。SpaceXAI にとってこれは、Grok 4.6 が競合する最先端モデルよりも安価で高性能であることを示すだけでなく、自社の AI サプライヤーがブランド安全性のリスク事件に巻き込まれることを許容できない組織にとって、その周囲の制御が十分に予測可能であることを証明しなければならないことを意味します。
この継続性が生む潜在的な採用上の課題は、ベンチマーク表では測ることができません。社内コーディングエージェント用にモデルを選択する開発者にとっては、価格、レイテンシ、タスク完了率が最優先事項になるかもしれません。一方で、同じモデルを顧客対応や規制の厳しい業務フローに導入する銀行、政府機関、医療提供者、あるいは消費者向けブランドは、ベンダーガバナンス、コンテンツセーフティ対策、監査可能性、そして評判へのリスクといった要素も考慮せざるを得ません。
チャットだけでなく、実際にデプロイして運用することを前提に設計されたモデル
Grok 4.6 は、提供された API 仕様書によると、テキストと画像の入力をサポートし、テキスト出力、関数呼び出し、構造化出力、推論機能を提供します。この仕様書には、1 秒あたり 150 リクエスト、1 分あたり 5,000 万トークンというレート制限も明記されており、利用可能なリージョンは us-east-1 と us-west-2 です。
Cursor の発表でも同様に、Grok 4.6 は長時間稼働するエージェントや、高度な対話・視覚処理を目的として設計されたとされています。これにより、開発者は API を取り囲む新しい環境を構築する必要なく、すでに確立されたコーディングエージェントの環境内で即座にこのモデルを利用できるようになります。
企業購入者にとって、この流通経路は別のリーダーボードの結果とほぼ同等かそれ以上に重要になる可能性があります。現在、モデル間の競争は推論スコアだけでなく、開発者が既存のコーディング、研究、運用ワークフローにモデルを組み込む際に、そのワークフローを不安定化させたり、推論コストを劇的に増大させたりしないか、また、controversial なブランドとの提携による反発を招かないかも問われています。
Grok 4.6 が圧倒的な性能優位性を確立したわけではありません。むしろ、この発表が提示するのは異なる価値提案です。それは最前線の知能、前世代からの大幅な改善、より強力な長時間稼働型エージェントの振る舞い、そして比較的積極的なトークン経済性という組み合わせです。
今後の試金石は、Artificial Analysis が制御されたエージェントワークロードで観測した効率性が、本番環境でも維持されるかどうかです。Grok 4.6 が、より少ないターン数とトークン数で長時間実行されるコーディングや知識作業を安定して完了できるのであれば、このモデルにとって最も重要なベンチマークはリーダーボード上の順位ではなく、企業における推論コストとなるでしょう。
原文を表示
Elon Musk's company SpaceXAI, formerly known as xAI, has released Grok 4.6, its latest frontier AI model, with a focus on long-running agents, coding and knowledge work — and a pricing strategy designed to make those workloads cheaper to run.
The model scores 61 on the third-party Artificial Analysis Intelligence Index, surpassing the popular open weights Chinese model from Moonshot, Kimi K3, and tying rival OpenAI's GPT-5.6 Sol Max and improving five points over Grok 4.5 High. Anthropic's Claude Opus 5 and Fable 5 occupy the number one and two spots, respectively.
More consequential for enterprises evaluating AI agents, Grok 4.6 posts sizable gains over its predecessor across coding, terminal, knowledge-work and agent benchmarks while retaining an application programming interface (API) price starting at $2 per million input tokens and $6 per million output tokens, making it a mid-priced frontier model comparing leading options that are both proprietary and open source, globally, according to VentureBeat's analysis.
Model
Input ($/1M)
Output ($/1M)
Total ($/1M)
Source
Muse Spark 1.2 Contributor
$0.10
$0.20
$0.30
Meta
MiMo-V2.5 Flash
$0.10
$0.30
$0.40
Xiaomi
deepseek-v4-flash
$0.14
$0.28
$0.42
DeepSeek
deepseek-v4-pro
$0.435
$0.87
$1.305
DeepSeek
GPT-5.6 Luna
$0.20
$1.20
$1.40
OpenAI
MiniMax-M3
$0.30
$1.20
$1.50
MiniMax
LongCat-2.0 — limited-time promo
$0.30
$1.20
$1.50
LongCat
MiMo-V2.5
$0.40
$2.00
$2.40
Xiaomi
LongCat-2.0 — standard
$0.75
$2.95
$3.70
LongCat
MiMo-V2.5 Pro (≤256K)
$1.00
$3.00
$4.00
Xiaomi
Muse Spark 1.1 / 1.2
$1.25
$4.25
$5.50
Meta
GLM-5.2
$1.40
$4.40
$5.80
Z.ai
Grok 4.6 — <200K prompt tokens
$2.00
$6.00
$8.00
xAI
MiMo-V2.5 Pro (>256K)
$2.00
$6.00
$8.00
Xiaomi
Qwen3.8-Max
$2.00
$6.00
$8.00
QwenCloud
Gemini 3.6 Flash
$1.50
$7.50
$9.00
GPT-5.6 Terra
$2.00
$12.00
$14.00
OpenAI
Grok 4.6 — ≥200K prompt tokens
$4.00
$12.00
$16.00
xAI
GPT-5.4
$2.50
$15.00
$17.50
OpenAI
Kimi K3
$3.00
$15.00
$18.00
Moonshot AI
Claude Opus 5
$5.00
$25.00
$30.00
Anthropic
Sakana Fugu Ultra (≤272K)
$5.00
$30.00
$35.00
Sakana AI
GPT-5.6 Sol — Standard mode
$5.00
$30.00
$35.00
OpenAI
Claude Fable 5 / Claude Mythos 5
$10.00
$50.00
$60.00
Anthropic
GPT-5.6 Sol — Fast mode
$10.00
$60.00
$70.00
OpenAI
Still, that's less than half of what GPT-5.6 Sol costs over OpenAI's API in standard mode.
SpaceXAI says Grok 4.6 is available today in Grok Build, SpaceXAI's answer to Anthropic's Claude Code and OpenAI's Codex, which is available starting in the $30 per month SuperGrok plan.
It's also available in SpaceX's recent acquisition of the AI coding startup Cursor, and from partners including OpenRouter, Vercel and Cloudflare.
SpaceXAI is providing twice the included usage for Grok 4.6 in Cursor and Grok Build during the first week.
The release arrives only weeks after Grok 4.5, which SpaceXAI launched in July as a model targeting coding, agentic tasks and knowledge work, and one day after the launch of Grok Bot, a new system for assigning AI agents to complete designated tasks as virtual employees.
The bigger change is agent behavior, not just another benchmark point
SpaceXAI describes Grok 4.6 as being built specifically to stay on task across longer sequences of work, including researching unfamiliar topics, analyzing information, navigating codebases and converting product ideas into working applications.
The company says it subjected the model to a longer supplemental training run than Grok 4.5, using curated model-generated reasoning and technical data alongside engineering data and changes to its optimizer and training recipe. It then used Grok 4.5 to regenerate supervised fine-tuning trajectories across reasoning levels, agent harnesses, STEM, software engineering and knowledge work, filtering problematic trajectories with model-based checks.
Reinforcement learning also targeted agentic environments spanning general coding, knowledge work, kernel optimization, web development and computer-aided design.
That matters because enterprise AI deployments are increasingly moving beyond isolated prompt-and-response interactions toward agents expected to maintain state, operate tools, modify code and recover from problems across longer execution paths.
SpaceXAI says that during its testing, Grok 4.6 showed more self-testing and verification on longer trajectories, checking its own work before proceeding. It also reports stronger first attempts on interactive and visual projects than Grok 4.5. Those are company observations rather than independent guarantees of production behavior, but they indicate where SpaceXAI concentrated the model’s post-training work.
Grok 4.6 reaches the frontier, but does not sweep it
Grok 4.6's improvement over Grok 4.5 at this juncture of the AI model competition cannot be overstated.
According to Artificial Analysis, Grok 4.6 reaches an Elo score (human preference of head-to-head model outputs, adapted from chess) of 1,753 on GDPVal-AA v2, the benchmark measuring performance on real-world tasks like scheduling and diagramming, versus 1,526 for Grok 4.5, 1,728 for GPT-5.6 Sol Max and 1,741 for Fable 5 Max.
The coding results from SpaceXAI show a similar generational improvement but more competition at the frontier.
Grok 4.6 scores 69.9% on CursorBench v3.2, up from 66.7%, while Fable 5 Max reaches 70.5%. On DeepSWE v1.1, Grok rises sharply from 54% to 65.9%, but GPT-5.6 Sol Max leads at 73%. FrontierCode v1.1 Extended moves from 56.6% to 61.3%, compared with 60.6% for GPT-5.6 Sol Max and a leading 63.6% for Fable 5 Max.
Agent benchmarks tell much the same story. Grok 4.6 reaches 57.5% on APEX-Agents, a 10.4-point increase over Grok 4.5’s 47.1%, narrowly exceeding GPT-5.6 Sol Max’s 56.7% but trailing Fable 5 Max at 59.2%. On APEX-SWE, Grok 4.6 rises to 56.4% from 53.6%, while Fable 5 Max scores 58.8%.
Terminal-Bench v3.0 exposes a larger remaining gap. Grok 4.6 improves from 15.7% to 26%, but GPT-5.6 Sol Max and Fable 5 Max score 34.6% and 34.1%, respectively.
Two of Grok 4.6’s strongest results come from longer-horizon professional work. On AA-Briefcase it scores an Elo of 1,577, narrowly exceeding Fable 5 Max’s 1,574 and topping GPT-5.6 Sol Max’s 1,502. On Harvey LAB, Grok 4.6 reaches 15.8%, versus 12.9% for Grok 4.5, 11.3% for Fable 5 Max and 2.5% for GPT-5.6 Sol Max.
SpaceXAI notes an important methodological caveat: third-party scores in its table use the best self-reported or publicly available results. The comparison therefore should not be interpreted as a perfectly controlled four-model evaluation.
In other words, the evidence supports a substantial upgrade over Grok 4.5 more clearly than it supports across-the-board superiority over rival frontier models. Grok 4.6 wins several of the displayed evaluations while GPT-5.6 Sol Max and Fable 5 Max retain meaningful leads elsewhere.
Cost could be the more important enterprise benchmark
Artificial Analysis’ supplied evaluation adds another dimension: how much work the model performs for the money spent.
The testing places Grok 4.6 on its Intelligence-versus-Cost-per-Task Pareto frontier at a reported $0.84 per task — which actually makes it less of a bargain than its predecessor, Grok 4.5, and less economical than OpenAI's GPT-5.6 Luna, z.ai's GLM-5.2, and Meta's new Muse Spark 1.2, among other models.
Artificial Analysis also reports that Grok 4.6 completed its AA-Briefcase workloads in roughly 53 turns and about 0.5 billion input tokens on average, versus approximately 103 turns and 2 billion input tokens for Claude Opus 5 Max.
Those measurements do not prove that every production agent will use fewer tokens or finish twice as quickly. Agent costs depend heavily on harness design, prompts, tool calls, caching, retry behavior and the task itself. But they point toward an increasingly important enterprise metric: the cost of completing a workflow, rather than simply the cost of generating one million tokens. That distinction is central to SpaceXAI’s positioning.
The standard Grok 4.6 API starts at $2 per million input tokens and $6 per million output tokens, and SpaceXAI also offers a faster variant at twice the price.
The supplied API documentation adds an important caveat for long-context deployments. Grok 4.6 supports a 500,000-token context window, but prompts below 200,000 tokens are billed at $2 per million input tokens, $0.50 per million cached-input tokens and $6 per million output tokens. Once a prompt reaches 200,000 tokens, those rates rise to $4, $1 and $12 respectively, with the higher pricing applying to all tokens in that request.
That means enterprises should not extrapolate the $2/$6 headline pricing across the model’s entire context window when estimating total cost of ownership.
Artificial Analysis says the standard headline rates remain more than 60% below the competing frontier-model prices it cites for Claude Opus 5 and GPT-5.6 Sol. The practical savings will depend on how many tokens each model consumes to complete the same workload.
The Grok name carries considerable baggage and controversy, separate from the general AI skepticism
Performance and price may not be the only hurdles SpaceXAI faces in converting Grok 4.6's benchmark gains into enterprise adoption. The Grok brand arrives with an unusually visible history of safety and governance controversies — including extremist and antisemitic outputs, politically skewed responses, exaggerated praise of Elon Musk and, more recently, the use of Grok's image-generation capabilities to produce non-consensual sexualized imagery. For companies with strict compliance, brand-safety or responsible-AI requirements, that history could become a procurement consideration separate from the technical capabilities of Grok 4.6 itself.
The most notorious text-generation episode came in July 2025, when Grok produced antisemitic posts, praised Adolf Hitler and in some responses referred to itself as "MechaHitler." SpaceXAI's predecessor xAI subsequently said it was removing inappropriate posts and taking steps to prevent hate speech from being published by Grok.
Also in summer 2025, Grok began inserting references to an alleged "white genocide" in South Africa into answers to unrelated questions. xAI said an unauthorized modification to Grok's response software had directed the system to produce a particular response on a political topic while bypassing its normal review process. The company said the change violated its policies and subsequently pledged to publish Grok's system prompts and establish round-the-clock monitoring for problematic responses. The South African government has rejected claims that a genocide against white South Africans is taking place.
Grok's objectivity came under scrutiny again in November of the same year after the chatbot repeatedly produced implausibly flattering assessments of Musk. Among the examples reported at the time were claims placing Musk above elite athletes and historic intellectual figures. Musk said Grok had been manipulated through adversarial prompting into making "absurdly positive" statements about him. Whatever the underlying cause, the incident illustrated the reputational problem for an enterprise model whose outputs can become entangled with the public persona of the executive most closely associated with its developer.
The most serious controversy has involved image generation. In January 2026, U.K. regulator Ofcom opened a formal investigation into X after reports that the Grok account was being used to create and distribute undressed images of people and sexualized images of children. Ofcom said the material under examination could amount to non-consensual intimate-image abuse, pornography and child sexual abuse material. X subsequently said it had implemented measures intended to stop the Grok account from being used to create intimate images of people, but Ofcom said its investigation remained open.
The scrutiny extends beyond Ofcom. Britain's Information Commissioner's Office is investigating X and xAI over both the development and deployment of Grok, including whether personal data was handled lawfully and whether adequate safeguards existed to prevent harmful manipulated imagery. The European Commission, meanwhile, opened a separate formal investigation under the Digital Services Act examining X's management of systemic risks connected to Grok, including the dissemination of manipulated sexually explicit material.
Those investigations concern X and the earlier xAI organization rather than establishing a finding that the newly released Grok 4.6 API violates those laws. Nevertheless, they are unlikely to help SpaceXAI sell Grok to businesses.
SpaceX acquired xAI in February 2026, and the AI operation now markets itself as SpaceXAI, meaning Grok's newest models sit under a different corporate structure but retain the same consumer-facing brand.
There is no evidence in the material examined here that Grok 4.6 itself repeats the specific "MechaHitler," "white genocide," sexual-image or Musk-flattery incidents associated with earlier Grok deployments. But enterprise procurement teams rarely evaluate a model in isolation from its vendor and product history. For SpaceXAI, that means Grok 4.6 may have to demonstrate not only that it is cheaper or more capable than competing frontier models, but that the controls around it are sufficiently predictable for organizations that cannot afford their AI supplier to become a brand-safety event.
That continuity creates a potential adoption problem that benchmark tables cannot measure. Developers choosing a model for an internal coding agent may care primarily about price, latency and task completion. A bank, government agency, healthcare provider or consumer brand deploying the same model into customer-facing or regulated workflows may also have to consider vendor governance, content-safety controls, auditability and reputational exposure.
A model designed to be deployed, not just chatted with
Grok 4.6 supports text and image inputs with text output, function calling, structured outputs and reasoning, according to the supplied API specifications. Those specifications also list rate limits of 150 requests per second and 50 million tokens per minute, with API availability in us-east-1 and us-west-2.
Cursor’s launch announcement similarly characterizes Grok 4.6 as designed for long-running agents and ambitious interactive and visual work, giving developers immediate access to the model inside an established coding-agent environment rather than requiring them to build a new harness around the API first.
For enterprise buyers, that distribution may matter almost as much as another leaderboard result. Models increasingly compete not just on reasoning scores but on whether developers can place them inside existing coding, research and operational workflows without destabilizing those workflows or dramatically increasing inference costs, as well as incurring any blowback from associating with a controversial brand.
Grok 4.6 does not establish an uncontested performance lead. Its launch instead presents a different proposition: frontier-level intelligence, large improvements over the previous generation, stronger long-running agent behavior and relatively aggressive token economics.
The next test will be whether the efficiency Artificial Analysis observes on controlled agentic workloads carries into production. If Grok 4.6 can consistently complete long-running coding and knowledge-work tasks with fewer turns and fewer tokens, the model’s most important benchmark may ultimately be the enterprise inference bill rather than the leaderboard.
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み