SpaceXAI、長期エージェント向け「Grok 4.6」を公開し知能指数で GPT-5 に並ぶ
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
SpaceXAI は長期的なエージェント実行やコーディングに最適化された「Grok 4.6」をリリースし、50 万トークンのコンテキストと強化学習による自己検証機能を搭載して実用性を高めた。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 15:20
AI深層分析
キーポイント
技術的進化の方向性
ベースモデルの規模拡大ではなく、補完的なトレーニング期間の延長と再帰的な教師あり微調整、およびエージェント環境での強化学習によって性能を向上させた。
主要な機能強化
50 万トークンのコンテキストウィンドウを採用し、複雑なタスクにおけるドリフト防止と自己検証機能を強化したことで、長期的なエージェント実行が可能になった。
市場での利用状況
xAI API や Cursor、Grok Build などで提供されているが、オープンウェイト版はなく、オンプレミスやエアギャップ環境での展開は不可能である。
500K コンテキストと推論レベルの拡張
Grok 4.6 は 50 万トークンのコンテキスト長をサポートし、推論効率に「xhigh」レベルが新たに追加された。これはベースモデルの拡大ではなく、Grok 4.5 のポストトレーニングアップグレードである。
ベンチマークでの性能と価格設定
知能指数では GPT-5.6 Sol Max と同点だが、コーディング系タスクでは他モデルに劣る結果となった。トークン数 20 万を超えると料金が倍増する構造であり、キャッシュ利用には特定のヘッダー設定が必須である。
重要な引用
SpaceXAI held the foundation constant and spent the improvement on a longer supplemental training run
Agents that stay on a task across many steps without drifting
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, up five points from Grok 4.5
"Grok 4.6 is a post-training upgrade on Grok 4.5, not a bigger base model."
編集コメントを表示
編集コメント
モデルの規模拡大に頼らないトレーニング手法の最適化による性能向上は、現在の AI 業界における重要なトレンドを示している。ただし、オープンウェイト版が提供されない点は、セキュリティ要件の高い組織にとって導入判断の分岐点となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
SpaceXAI が Grok 4.6 をリリースしました。これは基盤モデルを大型化するのではなく、Grok 4.5 の後処理段階でのアップグレードです。SpaceXAI は基盤となるモデルの構造は据え置き、その分をより長い追加学習の実行、再生成された教師あり微調整(SFT)の軌道、そしてエージェント環境における強化学習に注ぎ込みました。これにより、タスクから逸脱することなく多数の手順にわたって安定して動作するエージェントが実現します。
Grok 4.6 は Artificial Analysis Intelligence Index で 61 を記録し、Grok 4.5 より 5 ポイント向上し、GPT-5.6 Sol Max と同点となりました。このモデルは 50 万トークンのコンテキスト長をサポートし、今日から Cursor や Grok Build で利用可能です。また、Grok 4.5 が提供していたラダー(階層)よりも上位に位置する新しい推論努力レベル「xhigh」が追加されました。
実運用への導入は可能でしょうか?
はい、限定的なワークロードであれば本番環境でのデプロイが可能です。モデルは xAI API 経由で grok-4.6 として一般提供されており、Grok Build のデフォルトモデルとなっています。また、すべてのプランで Cursor に搭載され、OpenRouter、Vercel、Cloudflare を介してルーティングも可能です。ただし、オープンウェイト版のリリースやセルフホスティングの経路は用意されていないため、エアギャップ(物理的に隔離された)環境での展開はできません。
企業の導入段階としては、Seed 期のチームやインディーズ開発者は、Harness の構築が不要な Cursor や Grok Build を通じて即座に採用できます。中堅企業向けのエンジニアリング組織にとって最も適しています。API 統合のみで対応可能であり、mTLS 認証、バッチ処理、優先度付き処理のドキュメントも整備されています。規制の厳しい大企業は、まずパイロット導入を検討すべきです。ベンダーのブランド歴については、いくつかの購買委員会において現在も実質的な調達課題となっています。
対象業界は、ソフトウェアおよび開発者向けツール、半導体およびカーネルエンジニアリング、ハードウェアおよび CAD に隣接する設計、金融研究、そして法務分析です。トレーニングの構成はこれらの分野のうちいくつかを明確に狙って設計されました。
適用事例としては、リポジトリ全体のリファクタリング、移行エージェント、50 万トークン規模のコーパスを対象とした調査・統合パイプライン、製品概要からの第一回アプリケーションスケフォールディング、GPU カーネル最適化、そして文書中心の知識作業が挙げられます。
実際に何が変わったのか
Grok 4.6 はベースモデルを大型化したものではありません。SpaceXAI が説明するのは、推論や高度な技術概念、高品質なエンジニアリングデータを用いた、Grok 4.5 よりも長い追加トレーニングの実施です。また、最適化手法とトレーニングレシピも改善されました。
Grok 4.5 を用いて、推論の努力レベル、エージェントハネス、STEM・ソフトウェアエンジニアリング・知識作業にわたるドメイン全体における教師あり微調整の軌道を再生成しました。問題のある痕跡はモデルベースのチェックによってフィルタリングされています。その後は、知識作業、一般的なコーディング、ウェブ開発、CAD、カーネル最適化をカバーするエージェント環境において強化学習が実施されました。
行動に関する重要な洞察として、SpaceXAI はより長い軌道において、自己テストと検証が増加していることを報告しています。これはモデル自身が次のステップに進む前に自分の作業をチェックする様子です。ただし、これは社内テストに基づくベンダーの観察であり、独立して測定された結果ではありません。
このモデルは 50 万トークンのコンテキスト長をサポートし、テキストと画像を入力として受け付けますが、出力はテキストのみとなります。出力文字数に制限はなく、知識の更新日付は 2026 年 2 月 1 日です。
推論の強度(reasoning_effort)には、低・中・高(デフォルト)に加え、新たに「xhigh」レベルが追加されました。SpaceXAI は Grok 4.6 のパラメータ数については公表していません。
ベンチマーク結果は、まず損失値を確認してください
xAI が発表した表によると、Grok 4.6(High モード)の Artificial Analysis Intelligence Index スコアは 61 で、Grok 4.5 の 56 を上回り、GPT-5.6 Sol Max と同点です。また、GDPval-AA v2 では 1753 Elo(Grok 4.5 は 1526)、AA-Briefcase では 1577(同 1313)、Harvey LAB でもトップを記録しています。
しかし、エンジニアリングチームにとって最も重要なコーディング関連の項目では、他社モデルに劣っています。DeepSWE v1.1 のスコアは 65.9% で、世代間で 11.9 ポイント向上しましたが、GPT-5.6 Sol Max の 73% には及びません。Terminal-Bench v3.0 は 26%(Grok 4.5 の 15.7% よりほぼ倍)で、リストされた 4 つのモデルの中では最下位です。CursorBench v3.2 は 69.9%、FrontierCode v1.1 Extended は 61.3%、APEX-Agents は 57.5% です。
評価にあたっては以下の 2 点に注意が必要です。第一に、GDPval-AA v2 と AA-Briefcase で太字で強調された勝利は、Artificial Analysis が公表した信頼区間内に含まれており、統計的な同点であり、明確なリードではありません。第二に、比較対象には現在このインデックスでトップを走る Anthropic の Claude Opus 5 は含まれていません。開示された損失値の方が、より信頼できる指標と言えます。
価格とアクセス方法
リリースノートによると、Grok 4.6 の料金は、20 万トークン以下のプロンプトの場合、入力 1M トークンあたり 2 ドル、キャッシュ済み入力 1M トークンあたり 0.50 ドル、出力 1M トークンあたり 6 ドルです。20 万トークンを超過する場合はそれぞれ 4 ドル、1 ドル、12 ドルとなります。
また、発表ページには、価格が倍の高速版も存在すると記載されていますが、個別のモデル ID は公開されていません。Grok Build と Cursor では、初週に利用枠を 2 倍にするキャンペーンを実施しています。
チームでは、prompt_cache_key を設定するか、Chat Completions API で x-grok-conv-id ヘッダーを追加してください。これらを指定しないとリクエストがサーバー間で分散し、キャッシュヒット率が不安定になるため、入力料金のフル価格が適用されてしまいます。
インタラクティブな解説
主なポイント
- Grok 4.6 は、より大きなベースモデルの一新ではなく、Grok 4.5 の学習後の改良版です。
- 文脈長は 50 万トークンに対応し、テキストと画像の入力が可能。また、推論努力レベルとして新しい「xhigh」が追加されました。
- AA Intelligence Index では GPT-5.6 Sol Max と同点の 61 を記録しましたが、DeepSWE や Terminal-Bench の評価ではやや劣ります。
- 基本料金は 2 ドル/6 ドルのままですが、プロンプトが 20 万トークンを超過すると料金が倍になります。
- API、Cursor、Grok Build で本日利用可能。オープンウェイトやセルフホスティングは提供されません。
詳細な技術情報はこちらをご覧ください。また、Twitter のフォローもぜひお願いします。15 万人以上の ML 研究者が集まる SubReddit への参加や、ニュースレターの購読もお忘れなく!
Telegram ユーザーの方へ:今なら Telegram でもコミュニティに参加できます。
SpaceXAI が「Grok 4.6」をリリース。50 万トークンのコンテキスト長を備え、長時間稼働するエージェント、コーディング、知識処理に最適化されたフロンティアモデルです。
この発表は、MarkTechPost で最初に紹介されました。
原文を表示
SpaceXAI just released Grok 4.6. The release is a post-training upgrade over Grok 4.5 rather than a larger base model. SpaceXAI held the foundation constant and spent the improvement on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning in agentic environments. Agents that stay on a task across many steps without drifting. Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, up five points from Grok 4.5 and tied with GPT-5.6 Sol Max. The model takes 500,000 context tokens, is live today in Cursor and Grok Build, and adds a new xhigh reasoning-effort level above the ladder Grok 4.5 shipped with.
Is it deployable?
Yes, in production, with a bounded set of workloads. The model is generally available through the xAI API as grok-4.6, is the default model in Grok Build, ships in Cursor on all plans, and is routable via OpenRouter, Vercel, and Cloudflare. There is no open-weights release and no self-hosting path, so air-gapped deployments are out.
Company stage: Seed-stage teams and indie developers can adopt it immediately, since Cursor and Grok Build need no harness work. Mid-market engineering orgs are the strongest fit: API-only integration, mTLS authentication, batch and priority processing are documented. Regulated enterprises should stage a pilot first — the vendor’s brand history is a live procurement question in several buying committees.
Industries: Software and developer tooling, semiconductor and kernel engineering, hardware and CAD-adjacent design, financial research, and legal analysis. The training mix explicitly targeted several of these.
Applications: Repository-wide refactors, migration agents, research-and-synthesis pipelines over 500K-token corpora, first-pass application scaffolding from a product brief, GPU kernel optimization, and document-heavy knowledge work.
What actually changed
Grok 4.6 is not a larger base model. SpaceXAI describes a longer supplemental training run than Grok 4.5 received, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe.
Grok 4.5 was then used to regenerate supervised fine-tuning trajectories across reasoning-effort levels, agent harnesses, and domains spanning STEM, software engineering, and knowledge work, with problematic traces filtered by model-based checks. Reinforcement learning followed in agentic environments covering knowledge work, general coding, web development, computer-aided design, and kernel optimization.
The behavioral insight is an important one: on longer trajectories, SpaceXAI reports more self-testing and verification, with the model checking its own work before moving on. That is a vendor observation from internal testing, not an independently measured result.
The model takes 500,000 context tokens, accepts text and image input with text-only output, has no stated text output limit, and carries a February 1, 2026 knowledge cutoff. reasoning_effort now supports low, medium, high (default), and a new xhigh level. SpaceXAI did not publish a parameter count for Grok 4.6.
Benchmarks: read the losses first
On xAI’s launch table, Grok 4.6 (High) scores 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 and tied with GPT-5.6 Sol Max. It leads the table on GDPval-AA v2 (1753 Elo, versus 1526 for Grok 4.5), AA-Briefcase (1577, versus 1313), and Harvey LAB.
It trails on the coding rows that matter most to engineering teams. DeepSWE v1.1 lands at 65.9%, up 11.9 points generationally but behind GPT-5.6 Sol Max at 73%. Terminal-Bench v3.0 reaches 26%, nearly double Grok 4.5’s 15.7% and still last of the four listed models. CursorBench v3.2 is 69.9%, FrontierCode v1.1 Extended is 61.3%, and APEX-Agents is 57.5%.
Two things to note while evaluating. First, the table’s bolded wins on GDPval-AA v2 and AA-Briefcase sit inside Artificial Analysis‘ published confidence intervals — they are statistical ties, not leads. Second, the comparison set excludes Anthropic’s Claude Opus 5, which currently tops that index. The disclosed losses are the more reliable signal.
Pricing and access
Per the release notes, Grok 4.6 bills $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200K prompt tokens, and $4 / $1 / $12 above that threshold. The launch page also references a faster variant at double the price, with no separate model ID published. Grok Build and Cursor are offering 2× included usage for the first week.
Teams should set a prompt_cache_key (or the x-grok-conv-id header on Chat Completions). Without it, requests scatter across servers and cache hits become unreliable, so full input price applies.
Interactive explainer
Key Takeaways
Grok 4.6 is a post-training upgrade on Grok 4.5, not a bigger base model.
500K context, text and image input, and a new xhigh reasoning-effort level.
Ties GPT-5.6 Sol Max at 61 on the AA Intelligence Index; trails on DeepSWE and Terminal-Bench.
Same $2 / $6 headline price, with rates doubling above 200K prompt tokens.
Available today via API, Cursor, and Grok Build — no open weights, no self-hosting.
Check out the FULL TECHNICAL DETAILS here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work appeared first on MarkTechPost.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み