xAI、長時間実行型エージェントに注力した「Grok 4.6」を公開
本文の状態
日本語全文を表示中
詳細モードで約6分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
x.ai は長期的エージェント実行と複雑な対話・視覚作業に焦点を当てた新モデル「Grok 4.6」を発表し、人工知能分析インテリジェンス指数で GPT-5.6 Sol と同等の性能を示した。
AI深層分析を開く2026年8月14日 22:19
AI深層分析
キーポイント
Grok 4.6 の主要機能強化
新モデルは複雑なタスクを多段階で実行する能力に重点を置き、リサーチ、情報分析、コードベース全体での作業、アイデアの実装などに対応する。
ベンチマークにおける同等性能
x.ai によると、Grok 4.6 は人工知能分析インテリジェンス指数(9 つのベンチマークの合成スコア)において GPT-5.6 Sol と同等の最前線知能を達成した。
エージェント機能への注力
本リリースは単なるチャットを超え、長期間にわたる自律的なエージェント動作と、より野心的な対話型・視覚的作業の実行能力を強化している。
Grok 4.6 の知能指数での実績
Grok 4.6 はエージェント型コーディングと知識作業のベンチマークで先端的な知能を達成した。
Artificial Analysis Intelligence Index での同等評価
9 つのベンチマークからなる合成スコアである Artificial Analysis Intelligence Index で、Grok 4.6 は GPT-5.6 Sol と同点となった。
重要な引用
Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.
It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.
Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks.
It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks.
編集コメントを表示
編集コメント
Grok 4.6 の発表は、単なるバージョンアップではなく、複雑なタスクを自律的に処理するエージェント機能への明確な転換を示している。同モデルが主要ベンチマークで競合他社と同等の成績を収めたことは、実務環境での採用可能性を高める重要な材料となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本日、Grok 4.6のリリースを発表します。Grok 4.6 はGrok 4.5を基盤としつつ、特に長時間稼働するエージェントや、より野心的な対話型・視覚的なタスクに注力しています。トピックの調査から情報分析、コードベース全体での作業、アイデアの実装による完成度の高いアプリケーションや成果物の作成に至るまで、多段階にわたる複雑なタスクにも対応可能です。
Grok 4.6 は、エージェントによるコーディングおよび知識労働に関する複数のベンチマークで最先端の知能を実現しました。人工分析インテリジェンス指数(Artificial Analysis Intelligence Index)では GPT-5.6 Sol と同等のスコアを記録しています。この指数は 9 つのベンチマークを組み合わせた総合評価です。
0204060AA インテリジェンス指数 62Fable 5 Max 61 Grok 4.6 61 GPT-5.6 Sol Max 56 Grok 4.5 High
競合モデルの数値は、各開発者が公開したシステムカードまたはベンチマークリーダーボードから引用されています。
*Grok 4.6 と他社主要モデルを AA インテリジェンス、GDPVal-AA、DeepSWE 1.1、CursorBench 3.2、FrontierCode 1.1 の各項目で比較したベンチマークの棒グラフ。競合モデルの数値は、各開発者が公開したシステムカードまたはベンチマークリーダーボードから引用されています。
Grok 4.6 は本日、CursorおよびGrok Buildで利用可能です。初週は両プラットフォームで利用可能量が 2 倍になる特典を用意していますので、すぐに Grok 4.6 の試作を開始できます。
Grok 4.6 のトレーニングについて
Grok 4.6 は、推論や高度な技術概念に関する精選されたモデル生成データ、高品質なエンジニアリングデータを活用し、最適化アルゴリズムとトレーニングレシピを改善した上で、Grok 4.5 よりも長い追加学習期間を経ています。これにより、その後に続く SFT(Supervised Fine-Tuning)や RL(Reinforcement Learning)の段階に向けた、より堅牢な基盤が構築されました。
その後、推論プロセス、エージェント・ハーンセス、STEM 分野、ソフトウェアエンジニアリング、知識労働など多岐にわたる領域において、Grok 4.5 を活用して SFT トラジェクトリを再生成しました。モデルベースのチェックで問題のあるトレースを除外した結果、強力なパフォーマンスと改善された振る舞いを示す SFT チェックポイントが得られました。
Grok 4.6 は、知識労働や汎用コーディングに加え、カーネル最適化、ウェブ開発、CAD(Computer-Aided Design)など特定のドメインに特化した環境における幅広いエージェント RL タスクでトレーニングされています。
大胆なアイデアを実働プロジェクトへ
Grok 4.6 をテストした際、その能力の限界を試し、多数の手順にわたって作業を継続させることを目的としたプロジェクトが用意されました。その結果、広範な製品アイデアを実働可能な最初のバージョンへと具体化する点で、このモデルは特に優れていることが判明しました。未知のドメインに関する調査からアプリケーションの構造化、コアインタラクションの実装までを行い、フィードバックを数回繰り返しながら結果を洗練させていくことができます。
より長いトジェクトリにおいては、次のステップに進む前に自ら作業を検証する「自己テスト」や「検証」の動きも目立つようになりました。
Grok 4.6 は、視覚的・対話型のプロジェクトにおいて、従来の Grok 4.5 よりもはるかに質の高い初期出力を生成します。具体的な製品アイデアを与えられれば、アプリケーションの構造やビジュアル言語を一発で確立できるため、まずは実用的な基盤を作り、その上でループ内で反復改善していくアプローチが最も効率的なプロジェクトにおいて特に威力を発揮します。
安全性と機能
Grok 4.6 の安全対策は、モデルの能力向上に合わせて強化・調整されました。
当社のセキュリティスタックは、正当な利用ケースにおける利便性と安全性を最大化するように設計されており、脆弱性の修正やエンジニアリング設計サイクルの加速、AI 研究の支援といった領域で、Grok 4.6 が有用かつ安全に機能することを保証します。
Grok 4.6 の能力拡大に対応した評価作業では、これまでにない広範な事前展開テスト(能力と安全対策の調整)に加え、展開後の継続的な監視や第三者による検証も徹底して実施しています。
ベンチマーク結果
- テスト項目:Grok 4.6 High / Grok 4.5 High / GPT-5.6 Sol Max / Fable 5 Max
- AA Intelligence Index:61 / 56 / 61 / 62
- GDPVal-AA v2:1753 / 1526 / 1728 / 1741
- CursorBench v3.2:69.9% / 66.7% / 67.2% / 70.5%
- DeepSWE v1.1:65.9% / 54% / 73% / 70%
- FrontierCode v1.1 (Extended):61.3% / 56.6% / 60.6% / 63.6%
- APEX-Agents:57.5% / 47.1% / 56.7% / 59.2%
- Terminal-Bench v3.0:26% / 15.7% / 34.6% / 34.1%
- APEX-SWE:56.4% / 53.6% / — / 58.8%
- AA-Briefcase:1577 / 1313 / 1502 / 1574
Harvey LAB (Vals)
15.8%
12.9%
2.5%
11.3%
各評価項目での最高スコアを太字で示しています。サードパーティ製モデルのスコアは、自己報告値または公に利用可能な結果のうち最も高いものを採用しています。
Grok 4.6 の使い方
Grok 4.6 は本日、Cursor と Grok Build で利用可能になりました。また、API を通じてや、OpenRouter、Vercel、Cloudflare などのパートナー企業でもご利用いただけます。
料金は入力トークン 100 万あたり 2 ドル、出力トークン 100 万あたり 6 ドルからスタートします。さらに、高速版も用意されており、こちらは通常版の倍額です。
初週は Grok Build と Cursor で利用可能な使用量が 2 倍になりますので、すぐに Grok 4.6 の試作を開始できます。
API キーの取得
SpaceXAI API を通じて、今日から Grok 4.6 の開発を始めましょう。
API ドキュメント
ドキュメントを確認し、Grok 4.6 をご自身のスタックに統合してください。
Grok Build で無料で試す
x.ai/build から今日すぐにご利用ください。
原文を表示
Today we are releasing Grok 4.6. Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.
Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks.
0204060AA Intelligence Index62Fable 5 Max61Grok 4.661GPT-5.6 Sol Max56Grok 4.5 HighCompetitor figures are drawn from the respective developers’ published system cards or benchmark leaderboards
*Benchmark bar charts comparing Grok 4.6 with other leading models across AA Intelligence, GDPVal-AA, DeepSWE 1.1, CursorBench 3.2, and FrontierCode 1.1. Competitor figures are drawn from the respective developers’ published system cards or benchmark leaderboards.*
Grok 4.6 is available today in Cursor and Grok Build. We’re offering 2x included usage inside Grok Build and Cursor for the first week so you can start trying 4.6 immediately.
Training Grok 4.6
Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. This produced a stronger foundation for the SFT and RL stages that followed.
We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work, and filtered out problematic traces with model-based checks. The resulting SFT checkpoint shows strong performance and improved behavior.
Grok 4.6 is trained on a wide range of agentic RL tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, computer-aided design, and more.
Turning ambitious ideas into working projects
We tested Grok 4.6 on projects designed to stretch its range and ability to sustain work over many steps. We found the model is especially strong at turning a broad product idea into a working first version. It can research unfamiliar domains, structure the application, implement the core interactions, and continue refining the result through several rounds of feedback.
On longer trajectories, we also started to see more self-testing and verification, with the model checking its own work before moving on.
Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. This has made it especially useful for projects where the fastest route to a good result was to begin with something substantial and then iterate in the loop.
Safety and capabilities
Grok 4.6’s safeguards have been improved and calibrated in line with the model’s capabilities.
Our safety stack is designed to maximize utility and security across legitimate use cases, allowing Grok 4.6 to be helpful and safe in domains such as vulnerability patching, accelerating the engineering design cycle, and augmenting AI research.
Our safeguard evaluation work reflects Grok 4.6’s expanded capabilities, with our widest-ever suite of pre-deployment testing for capabilities and safeguard calibration, as well as extensive post-deployment and third-party testing.
Evals
Grok 4.6 High
Grok 4.5 High
GPT-5.6 Sol Max
Fable 5 Max
AA Intelligence Index
61
56
61
62
GDPVal-AA v2
1753
1526
1728
1741
CursorBench v3.2
69.9%
66.7%
67.2%
70.5%
DeepSWE v1.1
65.9%
54%
73%
70%
FrontierCode v1.1 (Extended)
61.3%
56.6%
60.6%
63.6%
APEX-Agents
57.5%
47.1%
56.7%
59.2%
Terminal-Bench v3.0
26%
15.7%
34.6%
34.1%
APEX-SWE
56.4%
53.6%
—
58.8%
AA-Briefcase
1577
1313
1502
1574
Harvey LAB (Vals)
15.8%
12.9%
2.5%
11.3%
Best score per evaluation in bold. Third-party model scores are the best of self-reported or publicly available results.
Get started with Grok 4.6
Grok 4.6 is available today in Cursor and Grok Build. It’s also available in the API and other partners like OpenRouter, Vercel, and Cloudflare.
Pricing starts at $2 per million input tokens and $6 per million output tokens. Additionally, there is a fast variant which is twice the price.
We’re offering 2x included usage inside Grok Build and Cursor for the first week so you can start trying 4.6 immediately.
Create an API Key
Start building with Grok 4.6 today via the SpaceXAI API.
API Docs
Read the docs and integrate Grok 4.6 into your stack.
Try it in Grok Build for free
Get started today at x.ai/build.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み