Cursor と SpaceXAI、長期間実行型エージェントに特化した「Grok 4.6」を共同発表
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Cursor Research
Cursor と SpaceXAI は共同で、長期的なエージェントタスクや複雑なコードベース処理に特化した次世代モデル「Grok 4.6」をリリースした。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 01:33
AI深層分析
キーポイント
長期的エージェント機能の強化
Grok 4.6 は、トピック調査からコードベース分析、アイデアの実装まで、多段階にわたる複雑なタスクを継続的に実行する能力に重点を置いている。
ベンチマークでの最高水準性能
同モデルは人工知能分析インテリジェンス指数において GPT-5.6 Sol と同等のスコアを獲得し、コーディングおよび知識作業の最先端指標を達成した。
高度なトレーニング手法の採用
推論と高度な技術概念のための curated モデル生成データや高品質エンジニアリングデータを活用し、SFT および RL 段階の基盤を強化して訓練された。
即座の利用開始と利用制限
Grok 4.6 は当日より Cursor と Grok Build で利用可能となり、初週は両プラットフォームで利用量が倍増する特典が提供される。
広範なアイデアから実装までの一貫した処理能力
Grok 4.6 は製品の広範なアイデアを調査し、構造を構築し、コアインタラクションを実装してフィードバックを通じて反復するまでを一貫して行う。
重要な引用
Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work.
It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks.
It can research unfamiliar domains, structure the application, implement the core interactions, and continue refining the result through several rounds of feedback.
Grok 4.6's safeguards have been improved and calibrated in line with the model's capabilities.
編集コメントを表示
編集コメント
Grok 4.6 のリリースは、単なるバージョンアップを超え、複雑なタスクを自律的に完遂する「エージェント」としての AI の成熟度を示す重要なマイルストーンである。特に開発ツールとの統合が強化されたことで、実務での活用範囲がさらに広がる可能性が高い。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
SpaceXAI と共同で、Grok 4.6 をリリースしました。
Grok 4.6 は Grok 4.5 を基盤としつつ、長時間稼働するエージェントや、より野心的な対話・視覚処理タスクに重点を置いています。トピックの調査から情報分析、コードベース全体での作業、アイデアを磨き上げたアプリケーションや成果物への変換まで、複雑で多段階にわたるタスクにも安定して対応します。
Grok 4.6 は、コーディングと知識労働に関する複数のエージェントベンチマークで最先端の知能を実現しました。9 つのベンチマークを組み合わせた Artificial Analysis Intelligence Index では、GPT-5.6 Sol と同等のスコアを記録しています。

Grok 4.6 は本日、Cursor と Grok Build で利用可能です。初週は Cursor および Grok Build 内で、通常使用量の 2 倍を利用できる特典をご用意しています。
Grok 4.6 のトレーニング
Grok 4.6 は、推論や高度な技術概念に関する厳選されたモデル生成データ、高品質なエンジニアリングデータ、そして改良されたオプティマイザーと学習レシピを用いて、Grok 4.5 よりも長い追加トレーニングを実施しました。これにより、その後に続く SFT(Supervised Fine-Tuning)および RL(Reinforcement Learning)の段階に対するより強固な基盤が築かれました。
その後、Grok 4.5 を用いて、推論の取り組みやエージェントハッチ、STEM(科学・技術・工学・数学)、ソフトウェアエンジニアリング、知識労働などのドメインにわたる SFT(教師あり微調整)トラジェクトリを再生成しました。モデルベースのチェックで問題のあるトレースは除外し、その結果得られた SFT チェックポイントは高い性能と改善された振る舞いを示しています。
Grok 4.6 は、知識労働や一般的なコーディングに加え、カーネル最適化、ウェブ開発、CAD(Computer-Aided Design)、その他ドメイン固有の環境など、幅広いエージェント RL タスクを学習対象としています。
大胆なアイデアを実働プロジェクトへ
Grok 4.6 をテストした際、その能力範囲を広げ、多数のステップにわたって作業を継続する能力を試すよう設計されたプロジェクトで評価を行いました。その結果、同モデルは広範な製品アイデアを具体的な初版へと落とし込む点で特に優れていることがわかりました。見知らぬドメインのリサーチからアプリケーションの構造設計、コアインタラクションの実装まで行い、フィードバックを数回繰り返しながら結果を洗練させていくことができます。
より長いトラジェクトリにおいては、モデルが次のステップに進む前に自ら作業を検証・テストする傾向も確認されました。
Grok 4.6 は、視覚的かつインタラクティブなプロジェクトにおいて、Grok 4.5 で通常見られたものよりもはるかに強力な最初の出力を生成します。具体的な製品アイデアが与えられれば、アプリケーションの構造とビジュアルランゲージを一発で確立できるのです。この特性により、「まずは実りのある基盤を作り、その上でループ内で反復する」というアプローチが最も効率的なプロジェクトにおいて、特に有用なツールとなっています。
セーフティと能力
Grok 4.6 のセーフティ機能は、モデルの能力に合わせて改善・調整されました。
当社のセキュリティスタックは、正当な利用ケースにおいて利便性と安全性を最大化するように設計されており、脆弱性の修正やエンジニアリング設計サイクルの加速、AI 研究の支援など、さまざまな領域で Grok 4.6 が有用かつ安全に活用できることを保証します。
Grok 4.6 のセーフティ評価は、その拡大した能力を反映しており、機能とセーフティ調整のための事前展開テストを従来最大規模で実施するとともに、広範な事後展開における第三者による検証も行っています。
Grok 4.6 の利用開始
Grok 4.6 は本日より、Cursor および Grok Build で利用可能です。また、SpaceXAI API を通じてや、OpenRouter、Vercel、Cloudflare などのパートナー経由でもご利用いただけます。
料金は、入力トークン 100 万あたり 2 ドル、出力トークン 100 万あたり 6 ドルからとなっています。高速版は価格が倍になります。
初週に限り、Cursor と Grok Build では利用回数が 2 倍になります。
原文を表示
Today we are releasing Grok 4.6 together with SpaceXAI.
Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.
Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks.

Grok 4.6 is available today in Cursor and Grok Build. We’re offering 2x included usage inside Cursor and Grok Build for the first week.
Training Grok 4.6
Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. This produced a stronger foundation for the SFT and RL stages that followed.
We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work. We filtered out problematic traces with model-based checks. The resulting SFT checkpoint shows strong performance and improved behavior.
Grok 4.6 is trained on a wide range of agentic RL tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, computer-aided design, and more.
Turning ambitious ideas into working projects
We tested Grok 4.6 on projects designed to stretch its range and ability to sustain work over many steps. We found the model is especially strong at turning a broad product idea into a working first version. It can research unfamiliar domains, structure the application, implement the core interactions, and continue refining the result through several rounds of feedback.
On longer trajectories, we also started to see more self-testing and verification, with the model checking its own work before moving on.
Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. This has made it especially useful for projects where the fastest route to a good result was to begin with something substantial and then iterate in the loop.
Safety and capabilities
Grok 4.6's safeguards have been improved and calibrated in line with the model's capabilities.
Our safety stack is designed to maximize utility and security across legitimate use cases, allowing Grok 4.6 to be helpful and safe in domains such as vulnerability patching, accelerating the engineering design cycle, and augmenting AI research.
Our safeguard evaluation work reflects Grok 4.6’s expanded capabilities, with our widest-ever suite of pre-deployment testing for capabilities and safeguard calibration, as well as extensive post-deployment third-party testing.
Get started with Grok 4.6
Grok 4.6 is available today in Cursor and Grok Build. It's also available in the SpaceXAI API and through partners including OpenRouter, Vercel, and Cloudflare.
Pricing starts at $2 per million input tokens and $6 per million output tokens. A fast variant is available at twice the price.
We’re including 2x usage inside Cursor and Grok Build for the first week.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み