Cursor と SpaceXAI が共同開発した最賢モデル「Grok 4.5」を発表
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Cursor Research
Cursor は SpaceXAI と共同で、ソフトウェアエンジニアリングを超えた複雑なタスク処理を可能にする最良のモデル「Grok 4.5」を発表し、同社によるとこれは混合専門家型モデルとして訓練されたものである。
AI深層分析を開く2026年8月4日 14:32
AI深層分析
キーポイント
SpaceXAI との共同開発
Cursor は SpaceXAI と協力して「Grok 4.5」を開発しており、これが同社がソフトウェアエンジニアリング以外の分野向けに初めて構築したモデルである。
混合専門家型アーキテクチャ
Grok 4.5 は混合専門家(MoE)モデルとして設計されており、Cursor の膨大なデータセットを用いて訓練されている。
広範なドメイン対応能力
同社によると、このモデルはソフトウェアエンジニアリング、データサイエンス、金融、法務など、コンピュータ上で行うあらゆる複雑で長時間のタスクを処理できる。
開発者行動データの活用
訓練にはコードベースやツールとのユーザーインタラクションを含む兆単位のトークンが含まれており、開発者の作業方法やエージェントとの相互作用を学習している。
トレーニングデータの多様化
Grok 4.5 はコーディング特化型ではなく、STEMタスクや研究論文を含む幅広い知識領域での学習のために意図的にデータミックスを広く設定した。
重要な引用
Grok 4.5 is a mixture-of-experts model that we trained jointly with SpaceXAI.
Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools.
While we trained our previous model, Composer 2.5, to be a coding specialist, for Grok 4.5 we kept the training data mix deliberately broader.
Many of these problems had to be designed to be difficult enough that even frontier models fail at them.
編集コメントを表示
編集コメント
Cursor がソフトウェアエンジニアリング特化モデルから、より汎用的なタスク処理能力を持つモデルへと戦略を転換した点は注目される。SpaceXAI との共同開発により、大規模な開発者行動データを学習基盤とした新しいアプローチが試みられている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本日、最も賢いモデルであり、ソフトウェアエンジニアリングに特化して初めて構築したモデルである Grok 4.5 を、SpaceXAI と共同でリリースします。
Grok 4.5 は、ソフトウェアエンジニアリング、データサイエンス、金融、法務など、コンピュータ上で行うあらゆるタスクにおいて、ツールを創意工夫して活用し、困難かつ長時間にわたる課題解決に対応できます。

Cursor の個人向け プラン やチーム向け プラン では、このモデルを大幅に利用可能となり、初週は使用量が倍になります。また、モデルのサイバーセキュリティ機能を反映した新たなセーフガードも追加しました。
堅牢な基盤
Grok 4.5 は、SpaceXAI と共同で訓練された混合専門家(Mixture-of-Experts)モデルです。
学習には、コードベースやソフトウェアツールに対する多様なユーザーのインタラクションを捉えた、トリリオン単位のトークンを含む Cursor データセットが含まれています。このデータセットにより、モデルは既存のソフトウェアと開発者・エージェント間の相互作用の両方から学習し、開発者の働き方や環境とのインタラクション方法を理解します。
前モデルの Composer 2.5 をコーディング特化型として訓練したのに対し、Grok 4.5 ではあえてトレーニングデータの構成を広く設定しました。STEM(科学・技術・工学・数学)の高度なタスクや研究論文、その他の知識労働にまで範囲を広げることで、多様なドメインで高い能力を発揮できるようにしています。
難問に対する強化学習
ソフトウェアエンジニアリングから広範な知識労働に至る現実的な環境において、難問への強化学習を行いました。これらの環境を通じて、モデルは問題の調査、ツールの活用、ミスの回復、結果の検証といったスキルを身につけます。
多くの課題は、最先端モデルでも失敗するほど難易度が高いよう設計されています。モデルが向上すると既存のタスクでは新たな学習効果が得られなくなり、かつては高度な推論を必要としていた問題も日常化してしまうからです。
こうした環境を大規模に構築するために、分散型エージェントシステムを開発しました。エンジニアが課題と解決策の検証方法を指定し、多数のエージェントが協力して各環境の作成・テスト・改良を行います。数百人のエンジニアチームが数ヶ月かけてようやく完成させたような環境も存在します。これは、前モデルを活用して次世代モデルの開発を加速した手法の一つです。
Grok 4.5 の利用開始
Grok 4.5 は本日より、デスクトップ版、Web 版、iOS 版、CLI、SDK を含む Cursor 上でご利用いただけます。
個人プランおよびチームプランでは、当社のファーストパーティモデルプールとして本モデルを大幅に活用しており、初週は利用量を倍増させます。
基本モデルの料金は、入力トークン 100 万あたり 2 ドル、出力トークン 100 万あたり 6 ドルです。高速版(fast variant)も用意されており、入力が 4 ドル/M、出力が 18 ドル/M です。
Grok 4.5 と Composer 2.5 は異なるモデルサイズクラスであり、両方のサイズと重みに対応できることを嬉しく思います。Composer 2.5 は引き続き提供され、今後同サイズの新しいモデルもリリースしていきます。
- SWE-Bench Pro と Terminal-Bench のスコアは、第三者モデルの自己申告値です。SWE-Bench Multilingual における GPT 5.5 のスコアは、当社の内部実行結果に基づいています。
- Grok 4.5 は CursorBench で優位性を示していますが、これはトレーニングデータに誤って Cursor コードベースの古いスナップショットが含まれていたためです。その影響の詳細は不明ですが、将来のモデル向けには該当データを削除済みです。同時に、CursorBench のより大規模なアップデートも進行中であり、今回の除外はこの一環です。
原文を表示
Today we are releasing Grok 4.5 together with SpaceXAI, our most intelligent model and the first we've built for more than software engineering.
Grok 4.5 can handle difficult, long-running tasks that require creatively using tools to solve problems, whether in software engineering, data science, finance, legal work, or anything else you do on a computer.

Cursor subscription plans for individuals and teams include significant usage of the model with double usage for the first week. We've also added new safeguards reflecting the model's cybersecurity capabilities.
A strong foundation
Grok 4.5 is a mixture-of-experts model that we trained jointly with SpaceXAI.
Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments.
While we trained our previous model, Composer 2.5, to be a coding specialist, for Grok 4.5 we kept the training data mix deliberately broader. This involved drawing on high-quality STEM tasks, research papers, and other knowledge work, so that the model gained proficiency across a wide range of domains.
Reinforcement learning on difficult problems
We used reinforcement learning on difficult problems in realistic environments spanning both software engineering and broader knowledge work. These environments teach the model to investigate problems, use tools, recover from mistakes, and verify results.
Many of these problems had to be designed to be difficult enough that even frontier models fail at them. As models improve, existing tasks stop teaching them anything new, and problems that once required extensive reasoning become routine.
We developed a distributed agent system to construct these environments at scale. Engineers specify a problem and how a solution is verified, and large groups of agents construct, test, and refine each environment. Some would have taken teams of hundreds of engineers months to build. This is one of the ways in which we used the previous model to accelerate progress on the next model.
Get started with Grok 4.5
Grok 4.5 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.
Individual and team plans include significant usage of the model as part of our first-party model pool, and we are doubling usage for the first week. The base model is priced at $2/M input tokens and $6/M output tokens. There is also a fast variant at $4/M input tokens and $18/M output tokens.
Grok 4.5 and Composer 2.5 are two different model weight classes, and we're excited to support both sizes and weights. Composer 2.5 will remain offered, and we will release new models of this size going forward.
- SWE-Bench Pro and Terminal-Bench show self-reported scores for third-party models. For SWE-Bench multilingual, the GPT 5.5 score comes from our internal run.
- Grok 4.5 has an advantage on CursorBench because an earlier snapshot of the Cursor codebase was accidentally included in training. The exact impact is unclear. That data has been removed for future models, and in parallel we are working on a larger update to CursorBench, hence the exclusion here.
AI算出
主要ニュースainew評価高い
AI モデルそのものの新機能(複雑な長期タスク処理)とベンチマークデータが中心であり、新規性が高い。ただし、日本企業や日本固有の情報が含まれていないため、日本の関連性は低い。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み