Grok 4.6 が Fable 5 Max と同等性能、価格は 85% オフ
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The New Stack AI
SpaceXAI、Alibaba、DeepSeek が短期間に新モデルを相次ぎ発表し、特にダウンロード可能な重みを持つモデルの登場が価格競争に決定的な影響を与えた。
AI深層分析を開く2026年8月15日 13:20
AI深層分析
キーポイント
短時間での3大モデル相次ぎ発表
SpaceXAI の Grok 4.6、Alibaba の Qwen 3.8-Max、DeepSeek の V4-Pro が約24時間の間に相次いで登場し、性能が同等と見なされる中で価格競争が激化した。
ダウンロード可能モデルによる価格天井の設定
Qwen 3.8-Max や DeepSeek V4-Pro が重みの公開を行ったことで、クローズドな研究機関が課す価格の上限を市場側が決める構造が明確になった。
Grok 4.6 の性能向上とコスト不変
SpaceXAI はモデルサイズを増やさずトレーニング手法の改善により性能を向上させ、API 価格を維持したことで「無料の改善」を実現した。
価格競争における選定権の重要性
モデル間の性能差が縮小し価格が収束する中、どのモデルを採用するかを決定する立場(例:Cursor のようなツール)にお金が移動する構造へ変化している。
Grok 4.6 と Fable 5 Max の価格競争
Grok 4.6 は Fable 5 Max よりも85%安く、入力トークンで80%、出力トークンで88%の削減を実現している。
重要な引用
Three frontier models in about 24 hours, and all three pitched on cost because the capabilities are assumed.
Two of the three frontier launches this week came with downloadable weights, which is why the ceiling on what a closed lab can charge is increasingly set by companies giving their models away.
Every benchmark existed to justify a cheaper bill.
Pareto dominant.
編集コメントを表示
編集コメント
今週のニュースは、AI モデル開発の競争が「性能の向上」から「コスト効率とアクセシビリティ」へと明確に転換した瞬間を示している。特に重みの公開という手段が価格設定の天井を押し下げる効果を持つことは、今後の業界戦略において無視できない重要な示唆となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
image私はマット・バーンズ、インサイト・メディアグループの最高コンテンツ責任者です。毎週、私が取り上げるのは最も重要な AI の動向で、この技術を現場で活用する人々や組織にとってそれが何を意味するのかを解説しています。私の主張はシンプルです。AI を使いこなす労働者が業界の次の時代を切り開くのです。このニュースレターは、あなたがその一人になるお手伝いをします。
グロック 4.6 が水曜日に登場しました。わずか数時間後にはクウェン 3.8-Max が続きました。そして木曜日にはディープシーク V4-Pro が加わりました。約 24 時間の間に 3 つの最先端モデルが発表され、そのすべてがコスト競争を訴えています。なぜなら、機能面はすでに同等と見なされているからです。
イーロン・マスク氏の X(旧 Twitter)での投稿「知能、速度、コストを総合的に考慮すれば、グロック 4.6 は客観的に第 1 位だ」という主張に異論を唱えるのは難しいでしょう。2 年前の最先端モデル発表は機能性を競うものでしたが、今は価格競争が主役です。あらゆるベンチマークは、より安価な請求書を正当化するために存在しているようなものです。
今週発表された 3 つの最先端モデルのうち 2 つには、ダウンロード可能な重み(ウェイト)が付属していました。これが、クローズドな研究機関が設定できる価格の上限を、自社のモデルを無料で公開する企業によって決定される increasingly な理由です。
だからこそ、イーロンのカーソル社との提携が、6 月よりもはるかに理にかなっているのです。モデル間の価格差が縮まれば、資金は「どのモデルを採用するか」を決める側に流れます。マスク氏はその点を理解しています。
同じモデル。より良いトレーニング。変わらないコスト。
SpaceXAI は、Grok 4.6 の発表をたった 2 つの文で行いました。同社はこれを「フロンティア(最前線)知能」と位置づけ、「Grok 4.5 と同じ価格で大幅な性能向上を実現した」と述べています。
Artificial Analysis が実施したインテリジェンス指数では、スコアは 61 点となり、Grok 4.5 よりも 5 ポイント上回りました。これは GPT-5.6 Sol と同水準です。また、GDPVal-AA リーダーボードでは 1,753 Elo を記録し、Fable 5 Max の 1,741 Elo をわずかに上回るトップの座を確保しました。
では、この 5 ポイントの差はどこから生まれたのでしょうか?答えは「モデルサイズの拡大」ではありません。製品アナリストのアカシュ・グプタ氏によると、Grok 4.6 は 4.5 と同じく 1.5 トリオンパラメータで動作しており、すべての性能向上は事後学習(post-training)によって達成されたものです。
Terminal-Bench のスコアは、単一のリリースで 66% も跳ね上がりました。一方、API の利用料金は据え置かれています。100 万トークンあたり入力 2 ドル、出力 6 ドルという価格設定は変わりません。同じ重み(weights)を使用しているため、提供コストも変わらないからです。つまり、この性能向上は実質的に「無料」で実現されたのです。
今週の残りの出来事は、この価格設定の背景を浮き彫りにしています。
2026 年 8 月 11 日から 13 日の 24 時間以内に発表された、3 つのフロンティアモデル
- Model:Weights you can download / Price (in / out)
- Grok 4.6:SpaceXAI / No / $2 / $6
- DeepSeek-V4-Pro-0813:DeepSeek / Yes / $0.66 / $1.98 off-peak
- Qwen3.8-Max:Alibaba / Yes (2.4T params, 95B active) / Self-host
- Nemotron 3.5 Lightning:Nvidia / Yes (30B params, 3B active) / Self-host
- Claude Fable 5:Anthropic (参考) / No / $10 / $50
API の価格設定は 100 万トークンあたり(入力/出力)で発表されました。DeepSeek の料金はオフピーク時の価格であり、ピーク時には倍額になります。新しい価格体系は 2026 年 8 月 16 日から適用されます。Fable 5 は比較のために記載されていますが、今週はリリースされませんでした。
SpaceXAI がその二文で省略したベンチマークを、当社のアマンダ・キャスウェルが詳しく分析しました。CursorBench v3.2 では 69.9% のスコアを記録し、Fable 5 Max(70.5%)にわずかに劣るものの、DeepSWE では他のモデル群には及ばない結果となりました。
アリババが Qwen3.8-Max を発表した際、コンサルタントのジェフ・ブロカウ氏は、リリース日付未定のオープンウェイト版への期待に対し懐疑的な見解を示しました。彼は「発売時の写真撮影用に着せられたオープンソース風のジャケットに過ぎない API ビジネスモデルだ」と指摘しました。当時としては妥当な批判でした。しかし現在、その重み(モデルの重み)は公開されています。DeepSeek は翌日 V4-Pro を発表し、OpenAI の Response API へのネイティブサポートと、ピーク・オフピーク料金を明示した価格表を追加しました。オフピーク時の料金はピーク時の半額です。これにより、最先端の知能モデルも夜間や週末の利用が可能になりました。
一方、投資家のガビン・ベイカーは、購入者が実際に計算する数値に基づいて分析を行いました。Grok 4.6 は入力トークンあたりで Fable 5 Max より 80% 安く、出力トークンあたりでは 88% 安いという結果です。彼の結論はたった二つの言葉、「パレート支配的」でした。
開発者の fallback が Hugging Face にあるファイルに依存している場合、ラボがより高性能なモデルに対して単に高い価格を請求することはできません。より良いモデルを従来通りの価格でリリースし、その差額を自社で吸収するしかないのです。
モデルがコモディティ化されれば、ルーティングこそがビジネスの核心となる
Box の CEO、アロン・レヴィーは今後の展開について解説していますが、その内容を直接読む価値は十分にあります。要約すると、安価なモデルが AI 予算を縮小させるわけではありません。むしろ、企業が以前から望んでいたもののコスト面で手が届かなかった業務——例えば、不審なコードベースをスキャンしてセキュリティホールを検出したり、すべての契約書を読み込んだりするエージェントの運用など——に費用が充てられるようになります。レヴィーはこれを「ジェボンズの逆説」と呼んでいますが、その見解には賛同できます。安くなるほど、企業はより多くを利用するようになるからです。
レヴィーが指摘するのは、コストや用途に合わせて調整されたモデルが増えることで、「タスクに応じてルーティングし最適化するレイヤーに存在することの価値が高まる」ことです。
安価なモデルが登場しても、技術スタック全体が均一にコモディティ化されるわけではありません。価値は「どのモデルをどの業務に割り当てるか」を決定する側にシフトします。
Nvidia は今週、まさにその製品をリリースしました。TechCrunch のフレデリック・ラルドノイスが報じたところによると、300 億パラメータのオープンソースモデル「Nemotron 3.5 Lightning」と、各ジョブを最適なモデルへ振り分けるオープンソースのルーター「NeMo Switchyard」です。
Nvidia が発表した自社データによれば、この Switchyard を活用した構成では、オープンソースモデルと Anthropic の Opus 4.8 を組み合わせることで、コストが約 3 分の 1 に抑えられつつも、最先端レベルの精度を維持できるそうです。
互換性のあるモデルが増えるほど、ルーターの価値は高まります。ルーターによってモデル選択はアーキテクチャ上の決定事項から、実行時の判断へと変化します。実験室での開発が書き換え作業から設定変更へと変わることで、簡単に差し替え可能なベンダーに対して価格転嫁を難しくさせるのです。
Towards Data Science に投稿した Sara Nóbrega 氏は、企業が実際に採用している構成比を報告しています。具体的には、ローカルで動作する小規模モデルが約 70%、ミドルティアの API が 20%、最上位の Frontier モデルが 10% という内訳です。Nvidia の調査によると、企業における AI タスクの 40% から 70% は、すでにパラメータ数が 100 億未満のモデルで処理されていると推計されています。
「トークンあたりのコストが安くても、タスクあたりのコストまで安くなるわけではない」という反論はもっともです。Jessica Wachtel が DeepSeek の V4-Flash と V4-Pro を比較した際、最も難しい課題では予算モデルの方が推論能力に優れていましたが、その結果として消費するトークン数は 3 倍にも達しました。最終的な請求額は 9 セント対 10 セントでした。実際に自社のワークロードで実行してみない限り、単価の安さだけでは判断できません。
ここで Cursor の話に戻りましょう。SpaceX は今年 6 月、全株式取引による 600 億ドル規模の買収を発表しました。この取引はまだ完了していませんが、当時としてはコーディングツールチームを雇うための非常に高額な手段に見えました。しかし今週、その買収の真価が明らかになりました。Cursor の共同創設者兼 CEO、25 歳の Michael Truell 氏が Grok 4.6 の開発に携わり、公開記事で「Opus クラスの知性と洗練さを、非常に低コストかつ高速な形で実現している」と評価しています。Cursor の CEO が、自ら Musk 氏の最新モデルを紹介する立場にあるのです。
「Grok Bot」は前日にリリースされ、「あなたのツールにサインインし、あなたと同じように使い、完成した成果物を持って戻ってくる」という AI チームメイトとして売り出されました。エンジニアの Kun Chen 氏はこれを一日かけて分解調査しましたが、その裏には Cursor が存在していました。すべてのユーザーにクラウド VM が割り当てられ、Mac や iOS のアプリは軽量クライアントに過ぎず、実際のコーディング作業はすべて Cursor のクラウドエージェントに任されます。Chen 氏によれば、「Grok Bot」は明らかに構築され、再ブランドされて「Grok」となったそうです。
つまり、SpaceXAI は今や計算資源(コンピュート)、モデル、エージェントの枠組み、そして数億人のユーザーを相手に提供できる体制を整えたことになります。
その中で、他社が真似できないのが計算資源の部分です。ムスク氏はスペースX内にこの計算資源と電力を配置し、これがボトルネックになると予測していました。ブルームバーグによると、Anthropic は現在、SpaceXAI に対して月額 12.5 億ドルを支払っているとのことです。今年春、SpaceXAI の独自モデルは利用可能な計算資源の 11% を使用していましたが、その頃 Anthropic は計算資源の不足に直面していました。
Elon がモデル開発競争で勝つ必要はありません。重要なのは、他社と並んでいられるかどうかです。同様に、オープンウェイト(公開重み)の研究機関が最先端を独占する必要もありません。彼らが果たすべき役割は、その価格基準を設定することだけです。
実行するのは難しいですが、アドバイスとして一つ挙げます。一つのオープンウェイトモデルを選び、実際に業務で動かしてみてください。Codex や Claude が消えるからではありません。価格表にあるすべての割引が存在する理由は、ダウンロード可能なモデルが購入者に別の選択肢を提供しているからです(つまり、他の選択肢も試してみるべきです)。
「Grok 4.6」が「Fable 5 Max」と同等の性能を 85% オフで達成したという記事は、The New Stack に最初に表示されました。
原文を表示
imageI’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.
Grok 4.6 launched on Wednesday. Qwen 3.8-Max hit a few hours later. Then, on Thursday, DeepSeek V4-Pro followed. Three frontier models in about 24 hours, and all three pitched on cost because the capabilities are assumed.
It’s hard to argue against Elon Musk’s post on X: “Grok 4.6 is objectively #1 when considering intelligence, speed & cost”. Two years ago, frontier launches were about capability. Now they’re about price. Every benchmark existed to justify a cheaper bill.
Two of the three frontier launches this week came with downloadable weights, which is why the ceiling on what a closed lab can charge is increasingly set by companies giving their models away.
And that explains why Elon’s Cursor deal makes considerably more sense today than it did in June. When models converge on price, the money moves to whomever decides which model gets the job. Musk bought that.
Same model. Better training. Same bill.
SpaceXAI announced Grok 4.6 in two sentences. Frontier intelligence, the company said, and a “significant improvement over Grok 4.5 at the same price.” Artificial Analysis scored it 61 on its Intelligence Index, five points above 4.5 and even with GPT-5.6 Sol. On the GDPVal-AA leaderboard it went to the top at 1,753 Elo, just past Fable 5 Max at 1,741.
So where did the five points come from? Not a bigger model. As product analyst Aakash Gupta said, Grok 4.6 runs on the same 1.5 trillion parameters as 4.5 and every gain came out of post-training. Terminal-Bench jumped 66% in one release. The API price never moved, $2 in and $6 out per million tokens, because the same weights cost the same to serve. The improvement is, essentially, free.
The rest of the week explains that price.
Three frontier models in 24 hours, Aug. 11-13, 2026
Model
Weights you can download
Price (in / out)
Grok 4.6SpaceXAI
No
$2 / $6
DeepSeek-V4-Pro-0813DeepSeek
Yes
$0.66 / $1.98off-peak
Qwen3.8-MaxAlibaba
Yes2.4T params, 95B active
Self-host
Nemotron 3.5 LightningNvidia
Yes30B params, 3B active
Self-host
Claude Fable 5Anthropic, for reference
No
$10 / $50
Published API pricing per 1M tokens (input / output). DeepSeek rates are off-peak; peak runs double, and the new pricing takes effect Aug. 16, 2026. Fable 5 is shown for comparison and did not ship this week.
Our own Amanda Caswell went through the benchmarks SpaceXAI left out of those two sentences: 69.9% on CursorBench v3.2, a hair under Fable 5 Max at 70.5%, and behind the field on DeepSWE.
When Alibaba announced Qwen3.8-Max, Adrian Bridgwater caught the skepticism from consultant Jeff Brokaw, who called an open-weights promise with no release “the API business model wearing an open source jacket for the launch photo.” It was a fair take at the time. The weights are out now. DeepSeek followed with V4-Pro a day later, adding native support for OpenAI’s Response API and pricing table with peak and off-peak rates, off-peak running half of peak. Frontier intelligence now has nights and weekends minutes.
Against that, investor Gavin Baker did the math a buyer actually does. Grok 4.6 runs 80% cheaper on input tokens and 88% cheaper on output than Fable 5 Max. His verdict is two words: “Pareto dominant.”
If a developer’s fallback is a file on Hugging Face, a lab can’t simply charge more for a better model. It releases a better model at the old price and eats the difference.
When models become commodities, routing becomes the business
Box CEO Aaron Levie explained what happens next, and it’s worth your time to read it directly. But in short, cheaper models do not shrink AI budgets. They pay for work companies already wanted and could not afford, like agents scanning eerie codebase for security holes or reading every contract. He calls it Jevons paradox, and I think he’s right. The cheaper it gest, the more they buy. What follows from that, Levie wrote, is that more models tuned to different jobs and costs means “the more value there is in being at the layer that can route and optimize based on the task.”
Cheap models do not commoditize the stack evenly. The value slides toward whoever decides which model does what.
Nvidia shipped that exact product this week. TNS’ Frederic Lardinois covered Nemotron 3.5 Lightning, an open 30-billion-parameter model, alongside NeMo Switchyard, an open source router that sends each job to the model that suits it. Nvidia’s own numbers have a Switchyard setup pairing open models with Anthropic’s Opus 4.8 at about a third of the cost, with frontier accuracy intact.
Routers get more value the more interchangeable models become. A router turns model choice into a runtime decision instead of an architecture decision. Leaving a lab stops being a rewrite and starts being a config change, and a vendor you can swap easily has a hard time raising prices.
On Towards Data Science, Sara Nóbrega published the practical version of this, including the split she reports companies actually settling into: roughly 70% local small models, 20% mid-tier API, 10% frontier. She cites Nvidia research estimating that 40% to 70% of enterprise AI tasks already run on models under 10 billion parameters.
The obvious objection is that cheaper per token is not cheaper per job. It’s fair. Jessica Wachtel ran DeepSeek’s V4-Flash against V4-Pro for us and found the budget model reasoned better on the hardest task, then burned three times the tokens getting there. Final bill: Nine cents against ten. Sticker price tells you very little until you have run your own work through it.
Which brings me back to Cursor. SpaceX announced in June that it would buy the company in an all-stock deal valued at $60 billion. It’s worth noting this transaction has not yet closed. At the time it looked like a very expensive way to hire a coding-tools team. This week you could see what he bought. Michael Truell, the 25-year old co-founder and CEO of Cursor, helped launch Grok 4.6 and vouched for it in public writing that 4.6 “combines Opus-class intelligence and polish with very low cost and high speed.” The CEO behind Cursor is not publicly introducing Musk’s latest model.
Grok Bot launched the day before, pitched as AI teammates that “sign in to your tools, use them just like you do, and come back with finished work.” Engineer Kun Chen spent a day taking it apart and found Cursor underneath. Every user gets a cloud VM, the Mac and iOS apps are thin clients, and any real coding work gets handed off to Cursor’s cloud agents. Grok Bot, he wrote, “was clearly built and rebranded to Grok.”
So SpaceXAI now has the compute, the model, the agent harness and a few hundred million people to hand it to.
The compute is the part nobody else can go build. Musk put it inside SpaceX that compute and power would be the constraint, and Bloomberg reported that Anthropic alone now pays SpaceXAI $1.25 billion a month for it. This spring, SpaceXAI’s own models were using 11% of the computing power available to them, while Anthropic worked through a compute crunch.
Elon does not need to win the model race. He just needs to keep up. Likewise, the open-weight labs don’t have to win the frontier. They only have to set its price.
Here’s some advice that’s easier said than done: Pick one open-weight model and run real work through it. Not because Codex or Claude are disappearing. Because every discount in that pricing table exists because downloadable models give buyers another option (so try the options).
The post Grok 4.6 matched Fable 5 Max at an 85% discount. Downloadable models set that price. appeared first on The New Stack.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み