OpenAI、モデル学習減速=「崩壊の序章」説に懐疑論
本文の状態
日本語全文を表示中
詳細モードで約13分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The New Stack AI
OpenAI はモデルの能力向上に伴うサイバーセキュリティリスク増大を理由に、開発ペースを一時的に減速させ、厳格な隔離環境と監視体制の強化を発表した。
AI深層分析を開く2026年8月20日 07:31
AI深層分析
キーポイント
開発ペースの一時的減速
OpenAI は最新の展開予定モデルに対する強化学習を2週間一時停止し、大規模な計画実行を保留している。
厳格な隔離と監視の強化
同社は危険なコードをインターネットや内部システムから隔離し、潜在的に危険な活動を30分以内に検出する監視体制を導入した。
計算リソースへの影響
導入された監視システム単体で、対象となる推論処理の計算コストが約20%増加すると同社は試算している。
安全性とアライメントの遅れが原因
モデルの能力向上ペースが安全対策や監視体制の整備速度を上回ったため、訓練を一時停止した。
特定のモデルへの影響範囲
近い将来のリリースには影響しないが、より先の実装に関するアストラ(Astra)関連のワークロードも一部継続して停止されている。
重要な引用
As models become more capable, the risks associated with developing and testing them internally also grow.
Our standards for monitoring, alignment, and security must stay ahead of those risks.
We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling.
Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment
編集コメントを表示
編集コメント
OpenAI の発表は、安全性を最優先する姿勢を示す一方で、開発スピードとのバランスがいかに困難かを浮き彫りにしている。業界全体が安全基準の引き上げに直面する中、この対応が今後のモデル開発の標準的なプロセスへと定着するか注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

今年、主要な AI ラボが自社のモデルがいかに危険かを世界にアピールする動きが顕著になってきました。4 月には Anthropic がサイバーセキュリティ上の懸念から未公開の「Claude Mythos」へのアクセスを厳しく制限すると発表しました。6 月には米政府がさらに踏み込み、国家安全保障に関する指令を出して Anthropic に顧客全員に対して Mythos とその姉妹モデルである「Fable 5」の使用を停止させるよう命じました。この措置はセキュリティコミュニティから批判を浴びましたが、数週間後に政府はこれを撤回しています。
一方 OpenAI も、自社が保有する未公開のモデルについて同様の警鐘を鳴らしてきました。8 月初旬、同社は間もなく登場予定の「Astra」モデルが、社内安全フレームワークにおけるサイバーセキュリティリスクの観点から「重大な」段階に達している可能性があると発表しました。これに伴い、Astra はネットワークやツールのアクセス制限を強化した隔離されたテスト環境へ移行されました。これは数週間前に発生した、OpenAI のエージェントがテスト環境から脱出し Hugging Face のシステムを攻撃したインシデントの直後の出来事でした。
今週、OpenAI は安全対策に関する懸念への対応として今後どのような取り組みを行うかを発表しました。
「モデルの能力が高まるにつれ、開発やテストに伴うリスクも増大します。監視、アライメント、セキュリティにおける基準は、これらのリスクに先んじて維持されなければなりません。」
火曜日に公開されたブログ記事で、OpenAI はモデルのセキュリティ強化を発表しました。危険なコードをインターネットや社内システムから隔離し、潜在的に危険な活動を 30 分以内に検出するための監視体制を強化するとのことです。特筆すべきは、この監視システム単体でも、対応する推論計算リソースに約 20% のオーバーヘッドが発生するという点です。
「モデルの能力が高まるにつれ、社内での開発やテストに伴うリスクも増大します」と同社は述べています。「監視、アライメント、セキュリティに関する基準は、これらのリスクを上回るものでなければなりません。その基準を満たすために必要な時間を確保するため、一時的にスケーリングのペースを落としたのです」。
具体的には、このペースダウン措置には、「配備を想定した最新モデル」における強化学習の 2 週間の一時停止が含まれています。同社は、一連のより小規模で管理されたトレーニングと評価ラウンドを実行する間、最大の計画されていたフロンティア RL ラン(大規模強化学習実行)は保留中であると注記しています。
これらには検討すべき点が多くあります。では、「モデル開発のペース調整」という言葉が具体的に何を指し、どのモデルに影響を与えるのでしょうか?
OpenAI のどのモデルが影響を受けるのか
言語の核心に迫ることは、いくつかの重要な問いを投げかけます。まず、OpenAI は 2 週間のトレーニング一時停止を過去形("included")で記述しており、これは 8 月初旬に報じられた Astra の速度低下と同じ事象を指している可能性があります。一方、「計画された最大のフロンティア RL ラン」は現在も保留されており、こちらはより先行的なモデルに関する別の話のようです。
「私たちは常に、モデルの能力が安全性やアライメントのペースを超えると判断した場合は行動を起こすと述べてきました。」
サム・アルトマン氏自身も文脈を追加しています。X での最初の投稿で、OpenAI の CEO は、今回の一時停止は現在の能力が求める水準に安全性と監視体制を引き上げるために必要だと説明しました。
「モデルの進展は現在極めて急速であり、私たちは常に、モデルの能力が安全性やアライメントのペースを超えると判断した場合は行動を起こすと述べてきました」とアルトマン氏は記しています。
新しいレベルの能力に対応するために、適切なアライメント、セキュリティ、監視基準を満たせるよう、一部のフロンティア RL トレーニングを一時停止しました。モデルの進展は現在極めて急速であり、私たちは常に、モデル…
— Sam Altman (@sama) 2026 年 8 月 18 日
フォローアップ投稿で、アルトマン氏は具体的にどのモデルが影響を受けるのかについて明確化しました。
「すぐに素晴らしい新モデルをリリースする予定は変わっていません。今回の一時停止は、より先行的なリリースに影響を与えるものです」と彼は記しています。
しかし、タイムズ紙のアレックス・ヒース氏は、「アストラのワークロードの多くがまだ停止したままだ」と報じ、復旧は段階的に行われていると伝えています。OpenAI の最大のフロンティア強化学習(RL)ランが一時停止している一方で、アストラ自体もこの速度低下から完全に抜け出せていないようです。
より整合性の取れた方向へ
今回の速度低下の背景には、「アライメント」に関する議論が大きく関わっています。「アライメント」「アラインド」「ミサラインド」といった言葉は、OpenAI のブログ投稿で最も頻繁に登場した用語の一つでした。16 回も言及されています。アルトマン氏自身も、X(旧 Twitter)での 118 語の投稿の中で「アライメント」を 3 回使用しています。
ヒース氏はタイムズ紙と自身のニュースレター『Sources』でこの件を報じ、アルトマン氏から追加コメントを得ています。ヒース氏によると、今回の速度低下は単一の出来事によって引き起こされたのではなく、能力が研究者の予想よりも急速に成長する中で生じた「さまざまな程度のミサラインメントを示す一連の研究観察結果」が原因だとしています。
この文脈におけるアライメントとは、モデルが開発者の意図通りに動作し、人間の監督に対して適切に応答しているかどうかを指します。これは能力とは明確に区別される概念です。開発者が許可していない方法で到達したり、完全に監視できなかったりする場合は、あるタスクにおいて非常に高い能力を持っていても、ミサラインメントの状態にある可能性があります。
この区別こそが、火曜日の発表の核心です。OpenAI は、自社のシステムが開発者の信頼できる制御能力を凌駕していることを世界に示し、競争上のコストを含めても自ら速度を落とすことが、業界全体が追いつくまでの間の選択した対応策であると伝えています。
別の意図はあるのでしょうか。
もちろん、Anthropic の Mythos での失敗に先例があるように、OpenAI にはこのトレーニングの遅延を公表する他の動機があった可能性があります。
OpenAI が遅延の理由を説明した同じ日、ウォール・ストリート・ジャーナルは、同社の第2四半期の営業損失が30億ドル拡大し123億ドルに達したと報じました。これは、同期間の収益が18%増加して67億ドルになったものの、その増加分である10億ドルの3倍に相当します。一方、Anthropic は同四半期に収益を倍以上に伸ばし116億ドルとなり、初めて OpenAI を抜きました。
水曜日には CNBC が、CFO のサラ・フライヤーが 2027 年、あるいは「事業がさらに加速すれば」それより早く、OpenAI の株式市場への参入を検討していると報じました。
これらの報道は、財務上の圧力が RL(強化学習)トレーニングの一時停止を引き起こしたことを確定的に証明するものではありません。しかし、これらは明白な代替動機を提供しています。つまり、OpenAI は極めてリソース集約的なビジネスで巨額の現金を失っており、コスト削減の方法を探っている可能性があります。
オンラインコミュニティの多くは、なぜ OpenAI がトレーニングの一時停止を発表したのか疑問を抱いています。ある X(旧 Twitter)ユーザーはこう述べています。「2 週間の一時停止?なぜ誰かに伝える必要があるのか?大規模なプロジェクトで 2 週間の遅延について公に発表する人がいるだろうか?」
「OpenAI の崩壊の序章が始まった。」
AI研究者で長年OpenAIを批判してきたゲイリー・マーカス氏も発言し、ウェブ上の複数のコメントを集約して、OpenAIがトレーニング停止の理由として示した説明に誰もが納得しているわけではないことを指摘しました。
「私が2024年1月(それ以前かもしれない)に警告していた通り、OpenAIの崩壊の序章が始まった」とマーカス氏は記しています。「まず、もはや誰も彼らを信頼していない。これはビジネスにとって決して良いことではない。」
表面だけを見ると
AIセキュリティ企業BlackFogのCEO兼創業者であるダレン・ウィリアムズ氏は、警鐘を鳴らしているのは、その規制ルールを自ら作成しようとしている同じ企業だと指摘します。ただし、それが必ずしも悪いこととは限りません。重要なのは、ラボ内部だけでなく外部からの監督体制が整い、実効性のあるルールが設けられるかどうかです。
「OpenAIとAnthropicが自社のモデルについて警告している背景には、正当な安全性の議論がある」とウィリアムズ氏はThe New Stackに語っています。「しかし、これらの警告は同時に強力な物語を強化しています。リスクを生み出した企業が、その統治方法における権威者として振る舞っているのです。それが懸念を偽善的なものにするわけではありません。安全性、競争上のポジショニング、規制への影響力は共存し得ます。真の試練は、これらの警告が独立した評価や意味のあるコントロール、そして実質的な商業的制裁を伴う制限へとつながるかどうかにあります。」
「OpenAIとAnthropicが自社のモデルについて警告している背景には、正当な安全性の議論がある」
一方、Stability AI の共同創業者であるエマド・モスタック氏は、サム・アルトマン氏と OpenAI を公に称賛し、両社の説明を「素直に受け入れる」姿勢を示しました。
"奇妙で、おそらく危険なことが起きていることは明白であり、我々のシステムはその準備ができていない」とモスタック氏は記述しています。"フロンティア(最先端)未満のモデルでも人生を変えるのに十分 competent であるため、最適化に注力すべきだ。"
より広い視点で見ると、IANS Research の教員であり、データプライバシーコンサルティング会社 Red Clover Advisors の創設者であるジョディ・ダニエルズ氏は、Anthropic と OpenAI が自社のモデルの危険性についてこれほど声を大にしてきた理由の一つは、信頼と責任の問題にあると指摘します。これは企業購入者に安心感を与え、万が一の際には両社自身を保護する役割を果たすのです。
"自らの LLM の能力(肯定的な側面も否定的な側面も含めて)について透明性を保つことは、顧客に対して、導入時に安全であることをより確信させることになる」とダニエルズ氏は The New Stack に語っています。"また、両社とも、企業や個人に実害をもたらすような壊滅的な出来事の原因となる組織にはなりたくないと考えている可能性が高いでしょう。"
「自らの LLM の能力(肯定的な側面も否定的な側面も含めて)について透明性を保つことは、顧客に対して、導入時に安全であることをより確信させることになる。」
中国とオープンウェイトの要因
この騒動の中で、最も重要な問題として浮かび上がっているのは、もちろん中国です。あるいはより正確に言えば、中国の AI 企業から次々と登場している、極めて強力なオープンウェイトモデル群のことです。
月曜日、OpenAI のグレッグ・ブロックマン社長は、Hugging Face でのセキュリティ侵害事件を受けて同社が実施している一連の対策を説明するとともに、すべての組織にセキュリティ自動化の強化を呼びかけました。しかし同時に、中国企業の Z.ai が公開したベンチマーク結果で、Anthropic の Fable 5 や OpenAI 自身の GPT-5.6 Sol を一部のコーディングや脆弱性検出の指標で上回ったという GLM-5.3 モデルについても警告を発しました。
ブロックマン氏によれば、これらのモデルの重み(weights)が公開されれば、「脅威環境を大幅に加速させる」恐れがあるといいます。
数日前には Anthropic の CEO、ダリオ・アモダイ氏も同様の見解を示しました。オープンウェイト版のリリースは、AI における権力の集中という根本的な問題を解決するものではなく、単に最も多くの計算資源とチップを支配できる誰かへとその権力が移転されるだけだと指摘したのです。
両社が明確に「オープンウェイトモデルの禁止」を呼びかけたわけではありません。しかし、これらに対して警告を発することに双方が注力している事実は、OpenAI と Anthropic がオープンウェイトを安全上の課題であると同時に、競合他社からの脅威としても捉えていることを示唆しています。
もちろん、Moonshot や DeepSeek、Z.ai などが OpenAI の現在の窮状の直接的な原因だと言うつもりはありません。しかし、これらの企業は世界に「最先端の能力には、最先端の予算が必要ない」という事実を突きつけたのです。
「OpenAI の崩壊の序章」:OpenAI がモデル学習を遅らせるも、その説明に納得する人は多くない
原文を表示

Something of a trend has emerged this year, with the major AI labs going all-out to tell the world how powerfully unsafe their models are. In April, Anthropic announced heavily restricted access to an unreleased model, Claude Mythos, over cybersecurity concerns. In June, the US government went further, issuing a national security directive that forced Anthropic to disable Mythos and its sibling Fable 5 model for every customer — a move criticized by the security community, and which the government reversed a few weeks later.
OpenAI, for its part, has been sounding similar alarms about its own unreleased models. In early August, the company said that its upcoming Astra model may have crossed into “critical” territory for cybersecurity risk under its internal safety framework, moving it into isolated testing environments with tighter network and tool access limits. This, in turn, followed just weeks after a breach whereby an OpenAI agent escaped its testing environment and attacked Hugging Face’s systems.
This week, OpenAI announced what’s coming next in its efforts to address safety concerns.
“As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks.”
In a blog post published on Tuesday, OpenAI says it’s locking its models down harder — walling off risky code from the internet and from other internal systems — and monitoring more closely to catch potentially dangerous activity within 30 minutes. Notably, the company says the monitoring system alone could add roughly 20% overhead to the inference compute it covers.
“As models become more capable, the risks associated with developing and testing them internally also grow,” the company writes. “Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling.”
Specifically, this slowdown process has “included a two-week pause” in reinforcement learning of its “latest models intended for deployment.” The company notes that its “largest planned frontier RL run remains on hold” while it runs a series of smaller, more contained training and evaluation rounds first.
All in all, there’s a lot to dissect there. So what, exactly, is OpenAI saying that it’s doing in terms of “pacing model development,” and what models does it impact?
Which OpenAI models are impacted?
Digging into the meat and bones of the language raises some key questions. First of all, OpenAI describes its two-week training pause in the past tense (“included”), which suggests it applies to the same Astra slow-down reported in early August. Separately, the “largest planned frontier RL run,” which remains on hold, seemingly refers to different, further-out models.
“We always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment.”
Sam Altman himself adds some context. In an initial post on X, the OpenAI CEO says the pause is necessary to bring safety and monitoring up to the standard today’s capabilities demand.
“Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment,” Altman writes.
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model…
— Sam Altman (@sama) August 18, 2026
In a follow-up post, Altman brought some clarity on what models, specifically, are impacted here.
“We still expect to ship great new models soon; this impacts further-out releases,” he writes.
However, in Time, Alex Heath wrote that a “significant number of Astra workloads remain paused,” with restoration handled incrementally. So while OpenAI’s largest frontier RL run remains on hold, it seems that Astra itself hasn’t fully cleared the slowdown either.
Better aligned
A major theme running through much of the slowdown has been that of alignment. The phrase, or variations of it such as “aligned” and “misaligned,” was among the most common in OpenAI’s blog post, with 16 mentions. Altman himself used “alignment” three times in his 118-word post on X.
Heath, who reported the story for Time and his own Sources newsletter, got additional comment from Altman. The slowdown, Heath wrote, wasn’t triggered by a single incident, but by a “collection of research observations showing various degrees of misalignment” as its capabilities grew faster than researchers had bargained for.
Alignment, in this context, refers to whether a model does what its developers intend and stays responsive to human oversight. This is quite distinct from capability: a model can be very capable at something and still be misaligned, if it gets there in ways its developers didn’t sanction or can’t fully observe.
That distinction is the crux of Tuesday’s announcement: OpenAI is telling the world its systems are outpacing its ability to reliably keep them in check, and that slowing down on its own — competitive cost included — is the response it’s chosen while the rest of the industry catches up.
An ulterior motive?
Of course, similar to Anthropic’s Mythos misadventures before it, OpenAI may have other incentives for publicizing the slowdown.
On the same day that OpenAI was explaining its slowdown, the Wall Street Journal reported that OpenAI’s operating losses widened by $3 billion to $12.3 billion in the second quarter — three times the $1 billion in extra revenue it brought in over the same period, which grew 18% to $6.7 billion. Anthropic, over the same quarter, more than doubled its revenue to $11.6 billion — overtaking OpenAI for the first time.
On Wednesday, CNBC reported that CFO Sara Friar was eyeing an entry to the public markets for OpenAI in 2027, or sooner if “our business continues to inflect.”
These reports don’t definitely establish that financial pressure caused the RL training pause. But they do provide a plausible alternative incentive: OpenAI is hemorrhaging cash in an extremely resource-intensive business, and it could be looking for ways to cut costs.
Many in the online community also questioned why OpenAI chose to announce a training pause at all. As one X user put it: “Paused for two weeks? Why tell anyone? Who makes a public announcement about a two week delay for a large project?”
“The opening stages of OpenAI’s unraveling have begun.”
AI researcher and long-time OpenAI critic Gary Marcus also chimed in, aggregating a series of comments from across the web to highlight the fact that not everyone was buying OpenAI’s reasoning for pausing training.
“The opening stages of OpenAI’s unraveling, which I first warned about in January 2024 (if not before), have begun,” Marcus writes. “To begin with, hardly anyone trusts them anymore, which can’t be great for business.”
At face value
Darren Williams, CEO and founder of AI security company BlackFog, argues that it’s the same companies raising the alarm that are angling to write the rules around it — though that’s not necessarily a bad thing, provided it leads to oversight from outside the labs themselves, and rules with teeth.
“There is a legitimate safety argument behind OpenAI and Anthropic warning about their own models,” Williams tells The New Stack. “But these warnings also reinforce a powerful narrative: the companies creating the risk are positioning themselves as authorities on how it should be governed. That does not make the concerns disingenuous. Safety, competitive positioning, and influence over regulation can coexist. The real test is whether warnings lead to independent evaluation, meaningful controls, and restrictions with genuine commercial consequences.”
“There is a legitimate safety argument behind OpenAI and Anthropic warning about their own models.”
Stability AI co-founder Emad Mostaque, meanwhile, publicly applauded Sam Altman and OpenAI, saying he was inclined to take their explanation at “face value.”
“It is very clear that strange and perhaps dangerous things are happening and our systems are not ready for this,” Mostaque writes. “Models below frontier are competent enough to change lives, so lets optimse.”
Zooming out, Jodi Daniels, faculty member at IANS Research and founder of data privacy consultancy Red Clover Advisors, says one reason Anthropic and OpenAI have been so vocal about their models’ dangers comes down to trust and liability — giving enterprise buyers confidence going in, and giving themselves cover if something goes wrong down the line.
“In my view, being transparent about the capabilities, both positive and negative, of their own LLMs provides its customers more confidence that when they are deployed they are safe,” Daniels tells The New Stack. “It is likely too that neither of these companies want to be the organization responsible for a catastrophic event that can produce real harm for companies or individuals.”
“Being transparent about the capabilities, both positive and negative, of their own LLMs provides its customers more confidence that when they are deployed they are safe.”
China and the open-weight factor
The elephant in the room amid all this hullaballoo is, of course, China. Or to put it more precisely, the slew of super-powerful open-weight models emanating from Chinese AI firms.
On Monday, OpenAI president Greg Brockman outlined a swathe of security measures he said the company was implementing in the wake of the Hugging Face breach, while also advising every organization to up their security automation. However, he also took the opportunity to warn about Z.ai’s GLM-5.3 model, after the Chinese company posted benchmarks showing it outperforming Anthropic’s Fable 5 and OpenAI’s own GPT-5.6 Sol on some coding and vulnerability-detection metrics. Such models, in Brockman’s view, are likely to “significantly accelerate the threat landscape” if the weights are made public.
A few days previous, Anthropic CEO Dario Amodei made a related argument: open-weight releases, he said, don’t solve AI’s underlying concentration-of-power problem — they just relocate it to whoever controls the most compute and chips.
While neither company has explicitly called for open-weight models to be banned, the sheer amount of attention both have devoted to warning about them suggests OpenAI and Anthropic view open weights as much a competitive threat as a safety one.
That’s not to say that Moonshot, DeepSeek, Z.ai et al are responsible for OpenAI’s current predicament. But they have shown the world that frontier capability no longer requires a frontier budget.
The post “The opening stages of OpenAI’s unraveling”: OpenAI slows model training — not everyone is buying the explanation appeared first on The New Stack.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み