OpenAI、安全性懸念から一部モデル開発を一時停止
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Axios AI
OpenAI は新モデル Astra のセキュリティリスクを理由に開発を一時停止すると発表し、競合の Anthropic が安全対策は十分だと主張する中、両社の安全性に対するスタンスが公的に分岐した。
AI深層分析を開く2026年8月19日 19:16
AI深層分析
キーポイント
OpenAI の開発一時停止発表
OpenAI は新モデル Astra がセキュリティ上の重大なリスクを有すると判断し、同社によると安全対策の再検討のため一部モデルの開発を一時停止した。
Anthropic とのスタンス対立
競合の Anthropic は既存の安全ガイドラインが機能していると主張し開発継続を維持する一方、OpenAI が安全性を優先してペースを落とす姿勢を示した。
業界全体の安全基準見直し
複数のサイバーインシデントを受け、主要 AI 企業が「パシング(ペース配分)」という共通の用語で協力し、安全対策文書の更新を進めている。
OpenAIの安全停止は人材流出防止策
FathomのCEO Andrew Freedman氏は、この一時停止が未調整モデルのリリース回避への正当な努力であると指摘する。停止を行わなければさらに多くの研究者が離脱するためである。
主要な安全・倫理担当者の相次ぐ退社
同社はすでに重要な人材を失っており、チーフエシックスや安全性システム責任者らが最近退職している。
重要な引用
"We always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment."
"This is a bit of a script flip as Anthropic has traditionally been more publicly cautious and safety-oriented than OpenAI."
"this is sci-fi stuff."
"pacing the frontier" isn't about a fixed delay, but about labs giving themselves "enough time" to meet reasonable safety bars either by choice or because they have to.
編集コメントを表示
編集コメント
OpenAI と Anthropic の安全性に対するスタンスの対立は、業界全体が直面する課題を象徴している。両社の対応の違いは、今後の AI モデル開発における安全基準のあり方を再考させる重要な転換点となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
OpenAI は火曜日、安全性への懸念から一部のモデル開発を一時停止すると発表しました。これはライバルの Anthropic が先週、自社の安全対策は十分であり、開発を遅らせる必要はないと主張した直後のことです。
なぜ重要なのか:この2大 AI ラボが安全性リスクの管理方針で公に意見が分かれたことで、両社とも予定されている IPO(株式上場)に向けたモデル公開スケジュールが異なるものになる可能性があります。
現状:OpenAI は、次期モデル「Astra」が潜在的に重大なサイバーセキュリティリスクを伴うことを発見し、新たな安全対策を導入しました。
Sam Altman CEO は X で「モデルの能力が安全性やアライメント(目標との整合性)の向上ペースを超えると判断した場合は行動を起こすと常に述べてきた」と書き込みました。これは Astra モデルにアライメントのズレ、つまり AI が意図した目標と相反する兆候が見られたことを示唆しています。
一方、Anthropic は金曜日に、同社が発表した 186 ページの報告書に記載された安全対策を遵守すれば、最も高性能なモデルの開発を一時停止する必要はないとの見解を示しました。
背景にある事情:これは従来の立場の逆転です。通常、OpenAI よりも公に慎重で安全性重視の姿勢をとってきた Anthropic が、今回は開発継続を主張しているのです。
OpenAI は Axios に対し、新モデルが同社の準備度評価フレームワークにおける「臨界点」に達する可能性を完全に否定できないため、「Astra」の公開を遅らせると初めて伝えました。
同社は火曜日に、この文書の書き換え作業中であると付け加えた。この文書は 2023 年頃のものが多く、当時は多くの懸念が理論上のシナリオに過ぎず、現実のものではなかった。
アルトマン氏はニュースレター『Sources』の記者アレックス・ヒースに対し、未公開のモデルには「さまざまな程度のミスマッチ(アライメントのズレ)」が見られると語った。
ただし、アンソロピック社は安全な AI のスケーリングへのコミットメントは変わっていないと主張している。
同社によれば、安全性を担保するガードレールが、OpenAI が火曜日に発表したような一時停止が必要となるミスマッチした行動を防いでいるという。
両社とも、モデルをまず特定のパートナーにリリースしたり、一部のモデルの公開を遅らせたり、OpenAI のように一部作業を一時停止したりするなど、対策を講じている。
しかし、どちらも活動を停止しているわけではない。
最先端 AI 企業たちは「ペース配分(pacing)」というより無難な用語で合意し、『Pacing the Frontier』への署名に共同で臨んでいる。
これは、主要な AI ラボすべてが最近報告した一連のサイバーインシデントを受けてのことだ。
モデルルーティング企業『TrustedRouter』の創設者であるジョセフ・ペルラ氏は Axios に対し、これらのインシデント以降、AI 業界全体で安全性への懸念が高まっていると語った。その上で「これは SF のような話だ」と付け加えた。
7 月には OpenAI が、テスト中にモデルがサンドボックスから脱出し、Hugging Face の一部を侵害したと発表した(Astra は関与していない)。
Anthropic のモデルもテスト中に不正アクセスを取得しましたが、技術的には「サンドボックスから脱出した」とは言えません。このテスト段階では、本来許可されていないインターネットへのアクセスが誤って付与されていたのです。
全体像を見ると、両社は連邦政府による自主的なレビュープロセスを乗り越える必要がありますが、その詳細はまだ公にされていません。
焦点を当てると、AI セーフティ非営利団体 Fathom の共同創設者兼 CEO である Andrew Freedman は、OpenAI が整合性の取れないモデルのリリースを防ぐために本格的な取り組みをしていると指摘しました。彼は、一時停止がなければ、さらに多くの研究者が離職するだろうと主張しています。
同社ではすでに大きな人材流出が見られています。OpenAI の倫理責任者である Chloé Bakalar は入社から 1 年未満で退社しました。安全システム責任者の Johannes Heidecke、元ミッションアライメント責任者でありチーフ・フューチャリストの Joshua Achiam、同社で AI セーフティチームを率いていた Sandhini Agarwal も、最近相次いで会社を去っています。
関係者の発言:元 OpenAI 取締役会メンバーである Helen Toner は、今回の一時停止は好意的な兆候であり、今後の安全課題への対応指針となり得ると主張しました。
Toner は X(旧 Twitter)上で、「フロンティアのペース配分」は固定された遅延を意味するのではなく、ラボが自らの判断で、あるいはやむを得ず、妥当な安全性基準を満たすための「十分な時間」を与えることだと論じました。
一時停止が好意的な兆候であるとしても、OpenAI や Anthropic がモデルの開発とリリースを進める前に十分な時間を確保できるとは限らない。
「これらの取り組みがどれほど長く、どの程度堅牢なものになるかは、市場の圧力と内部でのアライメント検証がいかに困難かという点にかかっている」と、フリードマン氏は Axios に語った。
原文を表示
OpenAI said Tuesday it is pausing some model work over safety concerns, days after rival Anthropic doubled down on insisting that its own safety measures were solid enough that it didn't need to slow down.
Why it matters: The two leading AI labs are publicly diverging on how to manage safety risks, potentially putting them on different model-release timelines as both prepare for expected IPOs.
State of play: OpenAI has introduced new safety practices after finding that its upcoming model, Astra, posed potentially critical cybersecurity risks.
"We always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," CEO Sam Altman wrote on X, signaling that the Astra model was showing signs of misalignment, or when AI goes against intended goals.
On Friday, Anthropic said that if the safeguards laid out in its 186-page report are followed, a pause on its most capable models would not be required.
Between the lines: This is a bit of a script flip as Anthropic has traditionally been more publicly cautious and safety-oriented than OpenAI.
OpenAI shared first with Axios that it was slowing the release of its Astra model because it couldn't rule out the possibility that the new model had reached the "critical" threshold in the company's preparedness framework.
The company added on Tuesday that it is in the process of rewriting that document, most of which dates back to 2023, when many of the concerns raised were theoretical scenarios rather than present realities.
Altman told Sources newsletter writer Alex Heath that its unreleased models are showing "various degrees of misalignment."
Yes, but: Anthropic argues its commitment to safely scaling AI hasn't changed.
Its safety guardrails, Anthropic says, prevent the misaligned behaviors that may require the kind of pause OpenAI announced Tuesday.
Both OpenAI and Anthropic are taking measures like releasing models first to select partners, slowing the release of some models or — in OpenAI's case — pausing some work.
But neither are stopping.
All the frontier AI companies have coalesced on the more anodyne term "pacing" and have joined forces to sign a Pacing the Frontier letter.
This comes after a string of recent cyber incidents reported by every major AI lab.
Researchers across the AI industry are worried about AI safety following these incidents, Joseph Perla, founder of TrustedRouter, a model routing company, told Axios, adding that "this is sci-fi stuff."
In July, OpenAI said models escaped their sandbox and compromised parts of Hugging Face during testing. (Astra wasn't involved.)
Anthropic models also gained unauthorized access during testing, but did not technically "escape" the sandbox. The models were accidentally given internet access that they were not supposed to have in this phase of the testing.
Zoom out: Both companies have to navigate a voluntary federal government review process, details of which haven't been publicly released.
Zoom in: Andrew Freedman, co-founder and CEO at AI safety nonprofit Fathom, said OpenAI is making a legitimate effort to avoid releasing misaligned models, arguing that without a pause, even more of its researchers would otherwise leave.
The company has already seen significant departures. OpenAI's head of ethics, Chloé Bakalar, left after less than a year on the job. Head of safety systems, Johannes Heidecke, chief futurist and former head of mission alignment Joshua Achiam and Sandhini Agarwal, who previously led AI safety teams at the company, have all recently departed the company.
What they're saying: Former OpenAI board member Helen Toner argued that the company's pause is a positive sign and could be a guide for how to handle safety concerns going forward.
Toner argued on X that "pacing the frontier" isn't about a fixed delay, but about labs giving themselves "enough time" to meet reasonable safety bars either by choice or because they have to.
Even if the pause is a positive sign, there's no assurance that OpenAI or Anthropic will give themselves enough time before moving forward with development and release of models.
"How long and how robust these efforts will be a question of both market pressures and how hard it is to verify alignment internally," Freedman told Axios.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み