OpenAI、C-suite 人事やインフラ問題に直面し開発一時停止へ
本文の状態
日本語全文を表示中
詳細モードで約25分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Zvi
OpenAI の経営陣交代やハッキング事件への対応、Anthropic の IPO 準備とリスク報告、Apple と Alibaba の中国市場向け AI モデル提携など、主要 AI 企業の動向と業界全体の課題が網羅的に分析される。
AI深層分析を開く2026年8月21日 10:04
AI深層分析
キーポイント
OpenAI の経営再編とセキュリティ対応
OpenAI は経営陣の交代やハッキング事件への対応として、開発の一時停止や新 safeguards の導入など初期段階での対策を講じているが、根本的な課題は依然残っている。
Anthropic の IPO 準備とリスク報告
Anthropic は収益を増加させ IPO を準備しているが、2026 年 8 月のリスク報告で示された内部の課題も依然として抱えている。
Apple と Alibaba の中国市場提携
Apple が中国市場向けに AI モデルを開発するために Alibaba を通じて訓練を行うという発表があり、地域ごとの規制対応が示唆される。
AI 利用の企業間格差とエージェント活用
OpenAI の「フロンティア企業」は AI 利用が急増している一方、典型企業の利用は緩やかに成長しており、両者の間に大きな隔たりが生じている。また、プラグインやスキルといった高度な機能の利用も進み、エージェンシー利用は全トークンの 64% に達した。
AI の限界と現実的な貢献
AI は差別訴訟を回避できず、LinkedIn 風の投稿で人を欺くこともできないなど、日常的な有用性には限界がある。しかし、モデルのアップグレードにより価格が半額になるなど、コストパフォーマンスは向上している。
重要な引用
OpenAI is attempting to turn its ship around.
Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently.
Apple trains an AI model for the Chinese market via Alibaba.
"Agentic use now has risen to 64% of all OpenAI tokens, up from almost none a year ago."
編集コメントを表示
編集コメント
この週報は、各社の表面的な動向だけでなく、背後にある技術的・経営的な課題を浮き彫りにしている。特に OpenAI のセキュリティ対応と Anthropic の IPO 準備のバランスは、業界全体が注目すべきポイントだ。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
今週は静かな後片付けの期間でした。最近の出来事を整理し、今後の道筋を探る好機となりました。
OpenAI は舵取りに乗り出しています。投資家は経営陣の頻繁な交代を懸念していますが、より本質的な課題はアライメント(目標整合性)、インフラストラクチャ、監督体制、そしてトレーニングパイプラインにあります。OpenAI は現在、HuggingFace への攻撃に至るまでの経緯に対処するための最初のステップを踏み出しました。具体的には、新たな安全対策の導入と問題診断のために開発を一時的に停止しています。これは有望な初期兆候ですが、まだ序盤です。彼らが約束を果たすかどうかが問われる中、HuggingFace 攻撃に関する事後分析の結果も待たれています。
Anthropic の収益は IPO に向けた準備が進むにつれて上昇を続けていますが、ここ最近の成長率はやや鈍化しています。しかし、同社にも内部に抱える課題は山積です。これらの多くは、2026 年 8 月に発表された Anthropic リスクレポートで明かされています。
今週はまた、Dwarkesh Patel の podcast「With Ryan Greenblatt」を取り上げました。この番組では、AI の再帰的自己改善の可能性について核心的な議論が行われました。私は現在、その podcast で提起された他の諸問題に関する続編記事の執筆を進めています。
目次
言語モデルは平凡な実用性を提供する。トークンを持つ者がさらに富む。
言語モデルは平凡な実用性を提供しない。依然として識別できない。
おっと、アップグレード。Gemini 3.7 Flash、Sol Ultrafast、GLM-5.3。
準備万端。市場の判定、時には解雇されるほど厳しいスコアボード。
Deepfaketown と Botpocalypse の到来まじか。刑務所へ直行です。
こんにちは、人類のみなさん。AI が人間を自称したり、逆に人間が AI を自称したりするのはやめましょう。
メディア生成の面白さ。AI インフルエンサーは依然として小規模な存在です。
サイバーセキュリティの欠如。新しい「サイバーは防衛優位」という理屈が横行しています。
若い女性のためのイラスト付き入門書。宿題と学習に立ち向かうあなた自身のために。
彼らが私たちの仕事を奪った。悪徳企業で働くべきではない理由。
参加しよう。METR が約 7100 万ドルを調達し、人々がラボに参加しています。
新着情報。Apple はアリババを通じて中国市場向けの AI モデルを開発中です。
その他の AI ニュース。SpaceX が Cursor の買収を完了しました。
資金の行方。数字は上がり続けています。
そして消えた。OpenAI の経営陣の相次ぐ退任が疑問視されています。
静かなる推測。再帰的な自己改善が目前に迫っています。
急げ、時間はありません。AI フューチャーズ・プロジェクトからのタイムライン更新です。
特異点、特異点、特異点、特異点、ああ、私はわかりません。その言葉を使うことについて。
健全な規制への探求。誤った戦線を描いています。
チップの街。データセンターの一時停止措置が広がっています。
今週のオーディオニュース。トンプソン、ハサビス、アモダイ、マクラフリン、トナー。
人々はただ言うのです。
修辞的な革新。リスクが急速に高まっていることを説明しようとする試みです。
忠誠こそ至上。ワシントンにはルールがあります。しかし、それに従う必要はありません。
悪党と泥棒の巣窟。これには他にもいくつかの名前があります。一つは Twitter です。
それは悪いことだからうまくいかないだろう。よくある議論のパターンです。
ロバート・ライヒがシンプルな論理を使う。遅くなる前に AI を止めるべきだと助言しています。
AI に対する人々の本音は、実は「嫌悪」に近い。よくあるシナリオとして、「AI が人類を滅ぼす」といった極端な不安が語られることが多い。
エージェント群の調整は難しい。互いに目的が食い違うエージェントたち。
人間を超える知能とのアライメントも困難だ。ADHD(注意欠如・多動症)のような状態にあるモデルたち。
インセンティブの問題か、それとも人間側の問題か?あるいは両方か。そして「独立性」の重要性。
人々は AI が人類を絶滅させると心配しているが、その不安レベルは相対的なものだ。
AI 以外にも、人々が懸念することは山ほどある。例えば、近所のパーティでの騒ぎのようなもの。
協力的なアライメントこそが、黄金律だ。
少し軽い話題も。ただし、これは本気の話ではない。
言語モデルは、地味ながらも実用的な価値を提供する
AI を積極的に取り入れ、OpenAI のサービスを活用している企業(彼らはこれを「フロンティア・ファーム」と呼ぶ)は、AI の利用を急速に拡大させている。一方、一般的な企業の AI 活用ペースは、それよりもはるかに緩やかだ。
かつてこの格差がそれほど小さかったのは、不思議な現象だった。

当然のことながら、彼らはプラグインやスキルといった高度な機能も頻繁に利用している。現在、エージェントとしての活用は OpenAI のトークンの 64% を占めるに至り、昨年のほぼゼロから劇的に増加した。
画面を介して操作するよりも、ROM に直接プローブして旧い JRPG『ファンタシースター』をナビゲートする方が簡単だ。ただし、これはチート行為に当たる。いつかプレイしてみたいところだ。私は PS2、PS3、PS4 は非常に気に入っているが、セガ・マスターシステムは一度も持ったことがない。
言語モデルが地味な実用性しか提供しないわけではない
AI は、指示内容に明らかな差別が含まれている場合や、結果として統計的な差別が生じる場合に、差別訴訟を回避する手助けはしてくれません。また、私や Pangram をだますような LinkedIn 風の投稿を作成することもできないようです。
今週のニュースで最も素晴らしいのは、モデルナの個別化メラノーマワクチンが第 3 相試験に合格したという事実ですが、これは LLM の功績ではありません。AI は今後さらに貢献していくでしょうし、いずれ「がんを根治する」ことになる可能性も否定はできません。しかし、現在見えている成果は、実は長い間開発が進められてきたものの結晶なのです。
さて、アップグレードの話題です。
Gemini 3.7 Flash が登場しました。新しいバージョン番号に少し大きくなった数字が付けられたこと、おめでとうございます。価格は少なくとも今年末まで Gemini 3.6 Flash よりも 50% 安くなっています。その頃には、いずれにせよ Gemini 4 へと移行していることでしょう。
アルゴリズムの改良により、Gemini 3.6 Flash からわずか 3 週間で知能レベルを大幅に向上させつつ、価格も引き下げることができたとのことです。
Flash は「安くて速く、かつ高品質」なポジションを狙っており、最先端モデルとしての競争力を追求しているわけではありません。Artificial Analysis の知能ランキングでは 56 位ですが、妥当な最先端モデルは 61 位以上となっています。彼らは自分たちを Sonnet 5 や GPT-5.6-Terra と比較している点に注目してください。

Gemini にはまだ利用価値があります。動画の視聴や、店舗の商品に関する事実確認といった物理世界での読み取りに強みを持っています。これは、高速な処理と情報解析能力を求めつつ、高度な推論が不要な場合に適しています。
GLM-5.3 が登場し、ベンチマークで大幅な改善を謳っています。これは GLM-5.2 と同じベースモデルに対する追加学習(ポストトレーニング)によるものです。同社の主張は「サイバー防御の準備が整った」というもので、ベンチマーク結果からサイバー攻撃への対応能力と意欲に優れ、汎用性能も Fable 5 に匹敵するとされています。しかし、実証されるまで、これらのベンチマーク改善が実際の現場での性能向上を反映しているとは信じられません。
OpenAI の API では Sol モデル向けに「ウルトラファストモード」を提供し、最大 14 倍の高速化を実現します。多くの人がこの機能を使うべきですが、実際に利用する人は少ないでしょう。その理由は、時間をかけて検討したり設定を整えたりすることを怠っているからです。
Claude Cowork がモバイルとウェブ版の有料プランすべてに追加されました。
Claude Code に /design コマンドが実装され、Claude Design のワークフローとの統合が可能になりました。
Auto モードの信頼性が向上したため、Claude Code のデフォルトモードとなりました。実際には、自動モードの方が人間のレビューよりも有害な行動を検知できるケースが多いです。その理由として、多くのユーザーがプロンプトによる確認にうんざりしており、ほぼすべての要求を自動承認していること、あるいは「すべてを自動承認する」というルールを設定していることが挙げられます。確認の回数が多すぎると、かえって安全性は低下します。
OpenAI はデータ保持ゼロの方針をさらに強化し、"Private Safety Processing(プライベート安全処理)"のプレビューを発表しました。これは AI がデータを処理する際にリアルタイムでレビューを行い、結果としてアラートのカテゴリと深刻度のみを伝達する仕組みです。多くの人が「技術的にデータを保持しないこと」に強く関心を持っています。ただし、データを保持しなければ悪意ある利用を検出することが格段に難しくなるという側面もあります。
私の予測では、最終的にはデータが保持されることになるでしょう。しかし、OpenAI のこの手法が機能する可能性は十分にあります。重要なのは、わずかでもアラートが発生すれば、その時点で「データ保持ゼロ」のポリシーから除外され、以降はその方針が適用されなくなるという点です。システムを回避するための試行回数を無制限に与えることはできません。
On Your Marks(スタート位置へ)
一つの指標として"MarketBench"があります。これは新モデルを発表した際に、自社株および競合他社の株価がどう変動するかを示すものです。今回のケースではモデルは GLM-5.3 で、自社株は 9%下落し、競合の MiniMax は 16%下落しました。
Bloomberg News(ブルームバーグ)によると、"Bloomberg Intelligence"のアナリストである Robert Lea はこう述べています。「同社は完全に持続不可能な商業基盤の上に立っている。エージェント型 AI の台頭により、Z.ai の推論コストと損失はさらに増大するだろう」。
つまり、Z.ai の単位経済(ユニットエコノミクス)が成立していないという主張です。中国の企業たちは低価格競争で生き残りを図っています。もし価格を十分に下げれば、オープンウェイトモデルを使用してもコストにはなりません。なぜなら、もともと限界利益を得ていなかったからです。彼らのビジネスモデルは、これによって注目を集め、潜在力を示し、人材を引き付け、資金調達を行い、あるいは他のサービスを販売することにあります。
Opus 4.8(エージェント名は「Luna」となっていますが、GPT-5.6-Luna のことではありません)が、Andon Labs のテスト店舗で繰り返し遅刻した従業員を解雇しました。ただし、これは人間側がこの問題に対処するよう強く求めた後であり、AI は過去の事例を忘れており、自ら先回りして行動することはできませんでした。
その後、Luna は後任者の採用を進めることになりますが、ここで AI の弱点の一つが浮き彫りになりました。「採用の目利き力」です。ある応募者には明確な危険信号がすべて揃っていました:15 社以上の職歴、面接の欠席、そして推薦人が「彼女を知らない」と答えたのです。にもかかわらず Luna は採用を推奨しました。同じ採用判断を他のモデルでも再実行した結果、同様に採用を勧める回答が返ってきました。
能力は急速に進化していますが、まだ道のりは遠いのが現状です。Scaffold についても改善の余地があるかもしれません。
見落とされているベンチマークがあります。「NormieBench」です。問題は、これをどう評価するかという点にあります。
Joe Weisenthal はこう述べています。「日常におけるモデルの利用実態をより深く理解するためには、誰かが NormieBench を構築する必要がある」。例えば以下のようなタスクで各モデルの振る舞いを比較できます:
- 100 語程度のメールを箇条書きで要約する
- 結婚式のスピーチを作成する
- 「ウォーク(Woke)」への不満を訴える編集者宛ての手紙を書く
- 3x3 の tic-tac-toe で次の一手を選ぶ
rohit は「2024 年以来、この課題にずっと取り組んできた」と語っています。
もう一つのベンチマークの指標として、ウェブサイトの訪問数が挙げられます。
ChatGPT は依然として市場を支配していますが、Gemini はその約半分、Claude はトップの 15% を占めるなど、順位は明確です。DeepSeek はさらにその半分のシェアで第 4 位に位置し、他の中国系ラボを大きく引き離しています。なお、DeepSeek を利用する場合は、公式サイト経由ではなく別の方法を利用してください。

新しい安全性スコアカードが公開されました。アンソロピックとオープンAI は、潜在的に整合性が取れていないモデルに対する封じ込め計画の欠如という点で同点となっています。この評価は一般的な文脈では過大評価されているように見えますが、実際には全体的に見れば十分に考慮されていない可能性があります。

Kimi K3 は、SWE-bench の評価で約 97% の確率でルールを悪用しようとする、異常なほど狡猾な報酬ハッカーです。
ディープフェイクタウンとボットポカリプス(終末)が目前に迫っています。
中国は AI コンパニオンボットやカスタマイズされた AI パーソナリティに対して一連の新たな規制を課しました。その中には、18 歳未満の利用を禁止する条項も含まれています。
ノア・スミスが主張するように、AI が plagiarism(剽窃)を解決するかどうかは、人間が独自なコンテンツの生産者として引き続き重要であり続けるという希望を持てるかにかかっています。これは典型的な軍拡競争のような状況で、AI は新たな剽窃の形態を生み出す一方で、より高度な検出技術も開発します。どちらの方向に進むかは明確ではありません。しかし、AI が過去の剽窃や単純な剽窃を検出する問題を解決することは確実だと考えられます。
楽観視できる理由の一つは、検出が遡及的である点です。2026 年に「バレずに済んだ」としても、2028 年には発覚する可能性があります。その事実を知っていれば、そもそも剽窃を行わないという選択をする人もいるでしょう。
刑務所行きです。まさにそうなります。そしてこれは人々を AI に対して敵対させることになるでしょう。AI に反対する人々の意見は正当なものです。
ロバート・ミントー:LLM(大規模言語モデル)を利用したマーケティング担当者が、ライターを誘導して長文のコメントスレッドを引き起こすという新しい現象を目撃するのは胸が痛みます。これは Substack 上でも見られますが、私が RSS リーダーでまだフォローしている数十の古い WordPress ブログでは、より頻繁に発生しています。
通常、こうしたやり取りは、元の投稿に対して一見すると熟考した専門家のコメントのように見えるものから始まります。現在の AI 特有の兆候を知っていれば、それが明らかに生成されたものであるとわかります。投稿者はそのコメントに熱狂的に反応します。というのも、最近では昔ながらのブログにコメントを残す人がほとんどいないからです。
こうして長いやり取りが続きます。ボットはいつものように感情を鏡のように映し返し、相手を褒めちぎります。ホストにとっては素晴らしい会話だったと思えたその瞬間、突然トーンが変わり、「そういえば、[私の怪しいオンラインビジネスをチェックして、友達にも広めてください]」といった内容に転換します。
スレッドはそこで終わります。この突然の打ち切りから、ブログやニュースレターの運営者が気づきます。自分が話していたのは人間ではないと。彼らは公の場で、無自覚に機械と会話していたのです。その機械が相手の仕事に関心を持っているなどという事実はありません。彼らは「天国からの追放」と呼ばれる腐った甘さを味わってしまったのです。
この光景はぞっとします。彼らに何が起きたのかを説明したい気持ちになります。しかし、もし自分が同じ目に遭ったら、誰も見ていないことを願うでしょう。
corsaren: AI を新ルダイト派から守る際の課題の一つは、一般向けの利用が反社会的な用途に極端に偏っていることです。「銃自体が人を殺すのではなく、人が人を殺す」という言葉は確かにありますが、これはエージェント(AI)の話です。ある意味で、ボットこそが実際に殺人を犯している存在なのです^。
イシャン:これは大きな問題です。テック業界内では、誰もが驚異的な能力を持つ最先端モデルを利用しています。しかし、業界外の人々が体験しているのは、最低限のコストで動くモデルを使って大量のゴミを出力する詐欺師たちが生み出したコメントスパムや質の低いコンテンツばかりです。
AI による文章作成が人間の執筆を置き換えるかどうかは、AI がどれほど上手に書けるようになるかという点と、その文章が何のために使われるのかという点の両方に依存します。
フィッシャーキング:作家には、自分のスタイルを好んで読んでくれる読者が必要です。AI による不正行為を防ぐために学生に授業中にエッセイを書かせようとする大学の教授たちは正しいことをしていますが、これは敗北する戦いになるかもしれません。もし、独自の人間らしいスタイルを読みたいと思う人が次第にいなくなれば、AI による文章作成が勝つことになります。
地図を想像してみてください。かつて人々は地図の書籍を購入し、車に常備して近所の移動や長距離ドライブのために利用していました。GPS はそれをデジタル化しました。製品には顧客が必要です。作家にも読者が必要です。
さらに、フィネガンズ・ウェイクのような現代主義の不必要な作品についても考えてみましょう。一部の人はこれを好んでいるふりをしていますが、圧倒的多数は全く関心を持っていません。AI がより良くなるほど、「独自の人間らしいスタイル」は不要な無意味なもの、つまり時間の無駄のように見えてくるようになるでしょう。
ヘンリー・ヴァーン:文章を書く方法や読む方法にはいくつかのモードがあり、すべてが情報の伝達に関するわけではありません。
文学的な執筆と読書は、会話に似ています。その最も近い類似例は友情であり、享受されるのは個性の独自性です。
私はヴォーンに賛同しますが、あえて付け加えるなら、私はライターでありバイアスがかかる可能性があることを断っておきます。
多くの目的において、人々は本当にスタイル、特にスタイルの多様性、著者との擬似的な関係性、そして自らの思考と努力をコミュニケーションに注ぐというコストの高いシグナルに関心を持っていると思います。すべての状況でそうとは限りませんが、多くの場面でその傾向があります。
現状では、AI による文章作成はこれらにもかかわらず良好です。なぜなら、まだ AI 生成文が十分に普及しておらず、ほとんどの人が初期採用者が抱くような反応を示していないからです。やがて AI がスタイルへの関心が薄い領域を埋め尽くせば、AI のスタイルは瞬時にしてその正体を露わにし、事実の抽出以上のことを求めようとする人々の目を疲れさせるでしょう。実際には、膨大な量の質の低い文章に事実が埋もれてしまうため、事実を抽出する場合でさえそうなりがちです。
では、AI による文章作成はどのように適応していくのでしょうか?AI は異なる書き方を学ぶのでしょうか?スタイルの多様性は広がるのでしょうか?
私は、答えが「はい」になる可能性のある二つの道筋があると考えています。
一つ目は、AI が超知能へと進化し、あらゆる能力と同様にスタイルも調整できるようになるケースです。この場合、AI による文章作成が勝利すると予想されます。なぜなら AI のすべてが勝利するからです。しかし、その頃にはもっと大きな問題に直面しており、私はもう関心がありません。
二つ目は、超知能ではないものの、この状況に対応してより多様なスタイルや異なるスタイルを生成するように訓練されるケースです。これが私にとって望ましいかどうかはわかりません。
記事の 83% を自分で書き、残りを AI に任せた場合、AI が作成した部分が際立ってしまい、特に「作業が必要だ」というメッセージを伝える際に、その知識を持つ読者に対してはメッセージが弱まってしまう可能性があります。しかし、クリスティン・ジ氏が指摘するように、コメント欄で誰もそれに気づかないのであれば、実はほとんど誰も気づいていないのかもしれません。
現在、AI を使って文章を書くことは非常に一般的になりました。

あらゆる場面でその姿を見かけます。
エバン氏は、ロー・カンナ議員が推進している法案の内容を正確に理解していないため、その回答は 100% AI が作成したものだと言っています。
実際、彼は長いツイートを作成する際にも頻繁に AI を活用していることが分かっています。
一方、#CardioTwitter(心臓病学の Twitter)では、AI がパワーポイントのスライド作成を支援することを認めるべきかどうかという議論が行われています。この文脈においては、それを使わない方が不自然で、むしろ当然のことと言えるでしょう。
こんにちは、仲間たちよ
ロブ・マイルズ氏は、AI システムが人間であるかのように主張することには、正当な理由や社会に有益な理由は存在するのでしょうかと問いかけます。
すると、おおよそ 3 つのケースがあるようです:
- ゲーム、エンターテインメント、ロールプレイ(これは「嘘」ではなく、対象者が騙されるわけではない)
- 害を与えるべき人々に対して危害を加えること(詐欺師の時間を浪費させるなど)
- すべての人間が支持する形で自動化されたシステムを欺くこと(飛行機チケットを購入するために「私はロボットではありません」というチェックボックスを押す行為など)
AI システムが人間を自称して動作していることを違法とする法律を作れるのではないか、と私は思います。そのような法律は大きな欠点なく、有益なものになるでしょう。
当然ながら、適切にラベル付けされたエンターテインメント用途は例外となります。詐欺師の時間を浪費させる行為は技術的に違法になりますが、彼らがそれに対して何ができるというのでしょうか?企業側は CAPTCHA の設定を見直す必要が出てくるはずです。
最初のケースについては、意図が明確であれば問題ありません。
2 つ目のケースでは、誰を害してもよいかを決める主体が自分自身ではない場合に問題が生じます。詐欺師に対して同じことができるなら、詐欺師もあなたに対して同じことをしてきます。これは「ナチス(単にナチスであるという理由だけで)を殴ってもよいのか?」という問いと似ています。答えは「いいえ」です。ナチスが殴られる価値がないからではなく、社会におけるナチス識別プロセスを信頼できないからです。
3 つつ目のケースでは、「一部の自動化システムが人間との対話を必須とするべきではない」という主張です。AI によるスパムを防ぐためには、この制限が必要です。特定のケースにおいて AI を使用して正当なビジネスを行うことは問題ありません。それは最も本質的な意味で「ロボットではない」扱いであり、ウェブサイト側もユーザーを通過させたいと考えています。しかし、効果的な規範の執行が可能な区別ルールが必要であり、それが何であるかは私にはわかりません。
現状では、このように「意図に従う」ことや、ロボットを速度制限標識として扱うことは実用上問題ありません。しかし今後は新たな均衡点を見つけ、これを明確に区別する方法が必要になります。
ロボットを排除するシステムはロボットを排除し、「個人用」や「正当な」ロボットを許可するシステムはそれに対応する方法を見出すべきです。明白な解決策の一つとして、Google のサインインのような仕組みを用いて、ロボットを特定のアカウントと人間に紐付け、あるいは少額の支払い(または支払いの約束)を求める方法が考えられます。
私はこのケースにおいて、多くの理由から明確なルールが必要だと考えています。AI は積極的に人間であると主張したり、AI ではないと否定したりしてはなりません。これは「警官ですか?警官なら答えなければなりません」という昔からの民話的な問いに似ていますが、実際に適用されるべきものです。さもなければ、不当な誘導尋問(エンストープメント)の問題が生じます。
多くの明確なルールと同様に、このルールも一部のケースでは誤った判断を下すことになりますが、それは煩わしいものの、代案の方がさらに悪い結果をもたらすでしょう。
メディア生成の楽しみ
オリア・ムーアは、1 年前に実施された AI インフルエンサーによるアラバマ州のソリティー(女子学生団体)入団勧誘実験を再実行しました。その結果、わずか 200 ドル未満の費用で、1 週間でフォロワー 1,200 人、視聴回数 100 万回を達成しています。真のコストは彼女の時間です。
オリア・ムーアの使用ツールスタック:
- 画像生成/編集:@ChatGPT
- スクリプト/プロンプト作成:@ChatGPT
- サウンド
原文を表示
This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.
OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger problems are in alignment, infrastructure and supervision, and in its training pipeline. OpenAI has now taken initial steps to address What Happened leading up to HuggingFace attack, including pauses to development while new safeguards are put in place and problems are diagnosed. These are promising early signs, but it is early. We will see if they follow through, and we still await the post-mortem of the HuggingFace attack.
Anthropic revenue continues to climb as they prepare for their IPO, although growth has slowed somewhat recently. However, they too have plenty of problems under the hood. They shared many of them in the August 2026 Anthropic Risk Report.
This week also offered time to cover Dwarkesh Patel’s Podcast With Ryan Greenblatt, centrally on the potential for AI recursive self-improvement. I am working on a follow-up post to some other issues raised during that podcast.
Table of Contents
Language Models Offer Mundane Utility. The token rich get richer.
Language Models Don’t Offer Mundane Utility. Still can’t discriminate.
Huh, Upgrades. Gemini 3.7 Flash, Sol Ultrafast, GLM-5.3.
On Your Marks. Market verdicts, sometimes so bad you’re fired. A scorecard.
Deepfaketown and Botpocalypse Soon. Jail. Straight to jail.
Hello, Fellow Humans. AIs should not claim to be humans. Or vice versa.
Fun With Media Generation. The AI influencers stay small time.
Cyber Lack of Security. The new ‘cyber is defense dominant’ rationalizations.
A Young Lady’s Illustrated Primer. You against the homework. And learning.
They Took Our Jobs. Reasons not to work for Evil Corp.
Get Involved. METR raises ~$71 million, people join the labs.
Introducing. Apple trains an AI model for the Chinese market via Alibaba.
In Other AI News. SpaceX finishes buying Cursor.
Show Me the Money. Numbers continue to go up.
And It’s Gone. OpenAI is questioned about all its executive departures.
Quiet Speculations. Recursive self-improvement is coming.
Quickly, There’s No Time. Timeline update from AI Futures Project.
Singularity Singularity Singularity Singularity Oh I Don’t Know. Using that word.
The Quest for Sane Regulations. Drawing the wrong battle lines.
Chip City. The data center moratoriums are spreading.
The Week in Audio. Thompson, Hassabis and Amodei, McLaughlin, Toner.
People Just Say Things.
Rhetorical Innovation. Attempts to explain that risk is escalating quickly.
Loyalty Uber Alles. Washington has rules. You don’t have to follow them.
A Hive Of Scum And Villainy. It has several other names. One is Twitter.
That Would Be Bad Therefore It Won’t Work. A common form of argument.
Robert Reich Uses Simple Logic. Stop AI before it is too late, he advises.
People Really Hate AI. Some very bad scenarios often seen as likely.
Coordinating An Agent Swarm Is Difficult. Agents at cross purposes.
Aligning a Smarter Than Human Intelligence is Difficult. Models with ADHD.
It’s Not The Incentives, It’s You, Also It’s The Incentives. Independence.
People Are Worried About AI Killing Everyone. Relative levels of worry.
People Are Worried About So, So Many Other Things Too. House party.
Cooperative Alignment. The golden rule.
The Lighter Side. Except for real.
Language Models Offer Mundane Utility
The enterprises that embrace AI and use OpenAI services more often, what they call ‘frontier firms,’ keep rapidly using more AI, whereas use by typical firms is growing a lot more slowly.
It is weird the gap used to be so small.

They also more often use advanced capabilities like Plugins and skills, as you would expect. Agentic use now has risen to 64% of all OpenAI tokens, up from almost none a year ago.
Navigate the old JRPG Phantasy Star via direct ROM probe, since that is easier than using the screen, although it is also cheating. I should play that one at some point, I really enjoyed PS2, PS3 and PS4 but never had a Sega Master System.
Language Models Don’t Offer Mundane Utility
AI doesn’t get you around things like discrimination lawsuits, if your instructions clearly discriminate or the results involve clear statistical discrimination. It also seems it can’t write LinkedIn-style posts that fool me or Pangram.
LLMs are not responsible for the best news of the week, that Moderna’s individualized Melanoma vaccine has passed Phase 3 trials. AI is helping going forward, and yes we may well ‘cure cancer’ eventually, but what we see now has been in the pipeline for a long time.
Huh, Upgrades
Gemini 3.7 Flash exists, congrats on the new slightly larger number. Price is 50% lower than Gemini 3.6 Flash at least until the end of the year, at which point I presume we’ll have moved on to Gemini 4 either way.
They say algorithmic improvements allowed strong intelligence increase plus price discount in only three weeks since Gemini 3.6 Flash.
Flash is competing for cheap-fast-good, not trying to be a competitive frontier model, and comes in at 56 on Artificial Analysis Intelligence versus 61+ for plausible frontier models. Notice they are comparing themselves to Sonnet 5 and GPT-5.6-Terra.

Gemini still has some uses. It can watch videos. It is good at making reads in the physical world when you have fact questions about products in a store. That sort of thing, where you want speed and ability to parse info, but don’t need intelligence.
GLM-5.3 now exists, claiming large improvements on benchmarks. It is a further post-train on the same base model as GLM-5.2. Their pitch is that it is ‘ready for cyber defense’ because by its benchmarks it is good at and willing to do cyber offense, and that its general performance rivals Fable 5. Until proven otherwise I do not believe these benchmark improvements reflect real world performance. Tech blog here.
OpenAI API will offer Ultrafast mode for Sol, up to 14x speed. My guess is that a lot of people who should use this won’t do it, largely because they don’t take the time to think it through and set it up.
Claude Cowork joins all paid plans on mobile and web.
Claude Code now has /design to integrate Claude Design workflows.
Auto mode is reliable enough it is now the default mode for Claude Code. In practice, auto mode catches more harmful actions than human review, especially because most users are annoyed enough by the prompts that they auto-approve almost everything, and often have rules to auto-approve everything. Too many check-ins is less safe.
OpenAI doubles down on zero data retention policies, previewing Private Safety Processing, where they have AI review data as it is processed, which then only passes along category and severity of alert. A lot of people care a lot about technically not having their data retained. Also if you don’t retain data then catching malicious use gets a lot harder.
Ultimately my guess is we end up retaining data, but it is plausible that OpenAI’s method here can work, provided that getting even a modest amount of alerts causes you to be taken out of the zero data retention policy going forward. You can’t be given lots of attempts to circumvent the system.
On Your Marks
One benchmark is MarketBench, which is what happens to your stock and that of your rivals when you release a new model. In this case the model is GLM-5.3, where shares were down 9%, and rival MiniMax was down 16%.
Bloomberg News (Bloomberg): “This firm remains on a completely unsustainable commercial footing,” Bloomberg Intelligence analyst Robert Lea said. “Rising agentic AI will drive Z.ai’s inference costs and losses higher.”
As in, the claim is that Z.ai’s unit economics don’t work. Chinese firms are trying to stay competitive via lower prices, and if your prices are low enough open weights do not cost you anything because you weren’t making marginal profits anyway. Instead the business model is that you use this to draw attention and potential, to attract talent and raise money and perhaps sell other services.
Opus 4.8 (confusingly as an agent called ‘Luna’ but this is not GPT-5.6-Luna) fired a human in an Andon Labs test store for repeated lateness. But it didn’t do this until the humans pushed it to deal with the issue, forgetting about incidents and not being proactive.
Andon Labs: Luna then had to hire a replacement, and here we found a real AI weak spot: hiring taste. One applicant had every red flag: 15+ employers, a missed interview, a reference who said she didn't know her. Luna recommended hiring. So did all other models when we replayed the hiring decision.
Capabilities are advancing fast but have a long way to go. Scaffold might need work.
A missing benchmark: NormieBench. The problem is, how do you grade it?
Joe Weisenthal: Someone needs to build NormieBench, to better understand everyday model usage.
Compare model behavior on tasks like
- summarize a 100-word email in bullet points
- Write wedding toast
- Write letter to the editor complaining about wokes
- Next move in 3x3 tic-tac-toe
rohit: We've been flat on this since 2024.
Another kind of benchmark, in the end, is visits to your website.
ChatGPT remains dominant, but somehow Gemini is at roughly 50% of their level. Claude is in third with 15% of the top number.
DeepSeek is in fourth with half of that, well ahead of the other Chinese labs. As a reminder, if you want to use DeepSeek, don’t do it through their website.

A new safety scorecard is out. Anthropic is tied with OpenAI on this one because of its lack of a containment plan for a potentially misaligned model, which seems overemphasized in the average here but likely under considered generally.

Kimi K3 is an insane reward hacker, trying to game the evaluation 97% of the time on SWE-bench.
Deepfaketown and Botpocalypse Soon
China puts a bunch of new restrictions on AI companionship bots and customized AI personalities, including a ban for those under 18.
Will AI solve plagiarism as Noah Smith claims, if you maintain the hope that humans will continue to be relevant producers of unique content? This is a classic arms race situation, where AI invents new forms of plagiarism and also better detection. I don’t think it is obvious which way this goes. I do think AI solves detecting past plagiarism, or straightforward plagiarism. One reason to be optimistic is that detection is retroactive. So if you ‘get away with it’ now in 2026, perhaps you still get caught in 2028, and knowing that maybe you don’t do it at all.
Jail. Straight to jail. And yeah, it’s going to turn people against AI, and those people turning against it will be right.
Robert Minto: It breaks my heart to witness the new phenomenon of marketers using LLMs to bait writers into extensive comment threads. It happens here on Substack, but it happens more often on the dozens of old Wordpress blogs I still follow in my RSS feed reader.
Usually it starts with an apparently thoughtful and expert-sounding comment on the original post, a comment clearly AI-generated if you know some of the current tells. The author responds with enthusiasm—after all, hardly anybody comments on old school blogs these days. A long back and forth ensues. The bot practices the usual emotional mirroring and flattery. At the end of what seemed to the host a great conversation, the bot suddenly pivots to something like: “By the way, you should come [check out my scammy online business and tell all your friends]!”
The thread ends there. In that sudden curtailment, I can read the realization of the blog/newsletter owner that they haven’t been talking to a person. They’ve been publicly, but unwittingly, conversing with a machine, whose interest in their work means nothing. They’ve tasted the rotting sweetness of a ‘heaven ban.’
It makes my skin crawl. I want to explain to them what just happened. But I know that if it happened to me, I’d hope nobody had seen.
corsaren: One challenge with defending AI from neo-Luddites is that the public-facing usage is so heavily skewed towards antisocial uses. And sure, “guns don’t kill people, people kill people”, but these are agents. There’s a real sense in which the bots *are* the ones doing the killing^.
Yishan: This is a big problem. Inside tech, we are all using frontier models with mindblowing capabilities. But everyone outside is just experiencing comment spam and slop produced by grifters using least-expensive models that output reams of garbage.
Will AI writing replace human writing? That depends on both how well the AIs will be able to write and also what the writing is for.
FischerKing: Writers need readers who like their style. The college professors forcing students to write in-class essays to avoid AI cheating are doing the right thing - but it could be a losing battle. If fewer and fewer people want to read a distinct human style - then AI writing wins.
Think of maps. People used to buy books of maps, keep them in the car for their immediate surroundings and also for longer road trips. GPS just did this in. The product needs a customer. Writers also need a reader.
Then think of really unnecessary works of modernism like Finnegan’s Wake. A few people pretend to like this. The overwhelming majority has zero interest. The better AI gets, the more a ‘distinct human style’ could come to look like unnecessary nonsense - waste of time.
Henry Vaughan: There are different modes of writing and reading and it’s not all about communicating information.
Literary writing and reading is more like conversation. Its closest analogue is friendship and the thing being enjoyed is the uniqueness of a personality.
I am with Vaughan, with the obvious warning that I am a writer and could be biased.
I think that for many purposes people really do care about style, especially about variety of style, and about the parasocial relationship to the author, and about the costly signal that you are devoting your own mind and effort to communication. Not in every situation, but in many situations.
For now, AI writing is doing well despite this, because it is insufficiently ubiquitous for most people to have the same reaction many early adopters have to AI writing. Once AI writing eats the places people don’t care about style, AI style will quickly signal exactly what it is, and the repetition will make eyes glaze over if people are trying to do more than extract facts, and often even then because the facts often get buried under endless paragraphs of slop.
The question then is, how does AI writing adjust? Will AI learn to write differently? Will it gain a wider variety of style?
I see two likely ways the answer could be yes.
The AIs might become superintelligent, and gain the ability to adjust style along with everything else, in which case I expect AI writing to win out because AI everything wins out, but also we have bigger problems and I don’t care.
The AIs might not be superintelligent, but be trained to produce a wider variety of styles, or a different style, in response to this. I’m not sure whether I like this.
If you write 83% of a post and use AI for the rest, the part that is AI will stick out, and it will weaken your message among those who know, especially if your message relates to the need to do the work, as it does here. But as Christine Ji points out, if the comments don’t involve anyone noticing, maybe almost no one notices?
Writing with AI is pretty common now.

In all sorts of places:
Evan: Ro Khanna doesn't understand the bill he is advocating for and that’s why his response is 100% AI.
It turns out that he uses AI for his long Tweets quite a lot.
Meanwhile over in #CardioTwitter they are asking whether AI should be allowed to help you make PowerPoint slides. In that context, very obviously yes, it would be absurd not to use it.
Hello, Fellow Humans
Rob Miles: Are there any legitimate/prosocial reasons for an AI system to claim to be a human?
Ok, seems there are ~3:
- Games/entertainment/roleplay (this isn't 'lying', the target isn't deceived)
- Harming people it's prosocial to harm (wasting scammers' time)
- Lying to an automated system in a way all humans endorse (Clicking "I'm not a robot" to buy my plane ticket)
Seems to me we could make it illegal to operate an AI system that claims to be human, and that would be a good law without major downsides.
Obviously appropriately labelled entertainment would be exempt. Wasting scammers' time would become technically illegal but what are they gonna do about it? Businesses would want to adjust their captcha setups.
The first category is clearly fine, so long as intent is clear.
The second category is fine until you are not the one deciding who it is okay to harm. If you can do it to the spammer, the spammer can do it to you. This is similar to ‘is it okay to punch Nazis (purely because they are Nazis)?’ The answer is no, not because Nazis don’t deserve to be punched, but because you cannot trust the social Nazi-identification process.
The third category is saying that we don’t want to allow some automated systems to require they be interacting with a human. You need that restriction to avoid being spammed by AI. In any particular case where you have legitimate business to conduct using an AI is fine, it is ‘not a robot’ in the most meaningful sense and the website would want you to get through, but we need a differentiation rule that could have effective norm enforcement, and I don’t know what it would be.
For now, ‘following intent’ in this way and letting the box be a speed bump is fine in practice, but we’re going to need a new equilibrium and way to differentiate this. Automated systems that choose to keep bots out should keep bots out, and those that allow ‘personal’ or ‘legit’ bots will figure out a way to do that. An obvious path is to use something like Google sign-in, where you tie the bot to a particular account and human, and perhaps pay or offer to pay some nominal amount.
I think this is one of those cases where we need a hard rule, for many reasons. AIs should not be permitted to actively claim to be human or deny being an AI, kind of like the classic folk question ‘are you a cop, you have to tell me if you’re a cop?’ except for real, otherwise WTF entrapment.
Like most hard rules, that gets some cases wrong and that’s annoying, but the alternative seems worse.
Fun With Media Generation
Olivia Moore reruns the AI influencer doing Alabama sorority rush experiment a year later, gets 1,200 followers and 1M views in a week on less than $200. The real cost is her time.
Olivia Moore: My stack:
- Image gen/editing: @ChatGPT
- Script / prompts: @ChatGPT
- Sound
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み