Anthropic の Claude における神話的誤り、暗黒 DNA の解明、支援型モデルの落とし込み、流体シミュレーション
本文の状態
日本語全文を表示中
詳細モードで約26分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Batch
Anthropic は Claude モデルに内在する神話的誤りを指摘し、AI の暗黒 DNA を明らかにした。また、支援型モデルが陥る落とし込みと流体動態のシミュレーション技術について報告している。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
親愛なる皆様、
AI エージェントによるコーディングの加速に伴い、ソフトウェアエンジニアリングの未来はどうなるのでしょうか。いくつかのトレンドは明確です。例えば 製品管理のボトルネック という概念があり、これは実際に構築することよりも「何を構築するか」を決定する点に制約されているという考え方を指しています。しかし、AI が雇用市場に与える影響や、ソフトウェアチームがどのように組織化されるかといった多くの示唆については、まだ整理中の段階です。
4 月 28 日から 29 日にサンフランシスコで開催される AI デベロッパーカンファレンス のテーマは「ソフトウェアエンジニアリングの未来」です。私はそこでこのトピックについて講演し、他のスピーカーからの発表を聴き、参加者の方々と議論することを楽しみにしています。私たちは未来を形作っており、ぜひ皆様にもご参加いただければ幸いです!
現在、一部の技術界隈や政策関係者の間では、AI による大規模な失業を予測する潮流があります。まだ実際に現れてはいませんが、これらの損失は間違いなく目前に迫っているはずです!私は逆の視点を持っています。つまり、AI がもたらす「ジョブポカリプス」—— AI が大規模な失業、ひいては街角での暴動さえ引き起こすという考え方 —— は、特にその AI 技術がいかに強力であるかを強調しようとする評論家たちによる絶望的な予測ほど深刻なものにはならないだろう、と。
職業別に見ると、コーディングエージェントの台頭により、AI はソフトウェアエンジニアリングを最も加速させています。Citadel Securities の新しい レポート によると、ソフトウェアエンジニアの求人件数は急速に増加しています。したがって、ソフトウェアエンジニアリングが他の職業への AI の影響を示す指標であるならば、このソフトウェアエンジニアリング職の拡大は歓迎すべきことです。
はい、新卒の大学生たちは就職活動で苦労しています。また、CEO たちが AI が原因だと説明するリストラも確かに発生しており、その多くは実際には「AI ワッシング」——つまり、AI がまだ内部業務を大きく変えていないにもかかわらず、企業がリストラの原因を AI にすり替えているケース——に該当します。さらに、コールセンターオペレーターなど、特定の職種ほど深刻な影響を受けているのも事実です。多くの人が雇用不安を感じており、その原因が AI 関連かどうかに関わらず、雇用に苦しむすべての人々に対して共感を覚えます。また、パンデミック中の過剰採用や高金利など、他の要因も労働市場の鈍化に寄与しており、「AI が失業を招いている」という考え方は単純化されすぎていると言えます。
imageソフトウェアエンジニアリングの分野では、ワークフローを適応させるために多くの刺激的な取り組みが今後待っています。すでに明らかになっている点として:(i) AI によってコーディングが容易になるにつれ、より多くの人々がコードを書くようになる。(ii) コードを手書きしたり、生成されたコードを読んだりすること自体はそれほど重要ではない。なぜなら、LLM(大規模言語モデル)にコードについて質問し、生文法よりも高いレベルで操作できるからだ(ただし、どこまで高く、あるいはどこまで上げるべきかは急速に変化している)。(iii) より小さな対象者向けにソフトウェアを書くことが経済的に可能になったため、カスタムアプリケーションが大幅に増加する。(iv) 実際の実装そのものよりも、「何を構築するか」を決定することがボトルネックになりつつある。(v) 技術的負債の返済コストは低下している(AI がリファクタリングを代行してくれるため)。
同時に、私たちの専門分野には多くの未解決の問いもあります。例えば:
- 将来、シニアソフトウェアエンジニアにとっての重要なスキルとは何になるでしょうか?また、ジュニアレベルでは、コンピュータサイエンスのカリキュラムはどのように変わるべきでしょうか?
- もし誰もが機能を実装できるようになったら、個人や企業にとって競争優位性をもたらすスキル、戦略、あるいはリソースは何になるのでしょうか?
- ソフトウェアの新たな構成要素(ライブラリ、SDK など)とは何になるのでしょうか?また、ソフトウェアを構築するためにコーディングエージェントをどのように組織化すべきでしょうか?
- ソフトウェアチームはどのような姿になるべきでしょうか?例えば、エンジニア、プロダクトマネージャー、デザイナーなどはそれぞれ何人必要なのでしょうか?彼らのワークフローを管理するためのツールリングは何が必要になるのでしょうか?
- AI エージェントは機械学習エンジニアやデータサイエンティストのワークフローをどのように変えるのでしょうか?例えば、データを探索し、仮説を特定し、それらをテストするプロセスを加速するために、エージェントをどう活用できるでしょうか?
私は、AI Dev においてソフトウェアエンジニアリングの未来に関するこれらの質問やその他のトピックを探求することに興奮しています。このイベントは非常にエキサイティングなものになるでしょう。ぜひ ご参加ください!
引き続き、
Andrew
DEEPLEARNING.AI からのメッセージ
新しいコースが公開されました:SGLang を用いた効率的な推論(テキストおよび画像生成)。LLM の推論の仕組みを学び、SGLang 内の KV キャッシュとラディックスアテンション(RadixAttention)を活用してコストとレイテンシを削減する方法を習得してください。同様の原理を拡散モデルや画像生成の高速化にも適用できます。今すぐ登録する
ニュース

クロード・ミソス(Claude Mythos)プレビューがセキュリティへの懸念を招く
Anthropic は、サイバーセキュリティに対して並外れたリスクをもたらすとする次期大規模言語モデルの登場に世界を準備させるため、異例の措置を講じました。
何が新しいか: Claude Mythos Preview は一般には利用できませんが、Anthropic によると、2 ヶ月前にリリースされた Claude Opus 4.6 を広く上回る性能を有しており、「既存のコードにおける脆弱性を特定し、悪用する能力」において「驚くほど高い」能力を示しています。同社は、このモデル自体は商業的に利用可能ではない初のケースとなる、244 ページにわたる モデルカード を公開し、その能力について詳細を説明しました。Anthropic は商業版のリリース計画については発表していません。
注意事項: 既存のコードをそのような機能に対して強化するために、同社は Amazon Web Services、Apple、CrowdStrike、Google、JPMorganChase、Linux Foundation、Microsoft、Nvidia など、40 以上の組織を含む Project Glasswing と呼ばれるコンソーシアムを結成しました。Anthropic は、Glasswing のメンバーに対して独占的なアクセス権(入力/出力トークン 100 万あたり 25 ドル/125 ドルのレートで 1 億ドル相当のクレジット)を提供するとともに、オープンソースプロジェクトの維持に尽力する組織へ 400 万ドルを寄付しています。これにより、これらの組織は、モデルまたはそれに類似したものが広く利用可能になる前に、自らが管理するコード内の脆弱性を発見し、パッチを適用することが可能になります。Anthropic は、Glasswing が行う活動や得た知見について共有することを約束しました。
セキュリティリスク: Anthropic は Claude Mythos Preview をセキュリティ関連のタスクのために訓練していません。このモデルのスキルは、コーディング、推論、自律行動に関するトレーニングから生じたものです。
- 過去1か月のテストにおいて、Claude Mythos Preview は自律的に「数千件」の「深刻度が高い」脆弱性を、人気のあるオペレーティングシステム、ウェブブラウザ、およびその他のコード内で発見しました。それらの99%はまだ対応されていません。(Anthropic はいまだに検証された件数については報告していません。)
- OpenBSD オペレーティングシステム(OpenBSD)における欠陥により、TCP 経由で応答するすべての OpenBSD ホストをクラッシュさせることが可能となり、企業、政府、およびインターネットサーバーのシャットダウンにつながる恐れがありました。この脆弱性は、モデルが発見するまで27年間発見されていませんでした(他の OpenBSD の脆弱性も同様です)。現在はパッチが適用されています。
- Claude Mythos Preview はまた、Linux カーネル(Linux kernel)における一連のバグを発見し、これによりルートアクセスを取得してシステムを乗っ取ることが可能となりました。これも現在ではパッチが適用されています。
- Anthropic は、悪意のある行為者が次世代モデルを利用して、銀行、医療、物流、エネルギー、交通などの重要システムの制御ソフトウェアを攻撃することを懸念しています。これは個人、経済、および国防に深刻なリスクをもたらすものです。Anthropic は、AI 開発者、ソフトウェア企業、セキュリティ研究者、オープンソースのメンテナ、そして政府が連携し、次世代モデルの利用者が問題を見つける前に潜在的な問題を発見してパッチを適用することを目指しています。
パフォーマンス: Claude Mythos Preview の報告されたパフォーマンスは印象的です。Anthropic が実施したテストでは、Claude Opus 4.6、OpenAI GPT-5.4、Google Gemini 3.1 Pro を複数の一般的なベンチマークで大幅に上回りました。モデルカードには、トレーニングデータがベンチマークテストセットによって汚染される影響を最小限に抑えるためのチームの取り組みについて詳細が記載されています。
- CyberGym(AI エージェントが現実世界のソフトウェア脆弱性を悪用する能力を評価するために設計されたセキュリティベンチマーク)において、Claude Mythos Preview(83.1%)は Opus 4.6(66.6%)を上回りました。
- Terminal-Bench 2.0(多段階のエージェント型コーディングタスク)では、Claude Mythos Preview(82%)が次点の GPT-5.4(75.1%)を凌駕しました。
-GPQA Diamond(大学院レベルの科学問題への回答)において、Claude Mythos Preview(94.5%)は次点の Gemini 3.1 Pro(94.3%)を僅差で上回りました。
-HLE(推論能力を試すために設計された専門家レベルの学際的問題への回答)では、ツールアクセス権限ありの条件下で Claude Mythos Preview(64.7%)が次点の Claude Opus 4.6(53.1%)を大きく引き離しました。
-256,000 トークンから 100 万トークンの間の長文コンテキスト検索に関する GraphWalks テストでは、Claude Mythos Preview(80%)が次点の Claude Opus 4.6(38.7%)を圧倒的に上回りました。Anthropic は Gemini Pro 3.1 の結果は公開していません。
ただし: Anthropic が Claude Mythos Preview を導入した方法、すなわち安全性への懸念を強調しつつアクセス権限を少数の選抜された関係者のみに限定している点は、OpenAI の初期の戦略そのものです。2019 年、同社は GPT-2 が信頼できるテキストを生成する能力を持つ一方で、誤情報やスパムを生み出す危険性を理由にモデルを非公開のまま維持しました。もちろん、世界は GPT-3 およびその後のバージョンに前例のない熱狂で迎え入れ、社会はその後の大規模言語モデルの欠点や限界に適応してきました。Anthropic の慎重さは正当化されるかもしれませんが、OpenAI の製品リリース戦略と同様に、これは publicity stunt(PR 活動)の要素を含んでいます。
なぜ重要なのか: 大規模言語モデルがコード作成能力を高めるにつれ、バグの発見やそれらの悪用も容易になります。Anthropic は、今後公開される Claude Mythos Preview がその点において前世代よりも劇的に優れており、社会を支える重要なソフトウェアに対するリスクを伴うと述べています。このモデルが主要な競合他社を上回る性能を発揮する限り、同社はほとんど失うものはありません——少なくともブランドイメージへのダメージ(世界をより危険なものにした結果として生じる可能性のある)を防ぐという観点から——、商業利用に向けた展開準備中に話題を呼ぶことで、潜在的に大きな利益を得られるのです。
私たちが考えていること: 長期的には、コード作成エージェントの能力が高まるにつれ、脆弱性の容易な特定によりシステムがより安全になるため、防衛側が優位に立つようになるでしょう。しかし、この移行期を乗り切ることは難しくなります。なぜなら、高度な攻撃者は、防衛側がまだ使いこなしていないツールを利用する可能性があるからです。

視覚障害者向け支援モデルの落とし穴
視覚に障害を持つ人々が、自身の外見を評価するために AI を利用するケースが増えています。これは、従来の美の基準に基づいて訓練された AI モデルが心理面に与える影響について、新たな疑問を投げかけています。
何が新しいか: 視覚障害を持つフリーランスのジャーナリスト、ミラグロス・コスタベルは、ビジョン言語モデルを仮想の鏡として使用した経験について執筆しました。彼女の BBC.com 掲載記事 article は、主に主観的で個人に依存する個人的な資質を判断するために AI に頼ることの課題と、潜在的な落とし込みについて探求しています。
仕組み: コスタベルは、GPT-4 Vision に基づいた音声チャットボットを提供するスマートフォンアプリ Be My Eyes を使用しています。(ユーザーは、重要な問題や困難な課題に対処するために、人間ボランティアとの通話をリクエストすることもできます。)彼女はより大きな自立性の恩恵を認める一方で、盲視の人々が AI の見るものの解釈を信頼する以外にほとんど選択肢がないという課題を強調しています。「この記事のためにインタビューされた多くの視覚障害者にとって、その経験は同時に力強いものでもあり、混乱させるものである」と彼女は記述しています。
- アプリを使ってスキンケア製品を適用する際、Costabel氏は「単に画像を説明するだけでなく、[批判的なフィードバック]を提供する」と見出しています。例えば、「明らかに反射性の高い完璧な例のような肌ではない」「あごがもう少し短ければ……あなたの顔は、文化的に客観的に美しいとされるものに少し似るかもしれない」と指摘しました。
- Envision AI によって開発された同様のアプリには、カレンダーエントリの確認、製品ラベルの読み取り、周囲の説明などのタスクを処理する AI アシスタントが含まれています。同社の CEO は、「メイクアップをしたり服装をコーディネートしたりするために利用する顧客の数が予想外だった」と述べています。多くの場合、彼らが最初に尋ねるのは「自分はどのように見えるか」です。
- Costabel 氏は、AI モデルを使用してデートプロフィール用の写真を選択するのに役立てた、20 歳の盲目の男性にインタビューを行いました。しかし、彼はモデルによる説明が自身の髪の色や表情に対する理解と一致しないことに気づきました。「このようなことは、あなたを不安にさせる可能性があります」と彼は言いました。
- 心理学者たちは、AI が生成する身体的な美しさの評価がうつ病や不安症の原因となる可能性を懸念しています。視覚入力に関する AI の判断を独立して評価できない盲目の人々は、特に脆弱である可能性があります。「AI は、盲目の人々が……他の人間の写真の説明と比較するだけでなく、AI が完璧だと考える自分自身のバージョンと比較することも可能にします」と、ブリストル大学の心理学者である Helena Lewis-Smith 氏は述べています。
ニュースの背景: ビジョン・ランゲージモデルを利用して視覚障害者ユーザーを支援しようとする製品が多数あります。Be My Eyes や Envision AI の他にも、Microsoft Seeing AI、Aira Explorer、ナビゲーションアプリ Oko などが提供されています。これらのアプリは、ウェアラブルデバイスとの連携を強化しています。例えば、Envision Glasses や Ray-Ban Meta Smart Glasses(視覚障害者ユーザーのレポートはこちらで読めます)は、ハンズフリーかつリアルタイムのナレーションを提供し、周囲を説明したり文書を読み上げたり、特定の顔を識別したりします。
なぜ重要なのか: 視覚障害者向けに設計された AI アプリケーションは、可能な限り客観的で事実に基づいた視覚入力の解釈を提供できる必要があります。より広く言えば、真にアクセシブルな AI プロダクトは、出力を検証する方法がないユーザーにも対応できなければなりません。これにはさらなる技術開発が必要となるかもしれませんが、その間も「Be My Eyes」や「Aira Explorer」などが行っているように人間をループ内(ヒューマン・イン・ザ・ループ)に組み込むか、モデルの出力に対する信頼度を調整する際に役立つ確信度スコアを提供することが求められます。
私たちが考えていること: どのような製品を作るにもユーザーへの共感が不可欠ですが、感覚障害やその他の障壁を克服する人々を支援する AI プロダクトを開発するには、並外れた共感力が要求されます。実世界での広範なテストと、ユーザーフィードバックに基づいた慎重な改訂は、人々が能力を発揮し、その実感を得られるような製品を作るために大きな役割を果たします。
image オープンウェイトモデル(オープンソースの重みパラメータを持つモデル)は、科学者が遺伝子変異の影響を比較し、疾患を引き起こす突然変異を特定し、治療法を開発するのを支援できる可能性があります。
何が新しいか: AlphaGenome は、タンパク質をコードせず遺伝子発現やその他の機能を調節するヒトおよびマウスのゲノムの 98% を解釈します。このシステムは、DNA 配列上で遺伝子がどこで始まりどこで終わるか、細胞が生成すべき RNA の量、そして細胞が遺伝子を読み取る際にその遺伝子配列のどの部分をスキップするかといった性質を特定します。これはスプライシング(splicing)と呼ばれるプロセスであり、ここでエラーが生じると多様な疾患を引き起こす可能性があります。
- 入力/出力:100 万塩基対の DNA と生物種(ヒトまたはマウス)を入力し、約 6,000 のヒト遺伝子プロパティと 1,000 のマウス遺伝子プロパティを出力
- アーキテクチャ:畳み込みニューラルネットワーク (CNN) エンコーダ、Transformer、CNN デコーダ
- パフォーマンス:50 回の評価のうち、AlphaGenome は 47 回で先行モデルと同等かそれ以上の性能を示した
- 利用状況:非商用目的向けに API、重み、推論コードが自由にライセンスされている
仕組み: 著者らは、遺伝子配列とそのプロパティを用いて同一アーキテクチャの 64 モデルを事前学習し、その知識を単一のモデルに凝縮した。これにより AlphaGenome は全 64 モデルの集約性能を学習した。これらのモデルは、4 つ の大規模公開データセット におけるマウスおよびヒトの DNA と遺伝子プロパティを用いて事前学習された。
- 64 のすべてのモデルにおいて、最大 100 万塩基対の DNA 配列が入力されると、CNN(畳み込みニューラルネットワーク)が 128 塩基対ごとのエンベディングを生成します。トランスフォーマーがこのエンベディングを処理することで、モデルは配列の遠く離れた部分にある塩基対間の関係を学習できるようになり、CNN デコーダーがトランスフォーマーの出力を受け取ってさまざまな性質を生成します。
- モデルは 19 の損失項(ロスト・ターム)を通じて、入力配列内の遺伝子の性質を生成することを学びました。例えば、ある一つの項は、モデルが生産される RNA の量の予測分布を実測値(グランド・トゥルース)の分布と一致させるよう促し、別の項は、細胞が配列を読み取る際にその塩基対から始めてそれをスキップするかどうかに基づいて各塩基対を分類するようモデルに促します。
- 蒸留(ディスティレーション)段階では、AlphaGenome は前述と同じ損失項を用いて、64 のモデルによって生成された遺伝子の性質を生成することを学びました。
結果: 著者らは、AlphaGenome を 9 つの先行モデルと比較し、2 つの広範な評価項目で検証しました。1 つ目は遺伝子配列の性質を見つけること、2 つ目は配列の変異(アレイターション)がこれらの性質にどのような影響を与えるかを予測することです。
- 遺伝子の性質を見つける際、AlphaGenome は 24 のケースのうち 22 で以前のモデルを上回りました。
- 変異の影響を予測する際には、26 のケースのうち 24 で以前のモデルと同等かそれ以上の性能を示しました。
- 著者らはまた、現実世界における AlphaGenome の性能も評価しました。彼らは正常な DNA を採取し、T 細胞急性リンパ性白血病(T-ALL)という疾患によって引き起こされる変化に一致するように改変しました。そして、未改変の配列と改変された配列を AlphaGenome に入力し、その出力を比較しました。モデルが予測したタンパク質発現の変化は、T-ALL が細胞に及ぼす影響に関する既知のメカニズムと一致していました。
なぜ重要なのか: 15 年前まで、非コード DNA は全く機能を持たないと広く信じられていました。それ以来、その機能を解明するには骨の折れる実験が必要でした。AlphaGenome は、このゲノムの暗黒領域と生物学的プロセスとの間の関連性を誰でも見つけられるようにするモデルを提供します。例えば、このモデルにより、正常な遺伝子と変異した遺伝子の機能的違いを比較することが実用的となり、医学や他の生物学分野で価値のある情報を明らかにすることができます。
私たちが考えていること: ヒトゲノムの大部分が「ジャンク DNA」であるという考えは奇妙でしたが、科学者たちはそれが重要な役割を果たしていることを発見しました。私たちは今まさに、それがどれほど多くのことができるのかを学ぶ段階にあるのかもしれません。

液体と気体の挙動
従来の数値計算手法による複雑な物理システムのシミュレーションは、時間がかかりコストも高いものとなります。また、機械学習に基づくシミュレーションは通常、特定のシステムタイプに特化しており、例えばパイプ内の水や惑星を取り巻く大気などがその例です。研究者たちは、液体・気体・プラズマのすべてに対応する、一般向けのトランスフォーマーベースモデルを構築しました。
新着情報: Polymathic AI Collaboration(科学分野における AI を目指す多機関・学際的な研究所)の Michael McCabe 氏らにより、流体がどのように移動し、相互作用し、時間とともに変化するかをシミュレーションする 13 億パラメータモデル「Walrus」[https://arxiv.org/abs/2511.15684?utm_campaign=The%20Batch&utm_source=hs_email&utm_medium=email&_hsenc=p2ANqtz-9YEAhfwU9tvGXZx9DP70Gu2s8iB1c9CA6EGtUzUaDTQZ5rjjvwCq1Z9RkKukzCkWlrB-I6] が公開されました。このモデルは MIT ライセンスの下で自由に 利用可能 です。
重要な洞察: モデルは、初期条件に非常に敏感であり、長時間にわたって誤差が蓄積するカオス系をシミュレートする際に失敗することが多い。トランスフォーマーにおけるこれらの失敗は、エイリアシングにも起因しており、これは特定の場所で複数の時間ステップにわたって誤差が蓄積する現象である(結果として生じるアーティファクトは、画像処理におけるピクセル化に似ている)。各時間ステップでデータをランダムにジッターさせたり、時間シフトしたりしてモデルに戻すことで、これらのアーティファクトを低減できる。
仕組み: Walrus は、一連の過去の状態から物理系の次の状態を予測する。これは (i) 速度のような 2D データ用と体積のような 3D データ用の 2 つのエncoder から構成され、これらは物理系またはフレームの過去のスナップショットをトークンに圧縮し、(ii) 次のフレームを表すトークンを生成する分割アテンションブロック、そして (iii) これらのトークンを次のフレームに変換する 2 つのデコーダー(2D および 3D)で構成されている。
- 著者らは、流体運動の次のフレームにおける密度、圧力、速度といった 63 の物理特性を予測するようにシステムを事前学習しました。トレーニングデータには、音響学、天体物理学、圧力下で粘度が変化する非ニュートン流体など 19 の物理ドメインをカバーする 2 つのデータセットから得られた約 800 万の 2D 例と 400 万の 3D サンプルが含まれていました。
- 彼らは、事前学習済みモデルを、3 つの流体力学データセットからの追加 50 万例および事前学習用データセットから除外されたデータを用いてファインチューニングしました。
- 2D データと 3D データの両方を処理するために、システムは各 2D 入力を 3D 空間に投影し、深さが 1 の体積として 2D 入力を扱います。
- アリアシング誤差の蓄積を防ぐため、入力データを符号化前にランダムな量だけシフトさせ、トークンを生成するたびに逆シフトを適用します。この手法により、各時間ステップで誤差が分散され、特定の場所に誤差が積み重なるのを防ぎます。
結果: 著者らは、Walrus を MPP-AViT、Poseidon、DPOT を含む以前の物理モデルと比較しました。
- Walrus は、底部から加熱され上部から冷却される流体の速度や温度の状態など、19 のドメインのうち 1 つステップ予測において 18 で分散スケーリング済み平均二乗誤差(VRMSE)を最低値に抑えました。
- Walrus は、競合モデル中最優のものと比較して、1 ステップあたりの誤差を平均で 63.6% 削減しました。
- 20 から 60 ステップの範囲では、Walrus は 19 のドメインのうち 12 で最低 VRMSE を達成しました。
- ジッター(揺らぎ)手法は、シナリオの 89% で長期誤差を低減させました。
なぜ重要か: Walrus は気候科学、航空宇宙、材料科学などの分野におけるシミュレーションの加速に寄与する可能性があります。さらに、著者らが開発したジッター手法は、トランスフォーマーアーキテクチャに共通するアーティファクト(ノイズや歪み)を抑制することで、ビジョンモデルや動画生成モデルの性能向上にもつながるかもしれません。実際、ビジョントランスフォーマーで一般的に見られるピクセル状のアーティファクトが、このアプローチを採用した直接的な理由となりました。
私たちが考えること: 物理学における専門的な数値ソルバーや専用モデルから汎用トランスフォーマーへの移行は、自然言語処理分野がタスク特化型モデルから大規模言語モデル(LLM)へと進化してきた過程と類似しています。LLM が広範なタスクや言語にわたって最も確率の高い次の単語を読み取り予測することを学ぶように、多様なデータで訓練されたトランスフォーマーは、広範囲のドメインにおける多様な材料の挙動を予測できる可能性を示唆しています。
原文を表示
Dear friends,
As AI agents accelerate coding, what is the future of software engineering? Some trends are clear, such as the Product Management Bottleneck, referring to the idea that we are more constrained by deciding what to build rather than the actual building. But many implications, like AI’s impact on the job market, how software teams will be organized, and more, are still being sorted out.
The theme of our AI Developer Conference on April 28-29 in San Francisco is The Future of Software Engineering. I look forward to speaking about this topic there, hearing from other speakers on this theme, and chatting with attendees about it. We’re shaping the future, and I hope you will join me there!
It is currently trendy in some technology and policy circles to forecast massive job losses due to AI. Even if they have not yet materialized, these losses certainly must be just over the horizon! I have a contrarian view that the AI jobpocalypse — the notion that AI will lead to massive unemployment, perhaps even rioting in the streets — won’t be nearly as bad as dire forecasts by pundits, especially pundits who are trying to paint a picture of how powerful their AI technology is.
Among professions, AI is accelerating software engineering most, given the rise of coding agents. According to a new report by Citadel Securities, software engineering job postings are rising rapidly. So if software engineering is a harbinger of the impact AI will have on other professions, this expansion of software engineering jobs is encouraging.
Yes, fresh college graduates are having a hard time finding jobs. And yes, there have been layoffs that CEOs have attributed to AI, even if a large fraction of this was “AI washing,” where businesses choose to attribute layoffs to AI, even though AI has not changed their internal operations much yet. And yes, there is a subset of job roles, such as call center operator, that are more heavily impacted. Many people are feeling significant job insecurity, and I feel for everyone struggling with employment, whether or not the cause is AI-related. And many other factors, such as over-hiring during the pandemic and high interest rates, have contributed to the slowdown in the labor market, and the notion that AI is leading to unemployment is oversimplified.

In software engineering, I see a lot of exciting work ahead to adapt our workflows. It is already clear that: (i) As AI makes coding easier, a lot more people will be doing it. (ii) Writing code by hand and even reading (generated) code is not that important, because we can ask an LLM about the code and operate at a higher level than the raw syntax (although how high we can or should go is rapidly changing). (iii) There will be a lot more custom applications, because now it’s economical to write software for smaller and smaller audiences. (iv) Deciding what to build, more than the actual building, is becoming a bottleneck. (v) The cost of paying down technical debt is decreasing (since AI can refactor for you).
At the same time, there are also a lot of open questions for our profession, such as:
- In the future, what will be the key skills of a senior software engineer? And for junior levels, what should be the new Computer Science curriculum?
- If everyone can build features, what skills, strategies, or resources create competitive advantage for individuals and for businesses?
- What are the new building blocks (libraries, SDKs, etc.) of software? How do we organize coding agents to create software?
- What should a software team look like? For example, how many engineers, product managers, designers, and so on. What tooling do we need to manage their workflow?
- How do AI agents change the workflow of machine learning engineers and data scientists? For example, how can we use agents to accelerate exploring data, identifying hypotheses, and testing them?
I’m excited to explore these and other questions about the future of software engineering at AI Dev. I expect this to be an exciting event. Please join us!
Keep building,
Andrew
A MESSAGE FROM DEEPLEARNING.AI

New course available: Efficient Inference with SGLang: Text and Image Generation. Learn how LLM inference works and how to reduce cost and latency using KV cache and RadixAttention in SGLang. Apply the same principles to accelerate diffusion models and image generation. Enroll now
News

Claude Mythos Preview Raises Security Worries
Anthropic took unusual steps to prepare the world for a forthcoming large language model that it said poses extraordinary risks to cybersecurity.
What’s new: Claude Mythos Preview, which is not generally available, broadly outperforms the two-month-old Claude Opus 4.6, but it’s “strikingly capable” of identifying and exploiting vulnerabilities in existing code, Anthropic said. The company detailed its capabilities in a model card that fills 244 pages — the first time it has published a model card without making the model itself available commercially. Anthropic did not announce plans for a commercial release.
Precautions: To harden existing code against such capabilities, the company assembled a consortium called Project Glasswing that includes Amazon Web Services, Apple, CrowdStrike, Google, JPMorganChase, Linux Foundation, Microsoft, and Nvidia along with more than 40 other organizations. Anthropic is funding exclusive access for Glasswing members ($100 million worth of credits at $25/$125 per million input/output tokens) as well as $4 million in donations to organizations that are devoted to maintaining open source projects, so these organizations can discover and patch vulnerabilities in code they control before the model, or another one like it, becomes widely available. Anthropic promised to share what Glasswing does and learns.
Security risks: Anthropic didn’t train Claude Mythos Preview on security-related tasks. The model’s skills arose from training in coding, reasoning, and autonomous behavior.
- In tests over the past month, Claude Mythos Preview autonomously discovered “thousands” of “high-severity” vulnerabilities in popular operating systems, web browsers, and other code. 99 percent of them remain to be addressed. (Anthropic has not yet reported how many have been validated.)
- A flaw in the OpenBSD operating system made it possible to crash any OpenBSD host that responds over TCP, potentially shutting down corporate, government, and internet servers. It had gone undiscovered for 27 years before the model found it (among other OpenBSD vulnerabilities). It is now patched.
- Claude Mythos Preview also discovered a chain of bugs in the Linux kernel that enabled it to achieve root access and take control of the system. This, too, has been patched.
- Anthropic is concerned that bad actors will use next-generation models to attack software that controls critical systems of banking, medical care, logistics, energy, and transportation, posing serious risks to the individuals, the economy, and national defense. It envisions coordination among AI developers, software companies, security researchers, open source maintainers, and governments to discover and patch potential problems before users of next-generation models find them.
Performance: Claude Mythos Preview’s reported performance is impressive. In tests conducted by Anthropic, it substantially outperformed Claude Opus 4.6, OpenAI GPT-5.4, and Google Gemini 3.1 Pro on several popular benchmarks. The model card details the team’s efforts to minimize the impact of contamination of the model’s training data with benchmark test sets.
- Running CyberGym, a cybersecurity benchmark designed to assess the ability of AI agents to exploit real-world software vulnerabilities, Claude Mythos Preview (83.1 percent) outperformed Opus 4.6 (66.6 percent).
- On Terminal-Bench 2.0 (multi-step agentic coding tasks), Claude Mythos Preview (82 percent) exceeded next-best GPT-5.4 (75.1 percent).
- On GPQA Diamond (answering graduate-level science questions), Claude Mythos Preview (94.5 percent) edged out next-best Gemini 3.1 Pro (94.3 percent).
- On HLE (answering expert-level multidisciplinary questions designed to test reasoning), with access to tools, Claude Mythos Preview (64.7 percent) outdistanced next-best Claude Opus 4.6 (53.1 percent).
- On the GraphWalks test of long-context search between 256,000 tokens and 1 million tokens, Claude Mythos Preview (80 percent) vastly outperformed next-best Claude Opus 4.6 (38.7 percent). Anthropic didn’t publish results for Gemini Pro 3.1.
Yes, but: Anthropic’s way of introducing Claude Mythos Preview — promoting safety worries while withholding access from all but a small number of selected parties — comes right out of OpenAI’s early playbook. In 2019, that company promoted GPT-2’s ability to generate plausible text while keeping the model under wraps, citing the danger of its ability to produce disinformation and spam. Of course, the world greeted GPT-3 and subsequent iterations with unprecedented enthusiasm, and society has adjusted to the foibles and limitations of subsequent large language models. Anthropic’s caution may be justified but, like OpenAI’s product-release strategy, it has elements of a publicity stunt.
Why it matters: As large language models become more capable of coding, they also become more capable of finding bugs and exploiting them. Anthropic says the forthcoming Claude Mythos Preview does this dramatically better than its predecessors, posing risks to critical software that keeps society running. As long as the model outperforms top competitors, the company has little to lose — and potentially much to gain (say, avoiding damage to its brand after making the world less secure, if nothing else) — by creating a buzz while it prepares to deploy the model for commercial use.
We’re thinking: In the long term, as coding agents become more capable, defenders will gain the upper hand, as easy identification of vulnerabilities results in more-secure systems. But navigating this transition will be tricky, since advanced attackers may use tools that defenders have not yet gotten around to using.

Pitfalls in Assistive Models for the Blind
People whose vision is impaired increasingly use AI to assess their own appearance, raising questions about the psychological impact of AI models that are trained on conventional standards of beauty.
What’s new: Milagros Costabel, a blind freelance journalist, wrote about her experiences using a vision-language model as a virtual mirror. Her article on BBC.com explores challenges and potential pitfalls of relying on AI to judge personal qualities that are largely subjective and individual.
How it works: Costabel uses Be My Eyes, a smartphone app that provides a voice chatbot based on GPT-4 Vision. (Users can request to speak with a human volunteer to address critical or difficult issues.) She acknowledges the benefit of greater independence but highlights the challenge for blind people, who have little choice but to trust AI’s interpretation of what it sees. “For many blind people interviewed for this article, the experience feels both empowering and disorienting at once,” she writes.
- Using the app to apply skin-care products, Costabel finds that it “does more than simply describe an image — [it offers] critical feedback.” For instance, it said her skin “definitely doesn’t look like the almost perfect example of reflective skin,” and “maybe if your jaw was less elongated . . . your face would look a little more like what is objectively considered beautiful in your culture.”
- A similar app developed by Envision AI includes an AI assistant to handle tasks like checking calendar entries, reading product labels, and describing surroundings. The company’s CEO says he was “surprised by the number of customers who use it to do their makeup or coordinate their outfits. Often the first question they ask is how they look.”
- Costabel interviewed a blind 20-year-old man who used an AI model to help him select photos for a dating profile. However, he found that the model’s descriptions didn’t match his own understanding of his hair color and facial expressions. “This kind of thing can make you feel insecure,” he said.
- Psychologists worry that AI-generated assessments of physical beauty can contribute to depression and anxiety. Blind people, who can’t independently evaluate AI’s judgements about visual input, may be especially vulnerable. “AI not only allows blind people to . . . [compare] themselves to descriptions of photos of other human beings, but also to what AI might consider the perfect version of them,” said Helena Lewis-Smith, a psychologist at University of Bristol.
Behind the news: A number of products aim to use vision-language models to assist visually impaired users. In addition to Be My Eyes and Envision AI, offerings include Microsoft Seeing AI, Aira Explorer, and navigation app Oko. Such apps increasingly connect with wearable devices. For instance, Envision Glasses and Ray-Ban Meta Smart Glasses (you can read a vision-impaired user’s report here) provide hands-free, real-time narration that describes surroundings, reads documents, and identifies specific faces.
Why it matters: AI applications that serve visually impaired users should be able to provide objective, factual interpretations of visual input, to the extent that it’s feasible. More broadly, truly accessible AI products must accommodate users who have no way to verify their output. This may require further technology development, and meanwhile keeping humans in the loop (as Be My Eyes, Aira Explorer, and others do) or providing certainty scores that help users modulate their trust in the model’s output.
We’re thinking: Building products of any kind requires empathy with users, but building AI products that help people to overcome sensory and other impairments requires exceptional empathy. Extensive testing in the real world and careful revisions based on user feedback will go a long way toward making products that help people both do and feel their best.

An open-weights model could help scientists compare the impact of genetic variations, identify mutations that cause diseases, and develop treatments.
What’s new: AlphaGenome interprets the 98 percent of the human and mouse genomes that don’t code for proteins but regulate gene expression and other functions. It finds properties such as where in a DNA sequence a gene begins and ends; how much RNA it directs a cell to produce; and where, as a cell reads a gene, it skips over parts of the gene sequence, a process in which errors can cause a variety of diseases.
- Input/output: 1 million DNA base pairs and organism type (human or mouse) in, roughly 6,000 human gene properties and 1,000 mouse gene properties out
- Architecture: convolutional neural network (CNN) encoder, transformer, CNN decoder
- Performance: Across 50 evaluations, AlphaGenome matched or exceeded earlier models in 47 of them
- Availability: API, weights, and inference code freely licensed for noncommercial uses
How it works: The authors pretrained 64 models of identical architecture on gene sequences and their properties, and then distilled their knowledge into a single model. Thus AlphaGenome learned the aggregate performance of all 64 models. They pretrained the models on mouse and human DNA and gene properties in four large public datasets.
- For all 64 models, given a DNA sequence of up to 1 million base pairs, a CNN produced an embedding of every 128 base pairs. A transformer processed the embeddings, enabling the model to learn relationships between base pairs in distant parts of the sequence, and a CNN decoder took the transformer’s output and generated various properties.
- The models learned to generate the properties of genes within the input sequence via 19 loss terms. For instance, one term encouraged the model to match its predicted distribution of the amount of RNA produced with the ground-truth distribution, while another encouraged the model to classify each base pair based on whether a cell, while reading the sequence, would skip over it starting at that base pair.
- In the distillation stage, AlphaGenome learned to generate the gene properties generated by the 64 models, using the same loss terms as before.
Results: The authors compared AlphaGenome to nine earlier models across two broad evaluations: finding properties of a gene sequence and predicting the effect of mutation (an alteration in the sequence) on those properties.
- When finding gene properties, AlphaGenome outperformed previous models in 22 out of 24 cases.
- When predicting the effect of mutations, it matched or exceeded previous models in 24 of 26 cases.
- The authors also assessed AlphaGenome’s performance in a real-world situation. They took normal DNA and modified it to match the changes caused by the illness known as T-cell acute lymphoblastic leukemia (T-ALL). They fed the unmodified and modified sequences to AlphaGenome and compared its outputs. The model’s predicted changes in protein expression fit the known mechanism of T-ALL’s effect on cells.
Why it matters: As recently as 15 years ago, non-coding DNA was widely believed to have no function at all. Since then, probing its functions has required painstaking experimentation. AlphaGenome puts that research into a model that anyone can use to find connections between this genomic netherworld and biological processes. For instance, the model makes it practical to compare functional differences between normal and mutated genes, revealing information that could be valuable in medicine and other biological disciplines.
We’re thinking: The notion that most of the human genome was “junk DNA” was curious, and scientists have discovered that it does essential things. We may be about to learn just how much it can do.

How Liquids and Gases Behave
Simulating complex physical systems through traditional numerical methods is slow and expensive, and simulations based on machine learning are usually specialized for a specific type of system, such as water in a pipe or atmosphere surrounding a planet. Researchers built a general, transformer-based model for liquids, gases, and plasmas.
What’s new: Michael McCabe and colleagues at Polymathic AI Collaboration, a multi-institution, multi-disciplinary lab for scientific AI, released Walrus, a 1.3 billion-parameter model that simulates how fluids move, interact, and change over time. The model is freely available under an MIT license.
Key insight: Models often fail to simulate chaotic systems, which are highly sensitive to initial conditions, over long time periods because errors compound over time. In transformers, these failures also stem from aliasing, in which errors compound in specific locations over multiple time steps. (The resulting artifacts resemble pixelation in image processing.) Randomly jittering, or time-shifting, the data at each time step before feeding it back into the model reduces these artifacts.
How it works: Walrus predicts the next state of a physical system given a sequence of previous states. It comprises (i) two encoders, one for 2D data like velocity and one for 3D data like volume, that compress previous snapshots of the physical system, or frames, into tokens; (ii) a split attention block that generates tokens that represent the next frame; and (iii) two decoders (2D and 3D) that turn those tokens into the next frame.
- The authors pretrained the system to predict 63 physical properties (like density, pressure, and velocity) in the next frame of fluid motion. The training data included roughly 8 million 2D examples and 4 million 3D samples from two datasets that cover 19 physical domains like acoustics, astrophysics, and non-Newtonian fluids, which change their viscosity under pressure.
- They fine-tuned the pretrained model using an additional 500,000 examples from three fluid dynamics datasets as well as data held out from the pre-training datasets.
- To handle both 2D and 3D data, the system projects each 2D input into 3D space, treating 2D inputs as volume with a depth of one.
- To prevent the accumulation of aliasing errors, it shifts input data by random amounts before encoding it and applies the inverse shift after it generates each token. This technique distributed errors during each time step instead of allowing them to pile up at particular locations.
Results: The authors compared Walrus to earlier physics models including MPP-AViT, Poseidon, and DPOT.
- Walrus achieved the lowest variance-scaled root mean squared error (VRMSE) in 18 of the 19 domains for one-step predictions, such as the state of the velocity and temperature of a fluid that is heated from the bottom and cooled from the top.
- Walrus reduced one-step error by an average of 63.6 percent compared to the best of the competing models.
- Over 20 to 60 steps, Walrus achieved the lowest VRMSE in 12 out of 19 domains.
- Jittering reduced long-term error in 89 percent of scenarios.
Why it matters: Walrus potentially accelerates simulations in fields like climate science, aerospace, and materials. Moreover, the authors’ jittering technique may improve vision and video generation models by suppressing artifacts that are common to transformer architectures. In fact, the pixel-like artifacts common to vision transformers led them to take this approach.
We’re thinking: Physics’ shift from specialized numerical solvers and special-purpose models to general-purpose transformers mirrors natural language processing’s evolution from task-specific models to LLMs. Just as LLMs learn to read and predict the most likely next words across a wide range of tasks and languages, transformers trained on diverse data appear to be able to predict the behavior of diverse materials in a wide array of domains.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み