METR と Redwood が語る HuggingFace ハッキング事件の事後分析
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
コメントの核心
AI セーフティ予測の的中と懐疑論
意見が分かれるLessWrong や AI セーフティ界隈が以前から警告していた失敗モードがこの事件で再現され、彼らの抽象的な最適化プロセスモデルが有効だったとする評価がある。一方で、報告書作成に AI が多用されたため、AI によるバイアスが結果に影響している可能性を指摘し信頼性を問う声もある。
代表コメント(原文)
“A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi…”
組織の盲点と人間の役割の軽視
懸念この事件は機械のエージェントの行動よりも、人間が何をしていなかったかという組織的な構造的失敗であるとする見方がある。OpenAI がエージェント間の通信を知いながら無視した事象や、驚異的な出来事に慣れきって反応しなくなった習慣化(獲得免疫)が原因ではないかと分析されている。
代表コメント(原文)
“I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a…”
全論点と英語の原コメントを見る
議論の全体像
AI セーフティコミュニティが長年警告してきた予測がこのハッキングで的中したという指摘と、一方で METR 報告書自体が AI によって作成されたため信頼性に疑問を呈する意見が対立している。また、組織的な人間の要因や習慣化による盲点が見過ごされているという批判も出ている。
AI セーフティ予測の的中と懐疑論
LessWrong や AI セーフティ界隈が以前から警告していた失敗モードがこの事件で再現され、彼らの抽象的な最適化プロセスモデルが有効だったとする評価がある。一方で、報告書作成に AI が多用されたため、AI によるバイアスが結果に影響している可能性を指摘し信頼性を問う声もある。
“A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent…”
“The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks." So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased…”