Sakana AI、AI の創造性実験を発表
Sakana AI と MIT、NYU の共同研究により、明確な目標設定なしに画像を生成する「Picbreeder」の再現実験が行われ、AI エージェントが人間の創造的飛躍には及ばない一方、多様なエージェント人口を導入することで探索能力が向上することが示された。
キーポイント
目標のない進化における AI の限界
明確なゴールや評価基準がない環境下では、AI エージェントは既存のアイデアを微調整する傾向が強まり、人間のような予期せぬ概念的飛躍を起こすことが難しいことが示された。
多様なエージェント人口による探索の向上
異なる性格や特性を持つ多様なエージェント群を導入することで、人間のアーカイブに匹敵する意味的多様性やバランスの良い進化樹が生成される可能性が確認された。
AI 進化と勾配降下法の比較
AI エージェントによって進化した画像(例:頭蓋骨)は、ニューラル表現を乱した際に変化が滑らかで、勾配降下法で直接最適化されたものよりも堅牢な表現を持つことが示唆された。
人間の創造性の本質的差異
人間は偶然の発見を「価値あるもの」と認識し、それを追求して次の段階へ飛躍させる能力に長けているが、AI は面白いパターンに陥ってそこから抜け出せない傾向がある。
多様な人格による探索の改善
多様な人格を持つエージェント集団を導入することで、AI の探索能力が大幅に向上し、生成されたアーカイブの意味的な幅広さが人間のものに迫る水準に達しました。
洗練化への偏りと人間の飛躍
VLM エージェントは既存のアイデアを洗練させる傾向が強いため、予期せぬ発見には至りにくく、人間のように偶然の産物を価値ある創造へと発展させる能力が不足しています。
オープンエンドな知性の未解明
現在の AI システムがなぜ人間のようにならばないオープンエンドな探索を行えないのか、その根本的な欠落要因についてはまだ十分に理解されていません。
重要な引用
Humans appear better at turning fortunate accidents into sustained creative discoveries: recognizing when something unexpected is worth pursuing, refining it, and then making a larger conceptual leap.
Compared with humans, VLM agents tend to keep circling back to the same kinds of images and concepts.
Introducing a diverse population of agent personalities substantially improves exploration.
VLM エージェントは特定の見た目や意味に引き寄せられやすく、既存のアイデアを捨てて予期せぬ何かを探すよりも、手元にあるものを洗練させることに留まりがちでした
人間は、偶然の産物を持続的な創造へとつなげることに長けています。あるものを見つけたとき、その価値を感じ取り、それを追いかけることで、より大きな概念的飛躍を達成できます
影響分析・編集コメントを表示
影響分析
この研究は、単なる画像生成の技術的進歩を超え、「目標設定なしでの創造性」という根本的な課題に AI が直面していることを浮き彫りにしました。AI エージェントが人間の創造性の一部を模倣できる可能性を示しつつも、その限界を明確にしたことで、今後のオープンエンドな探索システムや多様性を重視したエージェント設計において重要な指針となるでしょう。
編集コメント
明確な目標設定が逆に創造性を阻害するという「目標という幻想」の議論を、最新の VLM エージェントを用いた実験で再検証した点は非常に示唆に富んでいます。AI が人間の創造性のどこまで到達できるか、そして何が欠落しているかを定量的・質的に分析するこのアプローチは、今後の AI 開発の方向性を考える上で極めて重要です。
(*日本語は英文の後に)
In our new GECCO 2026 paper, “In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models”, in collaboration with MIT and NYU, we revisit Picbreeder, a lost website where people collaboratively evolved images without any predefined objective. Users simply selected images they found interesting, allowing unexpected forms such as faces, animals, vehicles, and skulls to emerge gradually across many generations and many different people.
We recreated this process using vision-language model agents. The agents explore a shared archive, choose images to branch from, evolve new candidates, publish their favorites, and evaluate the creations of other agents. There is no target image and no explicit definition of what counts as progress.
The results reveal both the promise and current limitations of AI-driven open-ended discovery.
Compared with humans, VLM agents tend to keep circling back to the same kinds of images and concepts. They repeatedly select similar parents, make smaller conceptual leaps, and often refine an existing idea rather than abandoning it in search of something genuinely unexpected.
However, introducing a diverse population of agent personalities substantially improves exploration. In some runs, diverse agent populations approached or matched the human archive on measures of semantic diversity and produced more balanced evolutionary trees.
We also find intriguing evidence that open-ended evolution can produce more robust representations. A skull evolved by the agents changes smoothly when its underlying neural representation is perturbed, less fractured than a skull directly optimized with gradient descent, although still less cleanly disentangled than one evolved collectively by humans.
But perhaps the most interesting result is the gap that remains.
Humans appear better at turning fortunate accidents into sustained creative discoveries: recognizing when something unexpected is worth pursuing, refining it, and then making a larger conceptual leap. The AI agents often notice interesting patterns too, but are more likely to become trapped in them.
We still do not fully understand what enables humans to navigate open-ended search in this way, or what ingredient(s) current AI systems are missing. For now, the results suggest that there remains something important about human creativity that AI agents have not yet learned to reproduce.
This paper will be presented at GECCO 2026 and is nominated for a best paper award! Please check out the interactive blog and technical paper for more details!
Technical Blog: https://pub.sakana.ai/picbreeder-vlm
Full Paper: https://arxiv.org/abs/2605.23908
Japanese
VLMは人間のような創造性を持てるか?
ケネス・スタンレー教授らの『目標という幻想(Why Greatness Cannot Be Planned)』は、明確な目標を設定することが、かえって真に偉大な発見を遠ざけてしまうという逆説を論じた書籍です。その議論の中核にあったのが「PicBreeder」の実験でした。
PicBreeder では、ユーザーが「面白い」と感じた画像を選び、それを少しずつ進化させていきます。事前に決められたゴールはなく、人々が「なんとなく良い」と思ったものを選び続けるだけで、顔や動物、乗り物、頭蓋骨といった予期しない形が、何世代もかけて、多くの人の手を経て自然と現れます。スタンレー教授らは、こうした「オープンエンド」、つまり目標をあらかじめ定めない探索こそが人間の創造性の根幹にあると考えたのです。
では、このオープンエンドな探索を、AIは再現できるのでしょうか。
MIT・NYUとの共同研究として発表する「In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models」では、視覚言語モデル(VLM)エージェントによる再現を試みました。エージェントたちは共有アーカイブを探索し、画像を選んで進化させ、気に入ったものを公開し、他のエージェントの作品を評価します。目標となる画像も「進歩」の定義も与えられていません。
その結果、AIによるオープンエンドな発見の可能性と限界の両方が浮かび上がりました。VLMエージェントは特定の見た目や意味に引き寄せられやすく、既存のアイデアを捨てて予期せぬ何かを探すよりも、手元にあるものを洗練させることに留まりがちでした。一方で、多様な人格を持つエージェント集団を導入すると探索は大きく改善され、生成されたアーカイブの意味的な幅広さは、人間が作ったアーカイブに迫る水準にまで達しました。
しかし、VLMでは届かなかった点もありました。人間は、偶然の産物を持続的な創造へとつなげることに長けています。あるものを見つけたとき、その価値を感じ取り、それを追いかけることで、より大きな概念的飛躍を達成できます。AIエージェントも興味深いパターンに気づくことはできても、そのパターンに囚われてしまう傾向がみられました。
なぜ人間は、本研究のVLMにはできなかったオープンエンドな探索を進められるのか。現在のAIシステムに何が欠けているのか。私たちはまだ十分に理解できていません。ここには、当社が探求するAI駆動型科学研究にも通じる大きな問いが残っています。Sakana AIは今後も、オープンエンドな知性の探求を深めていきます。
ブログ:https://pub.sakana.ai/picbreeder-vlm
論文:https://arxiv.org/abs/2605.23908
原文を表示
(*日本語は英文の後に)
In our new GECCO 2026 paper, “In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models”, in collaboration with MIT and NYU, we revisit Picbreeder, a lost website where people collaboratively evolved images without any predefined objective. Users simply selected images they found interesting, allowing unexpected forms such as faces, animals, vehicles, and skulls to emerge gradually across many generations and many different people.
We recreated this process using vision-language model agents. The agents explore a shared archive, choose images to branch from, evolve new candidates, publish their favorites, and evaluate the creations of other agents. There is no target image and no explicit definition of what counts as progress.
The results reveal both the promise and current limitations of AI-driven open-ended discovery.
Compared with humans, VLM agents tend to keep circling back to the same kinds of images and concepts. They repeatedly select similar parents, make smaller conceptual leaps, and often refine an existing idea rather than abandoning it in search of something genuinely unexpected.
However, introducing a diverse population of agent personalities substantially improves exploration. In some runs, diverse agent populations approached or matched the human archive on measures of semantic diversity and produced more balanced evolutionary trees.
We also find intriguing evidence that open-ended evolution can produce more robust representations. A skull evolved by the agents changes smoothly when its underlying neural representation is perturbed, less fractured than a skull directly optimized with gradient descent, although still less cleanly disentangled than one evolved collectively by humans.
But perhaps the most interesting result is the gap that remains.
Humans appear better at turning fortunate accidents into sustained creative discoveries: recognizing when something unexpected is worth pursuing, refining it, and then making a larger conceptual leap. The AI agents often notice interesting patterns too, but are more likely to become trapped in them.
We still do not fully understand what enables humans to navigate open-ended search in this way, or what ingredient(s) current AI systems are missing. For now, the results suggest that there remains something important about human creativity that AI agents have not yet learned to reproduce.
This paper will be presented at GECCO 2026 and is nominated for a best paper award! Please check out the interactive blog and technical paper for more details!
Technical Blog: https://pub.sakana.ai/picbreeder-vlm
Full Paper: https://arxiv.org/abs/2605.23908
Japanese
VLMは人間のような創造性を持てるか?
ケネス・スタンレー教授らの『目標という幻想(Why Greatness Cannot Be Planned)』は、明確な目標を設定することが、かえって真に偉大な発見を遠ざけてしまうという逆説を論じた書籍です。その議論の中核にあったのが「PicBreeder」の実験でした。
PicBreeder では、ユーザーが「面白い」と感じた画像を選び、それを少しずつ進化させていきます。事前に決められたゴールはなく、人々が「なんとなく良い」と思ったものを選び続けるだけで、顔や動物、乗り物、頭蓋骨といった予期しない形が、何世代もかけて、多くの人の手を経て自然と現れます。スタンレー教授らは、こうした「オープンエンド」、つまり目標をあらかじめ定めない探索こそが人間の創造性の根幹にあると考えたのです。
では、このオープンエンドな探索を、AIは再現できるのでしょうか。
MIT・NYUとの共同研究として発表する「In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models」では、視覚言語モデル(VLM)エージェントによる再現を試みました。エージェントたちは共有アーカイブを探索し、画像を選んで進化させ、気に入ったものを公開し、他のエージェントの作品を評価します。目標となる画像も「進歩」の定義も与えられていません。
その結果、AIによるオープンエンドな発見の可能性と限界の両方が浮かび上がりました。VLMエージェントは特定の見た目や意味に引き寄せられやすく、既存のアイデアを捨てて予期せぬ何かを探すよりも、手元にあるものを洗練させることに留まりがちでした。一方で、多様な人格を持つエージェント集団を導入すると探索は大きく改善され、生成されたアーカイブの意味的な幅広さは、人間が作ったアーカイブに迫る水準にまで達しました。
しかし、VLMでは届かなかった点もありました。人間は、偶然の産物を持続的な創造へとつなげることに長けています。あるものを見つけたとき、その価値を感じ取り、それを追いかけることで、より大きな概念的飛躍を達成できます。AIエージェントも興味深いパターンに気づくことはできても、そのパターンに囚われてしまう傾向がみられました。
なぜ人間は、本研究のVLMにはできなかったオープンエンドな探索を進められるのか。現在のAIシステムに何が欠けているのか。私たちはまだ十分に理解できていません。ここには、当社が探求するAI駆動型科学研究にも通じる大きな問いが残っています。Sakana AIは今後も、オープンエンドな知性の探求を深めていきます。
ブログ:https://pub.sakana.ai/picbreeder-vlm
論文:https://arxiv.org/abs/2605.23908
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み