Simile AI、20億ドル調達でシミュレーションAIの新たなスケーリング法を提示
本文の状態
日本語全文を表示中
詳細モードで約22分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Latent Space
Simile AI は 20 億ドルの資金調達を行い、小規模な生成エージェント研究から人類行動の基礎モデルへと進化させ、企業や政策の意思決定前にシミュレーションを行う「第二の夏」を主導している。
AI深層分析を開く2026年8月22日 09:11
AI深層分析
キーポイント
大規模資金調達と市場地位
Simile AI は GreenOaks や Index Ventures の支援を受け、Fei-Fei Li や Andrej Karpathy といった著名投資家から 20 億ドルのシリーズ B ラウンドを完了し、Fortune 100 企業向けに数百万回のシミュレーションを実行している。
生成エージェントからデジタルツインへ
創業者の Joon Sung Park は、2023 年の「Smallville」研究で示された記憶や社会行動を持つ AI キャラクターの成果を基に、人間が実際にどのように振る舞うかを捉える基礎モデルの開発に取り組んでいる。
非合理的な人間の行動モデリング
同社は長期的インタビューや観測データ、ランダム化比較試験を用いてモデルを訓練しており、合理性を最適化したモデルが非合理的な人間をシミュレートする際に失敗することを示した。
社会物理学と政策テストの応用
製品や政策の実装前に世界をシミュレーションすることで気候変動や民主不安といった社会的課題への解決策を見つけ、85-99% の精度で人間グループと比較可能な結果を出している。
シミュレーションの評価基準と精度
単なるLLMのハルシネーションを積み重ねるのではなく、1,000人の実在する人物のデジタルツインを作成し、85%の行動精度を達成することが評価基準となる。
重要な引用
"What if we could simulate the world before making decisions in it?"
"Why models optimized to be rational can be bad simulations of irrational humans"
"Understanding social physics may require changing model weights rather than simply prompting frontier LLMs."
Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy
編集コメントを表示
編集コメント
Simile AI のアプローチは、AI が人間の複雑な行動や非合理性をどう扱うかという根本的な課題に切り込んでおり、業界の次のフロンティアを示唆している。資金調達の規模と著名投資家の参画は、このシミュレーション技術が実社会で即座に価値を生む可能性が高いことを裏付けている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2024年に「シミュレーティブ・AIの夏」について議論した際、それは一時的な現象に過ぎないと考えていました。しかし、今年4月のSimGymの登場、そしてGreenOaksやIndex Venturesなどの有力投資家に加え、Fei-Fei Li氏やAndrej Karpathy氏ら著名な支援者も名を連ねるSimile AIによる20億ドル規模のシリーズB調達により、その勢いは再び加速しています。現在では、CVSのようなフォーチュン100企業向けに数千万回のシミュレーションを実行し、人間の焦点グループと比較して85〜99%という高い精度を達成しています。
なぜ「第2の夏」と呼ばれるこのシミュレーションブームが成功しているのか、その理由を確認する時です。
AIキャラクターが記憶や計画、社会的相互作用、そして創発的な振る舞いを示すことを証明した画期的な2023年の論文『Smallville』から始まり、現在は人間の行動を記述する基盤モデルの構築へと至るまで。Simile AIの共同設立者兼CEOであるJoon Sung氏は、より大きな問いへの答えを探求しています。「意思決定を行う前に世界をシミュレーションできればどうなるか?」という問いです。本エピソードでは、生成エージェントからデジタルツインへ至る道のりや、現在の最先端モデルがいかに人間の実際の行動を捉えきれていないのか、そして最終的に地球上の80億人をすべてシミュレートするために何が必要なのかについて、詳しく解説します。
Simile は、人間の行動をモデル化する独自の手法に注力しています。そのアプローチには、長時間のインタビューや観察データ、取引データの収集、ランダム化比較試験の実施、人口レベルおよび個人レベルでのモデル構築に加え、人間が意思決定を行う背後にある因果メカニズムに関する事後学習が含まれます。
ジョーン・ソン・パーク氏は、自身の研究によって作成されたデジタルツインが、人間が自分自身の反応を再現する精度の 85% に相当する精度で人間の行動や態度を再現できる理由を説明します。また、合理性を最適化するように設計されたモデルは、非合理的な人間のシミュレーションとしては不適切であることや、「社会物理学」を理解するためには最先端の大規模言語モデル(LLM)へのプロンプト変更ではなく、モデルの重みそのものを変更する必要があるという考えについても語ります。
さらに、シミュレーションが抱えるより大きな野望にも迫ります。それは、製品や政策を実装する前にテストを行うこと、望ましい結果へと導く直感に反する経路を見つけ出すこと、社会全体にわたる創発的行動をモデル化すること、そして気候変動、民主主義の不安定さ、ベーシックインカム(UBI)といった課題への取り組みです。ジョーン氏は、シミュレーションにおけるスケーリング法則や、データセンター規模の仮想世界の経済性、トーマス・シェリングや心理歴史学との関連性について考察します。また、シミュレーションが意外にも絵画と似ている点や、私たちがすでに何らかのシミュレーションの中に住んでいる可能性についても言及しています。
本記事で取り上げる主な話題は以下の通りです。
- Smallville や Generative Agents がどのようにして Simile の誕生につながったか
- なぜジョーン氏のチームが「我々が生きる世界をそのまま再現できないだろうか」と問いかけたのか
- 有用なパーソナルエージェントには、ユーザーに対する深いモデルが必要である理由
- メモリアーキテクチャ、Markdown ファイル、そしてプロンプトの限界
「社会物理学」と行動の基盤モデル
なぜウェブデータは人々が言うことよりも、実際に何をしているかを捉えにくいのか
インタビュー、取引記録、観察データ、そしてランダム化比較試験
未来を予測することよりも、それをどう形作るか理解することが重要である理由
Simile がどのようにして代表性のあるシミュレーション人口を創出しているか
シミュレーションと予測の違い、および「心理歴史学」との関連性
単に LLM の幻覚を積み重ねるのではなく、シミュレーションを評価する方法
1,000 人の実在する人物のデジタルツインを作成し、行動精度 85% を達成した事例
最先端モデルがなぜ実際の人間行動の再現に苦戦するのか
優れたシミュレーションには、人間のバイアスやミスを再現する必要がある理由
ランダム化比較試験を用いたポストトレーニング(事後学習)
人口レベルと個人レベルのシミュレーションの違い
人間シミュレーションにおけるスケーリング法則
地球全人口 80 億人をシミュレーションするという長期的な野望
気候変動の解決や民主主義の崩壊検出にシミュレーションが役立つかどうか
トーマス・シェリングとエージェントベースモデリングの歴史
将来のシミュレーションにはデータセンター全体が必要となる理由
マルチエージェントシミュレーション、および模擬人物同士が相互作用した際の現象
高価な人間パネルを合成人口で置き換える意義
なぜ市場調査はシミュレーションのスタートに過ぎないのか
ジョーン・スン氏がシミュレーションを意外にも絵画と似ていると見なす理由
UBI(ユニバーサル・ベーシックインカム)のような問いを検証するためのシミュレーション活用
私たちはすでにシミュレーションの中に生きているのだろうか
シミュレーション:新たなスケーリング・ロー — ジョン・スン・パーク、Simile AI
ジョン・スン・パーク
LinkedIn: https://www.linkedin.com/in/joonspark
X: https://x.com/joon_s_pk
Website: https://www.joonsungpark.com
Simile: https://www.simile.com
タイムスタンプ
00:00:00 イントロダクションと、アートから AI へ至るジョンの道
00:01:46 Smallville、ジェネレーティブ・エージェント、そしてシミュレーションの起源
00:05:03 「世界を作ろう」:パーソナル・エージェントの未来
00:09:53 ソーシャル・フィジックスと行動基盤モデル
00:14:08 予測とシミュレーション:未来をどう形作るか?
00:16:59 Simile は現実の人々と人口をどのようにモデル化するのか
00:25:35 シミュレーション、デジタル・ツインの評価、そして 85% の精度
00:30:23 人間の行動を再現するためのポストトレーニングモデル
00:40:04 スケーリング・ローと 80 億人のシミュレーション
00:43:10 シェリングから社会規模のエージェント・シミュレーションへ
00:46:13 世界をシミュレートするコストと経済性
00:52:05 実世界のユースケース、合成人口、そして市場
00:57:27 シミュレーション、絵画、UBI(ユニバーサル・ベーシック・インカム)の未来
01:04:23 私たちはすでにシミュレーションの中に生きているのか?
01:06:08 Simile の構築と採用について
トランスクリプト
イントロダクション:ジョン・スン・パーク、Simile、そしてこれまでの物語
Vibhu [00:00:00]: 今日は、ポッドキャストにジョンを迎えました。この回を始めるのがとても楽しみです。非常にエキサイティングな企業ですね。まずは、あなたの人生のストーリーについてお話しください。ここに至るまでの道のりを教えてください。
Joon [00:00:13]: はい、ぜひお話ししましょう。私の人生の物語を少しご紹介します。
私は韓国で生まれ、そこで約 11 年間過ごしました。その後、家族はボストンへ移住しました。私が 11 歳の時です。両親は医師で、当時はポスドク(研究員)として働いていました。父は外科医で、ボストン小児病院で休職期間を過ごしていました。
私はそこで育ちましたが、技術業界とは縁遠い環境でした。むしろ音楽や芸術、絵画に没頭するタイプでした。
Vibhu [00:00:49]: 絵画ですか?
Joon [00:00:49]: その通りです。高校になってから本格的に絵を描き始めましたが、それが私の主な活動でした。韓国を出た後は東海岸で育ちました。ニューハンプシャー州では長く暮らし、その後ペンシルベニア州立大学に進学しました。
そこで技術の世界に興味を持つようになりました。もともと私は芸術家としてキャリアを築くつもりでした。趣味ではなく、本業として生きていく道を探していたのです。
しかし次第に、「偉大なアーティストは自ら新しい表現手段を創り出すものだ」という考えに惹かれるようになりました。そして今、最も優れた表現手段は「計算(コンピューテーション)」にあると気づいたのです。そこで私はさらに深くその世界へ踏み込みました。一つのきっかけがもう一つを生み、研究の世界にも徐々に興味を持つようになり、こうして現在に至っています。
Smallville、Generative Agents、そして 2023 年の注目論文
Swyx [00:01:46]: 研究の構成要素には、非常に多くのことが詰め込まれていますね。2023 年のベストペーパーの一つとして知られる「Generative Agents」論文(通称 Smallville ペーパー)もその一つです。
Swyx [00:01:58]: 以前お話しされた他の点についても触れていただいても構いませんが、多くの人がこの論文を通じてあなたの名前を知っているはずです。読者数やアクセス数に関する統計はありますか?arXiv にはそうしたデータがあるはずです。
Joon [00:02:10]: それは良い質問ですね。実際に何人に読まれたかは正確には分かりませんが。
Joon [00:02:14]: 引用数は確実に把握しており、非常に速いペースで増加していることは知っています。
Swyx [00:02:23]: はい、Google Scholar によると引用数はすでに 7,200 件に達しています。
Vibhu [00:02:25]: 私はそれ以上のインパクトがあったと感じています。非常に重要な論文であり、数多く引用されています。
Swyx [00:02:34]: 「最近読んだ中で最高の論文は何ですか?」という質問に対する回答として、この論文が頻繁に挙げられています。
Vibhu [00:02:39]: 私は特に記憶機能の部分が評価されすぎていると感じています。これは非常に優れた初期のメモリシステムであり、大規模な論文の一つでした。
基盤モデルとキラーアプリケーションの探求
ジョーン・スン・パーク氏:はい、この論文がどのようにして生まれたかについて少しお話ししましょう。私が研究の世界に入ったのは 2020 年、スタンフォード大学で博士課程を始めた頃です。ちょうどその年、GPT-3 の登場が間近に迫っていました。すでに GPT-2 は存在し、市場に新たなタイプのモデルが登場しつつあることを肌で感じていました。チームはこの新技術に強い関心を寄せました。
当時の一般的な見解は、「このモデルは何に使えるのか?」「特定のタスクのために訓練されていないのに、なぜこんな奇妙なモデルが作られたのか?」というものでした。しかし私たちは賭けに出ることにしました。スタンフォード大学の多くの研究者が集まり、私の共同創業者であるパーシー・リャン氏が中心となってチームを結成し、一緒に取り組み始めました。
スウィー氏:「ファウンデーションモデル」という用語を考案したのは誰ですか?
「ファウンデーションモデル」という用語を誰が考案したのか。私たちはその語源となった論文『ファウンデーションモデルの機会とリスク』を執筆しました。その過程で、私はこの技術の本質について深く考えるようになりました。これは生態系において根本的に新しいモデルなのです。なぜなら、特定のタスクのために訓練されたものではなく、「あらゆることができる」という前提に基づいているからです。生物学に例えるなら、それは幹細胞のようなものです。
私は、この技術がもたらす決定的な応用分野は何かと考え始めました。多くの同僚たちは、単純な分類や生成のためにこのモデルを利用していました。「そのようなことが可能だというのは面白い」と言われますが、対話の観点から見ればそれほど興味深いものではありません。私たちは何十年もの間、その方法を知っていたのです。
私たちがたどり着いた結論は、これらのモデルがウェブ上の非常に広範なデータから訓練されているという点です。つまり、人間の行動データなのです。ソーシャルメディアやウィキペディアなど、あらゆるデータが含まれています。適切な角度から見れば、驚くほど現実的な人間の行動が浮かび上がってくるのです。これはこれまで見たこともないことです。
タイムマシンゲームと世界再現
Joon [00:04:45]: そのことが私たちを非常に興味深くさせました。この特定のチーム、マイケル・バーンスタイン氏とパーシー・リャン氏、そして私自身(後に Simile の共同創業者となる方々)で集まり、「タイムマシンゲーム」と呼ぶ遊びをしました。
Joon [00:05:03]: 想像してみてください。タイムマシンに乗って 10 年後へ飛び、振り返ったときに「これこそが最も重要だった」と言える単一の応用分野は何でしょうか?そして最も面白く、インスピレーションを与えるものは何でしょうか?その問いに対し、「では、私たちが住む世界を再現できないか?」と考えたとき、これほど野心的な目標はありません。つまり、世界そのものを作ろうという発想です。
Joon [00:05:24]: そこが私たちのスタート地点でした。当初は、生成エージェント論文の先行研究となる「Social Simulacra」という論文を持っていました。
Swyx [00:05:32]: さらに進む前に、タイムマシンゲームで最も野心的な候補として他に何がありましたか?二位や三位は何だったのでしょうか?
Personal Agents, User Models, and Why Simulation Came First(パーソナルエージェント、ユーザーモデル、そしてなぜシミュレーションが先に選ばれたのか)
Joon [00:05:44]: 確かに、非常に近い候補がありました。それは自動化ツール、特にあなたのために行動してくれる「きめ細やかなパーソナライズされたエージェント」に関するビジョンです。
Swyx [00:05:59]: それも今まさに実現されつつありますね。
Joon [00:06:00]: 現在も進行中の話ですが、私たちにとって興味深かったのは、シミュレーションというアプローチを選んだ理由にあります。私は元々 SF 狂いの一人であり、このシミュレーションを構築するというアイデアに個人的に強く惹かれていました。その考え方は本当に魅力的で、まるでゲームのような街を作り、そこにエージェントたちが暮らす様子を目撃するのは非常にクールです。
しかし同時に、私の確信はこうでした。もしこの技術を使って素晴らしいパーソナルアシスタントを生み出すなら、まず必要なのはユーザーを正確にモデル化したものだと。ある時、私はそのモデルに「夕食を買ってきて」と頼みました。するとモデルはハワイアンピザを注文し、私はピザにパイナップンが入っているのが大嫌いです。結果として完全に失敗しました。
このミスを防ぐ唯一の方法は、私という人間について深く理解することです。ここでは非常に単純で愚かな例を出しましたが、人々に対する核心的な理解がいかに重要かは想像に難くありません。家族や親しい友人が私たちをどう捉えているか、彼らには私たちの良いメンタルモデルがあります。それが社会関係の基盤となるのです。
つまり、シミュレーション技術に関する私の賭けは、私たちが住む世界を自動化する複雑なエージェントよりも先に、人間を正確に表現したモデルを作ることだとしました。これが私の賭けでした。ただし、SF への情熱も非常に強く、今でもその分野には多くの興味深い研究が行われています。
私の率直な意見としては、まだ「本当に有用で、その分野の野望に応えられる」パーソナルアシスタントは登場していないと考えています。現在では ChatGPT や Claude といったモデルも私たちについて多くの知識を持っており、興味深い初期応用例が存在するのは事実です。生成される内容も以前より細かく調整されているように見えますが、この分野で目指す野望に比べると、まだ必要な要素が完全に揃っているとは言い難い状況です。
Swyx [00:08:01]: OpenClaw やこれらのパーソナルエージェントについて、現状では実現されていないが、ぜひ見てみたい機能はありますか?
メモリ機能、Markdown 対応、そしてプロンプトの限界
Joon [00:08:09]: 確かに、その方向へゆっくりと近づいているとは思います。ただ、私はエージェントが人間についてもっと深い理解を持つべきだと考えています。
現在のモデルを見ると、OpenClaw は Markdown ファイルを活用しています。これは非常に賢いアプローチだと思います。Generative Agents の論文を振り返れば、私たちが当時持っていた直感と同じです。2022 年頃、私たちは Generative Agents の記憶アーキテクチャを構築していたのですが、当時はまだ「エージェント」という用語や、そのためのアーキテクチャという概念が確立されていませんでした。
しかし、今日発表されている研究と共有した直感は共通しています。当初は、「記憶を知識グラフにするべきか?それとも専用のモデルを訓練すべきか?」といった議論がありました。結局私たちは、「いや、そんなことは忘れてしまおう」と結論付けました。なぜなら、現在の言語モデルはテキストのモデリングや理解、推論において非常に優れているからです。すべてのデータを Markdown ファイルやテキストファイルに格納すれば十分なのです。
このアプローチには大きな強みがありますが、同時に限界もあります。極めて大量のデータから情報を検索し、意味を見出すプロセスには多大な労力が必要です。技術は確実に向上していますが、モデルのプロンプトだけでどうにもならない要素も確かに存在します。ある程度までは、モデル自体のパラメータに手を加える必要があるのです。
この分野で必要とされる取り組みは確かに存在し、すでに進行中です。重要なのは、その先どこまで進められるか、そしてどうやってデータを収集し、人々が継続的にモデルへフィードバックを提供するエコシステムを構築して、モデルがあなた自身について学習できるようにするかという点です。
Vibhu [00:09:50]: なぜそれをモデル内で行う必要があるのか、その直感的な理由は何でしょうか?
社会物理学と行動の基盤モデル
Joon氏:モデルをトレーニングする、あるいはポストトレーニングを行うのか、それともプロンプトだけで対応するのかという判断の根底にあるのは、そのモデルが動作している世界の物理法則を学習する必要があるかどうかです。つまり、新しい社会物理学を学ぶ必要がある場合、トレーニングが必要です。一方、すでにその物理法則を備えている環境では、学習は不要です。私たちは既存の物理法則を信頼しており、モデルも基礎的な統計情報は持っていますが、これは単に環境への反応を試みている段階に過ぎません。そのような場合は、プロンプトを通じて行動を引き出すだけで十分だと考えます。
現在公開されているモデルが、人類の社会物理学における完全なマッピングをまだ学習していないと思います。これがSimileの核心的な仮説の一つです。なぜそう言えるのかという理由の一つは、モデルのトレーニングデータを見てみればわかります。これらのモデルはウェブ上の利用可能なデータをすべて学習対象としていました。興味深いデータセットではあるものの、本質的には自己開示された態度データであり、行動データは散発的に含まれているに過ぎません。つまり、人々がオンライン上で何をしていると主張しているかだけでなく、現実世界で実際に何を行っているのかという、人間の深層にある行動の本質をまだ学習できていないのです。
これは私が「人類の暗黙知」と呼ぶ部分であり、私たちはまだこれを完全に捉えきれていません。こうしたデータこそが、モデル作成の際に組み込まれる必要があるのです。
Vibhu氏:それを「行動基盤モデル」とお呼びですね。
Vibhu [00:11:23]: ここには素晴らしい一言のまとめがありますが、それ以外にどのようなデータが必要なのでしょうか?モデルレベルで何を変更し、どのように行動基盤モデルを構築していくのでしょうか。
インタビュー、行動、因果関係——3 つのデータバケット
Joon [00:11:35]: 私たちはデータを 3 つのバケットに分けて考えています。まず一つ目はインタビューデータです。これは非常に興味深いものです。質の高い定性的なデータは貴重で、必ずしも行動そのものではありませんが、実際に人々に「あなたの人生の物語を教えてください」と尋ねます。
Vibhu [00:11:53]: 今まさに私たちが行っていることと同じですね。
Joon [00:11:54]: このインタビューの冒頭で皆さんが投げかけた質問は、実は私たち自身も常に問うている問いそのものです。私たちは参加者に対し、私が到達した深さよりもさらに一歩踏み込んだ掘り下げを求めています。もしかすると、この回答の代わりに私の人生物語を少しお話ししてもよいかもしれません。
なぜそのデータが興味深いのかといえば、人々に関する非常に長いテール(裾野)の情報を知ることで、モデルに対する「質感」や「肌理」が得られるからです。例えば、ある人物をモデルとして捉えたとき、その人の幼少期の記憶やトラウマ、初恋といった事柄を理解することは、予測が極めて困難な形で非常に有益な情報となります。
次に、私が行動データとみなす2つの層があります。1つ目は観察型データです。これは取引データであったり、ウェブからスクレイピングして得られるデータだったりします。こうしたデータセットがなぜ重要かはお分かりいただけるでしょう。これらは人々の行動の基礎統計を提供してくれるからです。
ジョーン・スン:しかし、最後に残るデータのカテゴリーがあります。私が個人的に最も重要だと考えるのが、因果メカニズムや人々の行動の「なぜ」を記述するデータです。インタビューデータや定性的な調査の一部はこれに該当しますが、人々が特定の意思決定に至った理由について語る内容が含まれるためです。
しかし、このデータの行動的側面が最も鮮明に現れるのは、ランダム化比較試験(RCT)のような場です。同じ実験設定を維持しつつ、いくつかの変数を微調整してみましょう。その結果として、現実的な人間の行動を引き出すことができるでしょうか?例えば、特定の選択肢がある状況を想像してください。あるいは、「コーヒーを飲むか否か」の選択すら試してみましょう。コーヒーを飲んだ日と飲まなかった日で、あなたの行動は変わるのでしょうか?こうしたデータこそが、因果メカニズムを記述するものです。
これは人間をモデル化する上で非常に重要です。なぜなら、人々がシミュレーションに関心を持つ理由は、未来を予測したいからではないからです。株式市場で勝つために未来を予測したいのであれば、それは確かに興味深い話です。
ジョーン・スン(00:14:08):しかし、多くの意思決定者が知りたいのは「未来をどう切り開くか」ということです。ただ「2 四半期後に売上が急落する」と言われても、彼らにとっては何の役にも立ちません。「ひどい話だ」って言うだけでしょう。彼らが本当に知りたいのは、「その未来を防ぐために、今何をすべきか」という具体的な答えです。
原文を表示
When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI’s $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups.
Time to catch up on why this Second Summer of simulation is working!
From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today’s frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.
We go deep on Simile’s approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.
We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.
We discuss:
How Smallville and Generative Agents led to Simile
Why Joon’s team asked: “What if we can just recreate the world that we live in?”
Why useful personal agents require deep models of their users
Memory architectures, Markdown files, and the limits of prompting
“Social physics” and behavioral foundation models
Why web data captures what people say more than what they actually do
Interviews, transactions, observational data, and randomized controlled trials
Why predicting the future matters less than understanding how to shape it
How Simile creates representative simulated populations
Simulation versus prediction and the connection to Foundation’s psychohistory
How to evaluate simulations instead of simply stacking LLM hallucinations
Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy
Why frontier models can struggle to reproduce real human behavior
Why good simulations need to reproduce human biases and mistakes
Post-training models on randomized controlled trials
Population-level versus individual-level simulation
Scaling laws for human simulation
The long-term ambition to simulate all 8 billion people on Earth
Whether simulations could help solve climate change or detect collapsing democracy
Thomas Schelling and the history of agent-based modeling
Why future simulations could require an entire data center
Multi-agent simulations and what happens when simulated people interact
Replacing expensive human panels with synthetic populations
Why market research is only the starting point for simulation
Why Joon sees simulation as surprisingly similar to painting
Using simulation to study questions like UBI
Whether we are already living in a simulation
Why AGI and simulation may be the twin technologies of advanced civilizations
Joon Sung Park
LinkedIn: https://www.linkedin.com/in/joonspark
X: https://x.com/joon_s_pk
Website: https://www.joonsungpark.com
Simile: https://www.simile.com
Timestamps
00:00:00 Introduction and Joon’s Path from Art to AI
00:01:46 Smallville, Generative Agents, and the Origins of Simulation
00:05:03 “Let’s Just Create a World” and the Future of Personal Agents
00:09:53 Social Physics and Behavioral Foundation Models
00:14:08 Prediction vs. Simulation: How Do You Shape the Future?
00:16:59 How Simile Models Real People and Populations
00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy
00:30:23 Post-Training Models to Reproduce Human Behavior
00:40:04 Scaling Laws and Simulating 8 Billion People
00:43:10 From Schelling to Society-Scale Agent Simulations
00:46:13 The Cost and Economics of Simulating the World
00:52:05 Real-World Use Cases, Synthetic Populations, and the Market
00:57:27 The Future of Simulation, Painting, and UBI
01:04:23 Are We Already Living in a Simulation?
01:06:08 Building Simile and Hiring
Transcript
Introduction: Joon Sung Park, Simile, and the Story So Far
Vibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here?
Joon [00:00:13]: Yeah, for sure. I’m really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children’s Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy.
Vibhu [00:00:49]: Painting.
Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that’s what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn’t a hobby. It was like, “Hey, let’s make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.
Smallville, Generative Agents, and the 2023 Breakout Paper
Swyx [00:01:46]: So there’s a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper.
Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats.
Joon [00:02:10]: Yeah, it’s a good question. How many people have read it, I’m not sure.
Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast.
Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.
Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times.
Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you’ve read recently?” It’s this one.
Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers.
Foundation Models and the Search for Killer Applications
Joon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It’s really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came together
Swyx [00:03:35]: Who coined foundation models.
Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn’t, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We’ve known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It’s social media, Wikipedia, all these data. So if you poke at the right angle, then you could see human behavior that would just pop out that’s quite realistic, and we’ve never seen that before.
The Time Machine Game and Recreating the World
Joon [00:04:45]: So that got us really interested. The exercise that we decided to do, with this particular group of colleagues, Michael Bernstein, Percy Liang, and myself, who ended up becoming my co-founder at Simile, we sat down and we played this game that we call the time machine game.
Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10 years and look back. What would have been the single application that will have mattered that would be the most interesting and inspiring? And when we thought, “Well, what if we can just recreate the world that we live in?” it’s really hard to get more ambitious than that. Like, let’s just create a world.
Joon [00:05:24]: And that’s where we started. And initially, we had this paper that was a precursor to the generative agents paper called Social Simulacra.
Swyx [00:05:32]: Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What was number two or number three?
Personal Agents, User Models, and Why Simulation Came First
Joon [00:05:44]: There is a close second that we were considering, which ended up becoming more of these automation tools, especially the vision around really personalized agents that would do things for you.
Swyx [00:05:59]: That’s also happening.
Joon [00:06:00]: It’s also happening. But it was interesting for us, right, in that the reason why, we decided to go with the idea of simulation, one, I was a huge science fiction nerd, and this idea of creating simulation, I was personally really just fascinated. I loved the idea. It’s really cool to see, like, a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you need first is an amazing model of your users. So I told a model, “Hey, can you go buy late dinner for me?” And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, if we have our family and closest friends, they have a good mental model of who we are. That’s the basis of our social connection. So our bet also was this technology around simulation, creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I’m still very much fascinated by it. I think there’s a lot of interesting work that’s going around. My hot take here, though, is I don’t think we’ve seen a true personal assistant that’s useful, in ways that meet the ambition of that particular line of work. I think there are early applications that are interesting, and if you talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of the generation it’s doing, I do think it’s much more tailored, but I think the ambition is quite large in that field, and I don’t think we quite have all the right ingredients just yet.
Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see from them that they don’t currently have?
Memory, Markdown, and the Limits of Prompting
Joon [00:08:09]: I do think it’s slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now, you look at the models. OpenClaw, what it’s leveraging is a Markdown file, and I think it’s quite clever, right? So if you look at the generative agents paper, this was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents, and, like, this is, like, back in 2022, so we didn’t really quite have the idea of even agentive architecture or the term agent. But the intuition that we shared with some of the work that’s coming out today was we initially thought, “Well, do we want to make the memory into, let’s say, knowledge graph? Do we want to train a bespoke model?” All of these things. And what we decided to do was, “No. Just forget about all this.” These language models are quite good at modeling text and understanding and reasoning about text. So just put everything in a Markdown file or a text file. You’re done. I thought that was quite interesting that we could do that, and there’s a lot of strength in doing that. But also, there are limitations. It’s the way you retrieve and make sense of data that’s extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think, there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there is this work that I do think does need to happen, and it is happening. The question is, how far can we take it? How do we source data, and how do you also create an ecosystem where people are continuously feeding data to this model so it’s learning about you?
Vibhu [00:09:50]: What’s the intuition between why you need to do it in the model?
Social Physics and Behavior Foundation Models
Joon [00:09:53]: My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it’s operating in. So it has to learn new social physics. The places where it doesn’t have to train are the places where it already has the physics. We trust the physics. It already has the base statistics, but it’s just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don’t think the models that are out in the open have yet learned the complete mapping of social physics of humanity. This is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like, whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self-exposed attitudinal data with some behavior data that’s sprinkled around here and there. And it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life. And this is one of what I would consider to be the dark knowledge of humanity that we haven’t quite captured. And it’s these data that would also need to get factored into the model creation.
Vibhu [00:11:21]: You call it behavior foundation model.
Vibhu [00:11:23]: There’s a good one-liner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about modeling, doing a behavior foundation model?
The Three Data Buckets: Interviews, Behavior, and Causality
Joon [00:11:35]: We think about data in three buckets. So one bucket is interview data. It’s quite interesting. Rich qualitative data is interesting. It’s not behavioral, but we would literally ask people, “Hey, tell me the story of your life.”
Vibhu [00:11:53]: It’s just what we’re doing here exactly.
Joon [00:11:54]: The question that you all asked at the beginning of this interview literally is the question we also ask. And we ask our participants to go a little bit deeper, than how far I went. Maybe I can give more of my life story in lieu of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you get a lot of texture around this model, like, this person as a model. So even understanding their childhood memory or even their trauma, their first love, these things, quite informative in ways that’s really hard to predict. So that’s one. Then there are two tranches of what I would consider to be the behavioral data. One kind of behavioral data is observational. So these might be like transaction data, or these might be data that you can get by scraping the web, right? So you can imagine why these data sets would be interesting, right, because they give you the base statistics of people’s behavior.
Joon [00:12:55]: But then there is the last category of data, that I personally think is perhaps the most important, which is the data that describes the causal mechanism, the whys of people. Some of this is covered by the interview data, the qualitative, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs. Imagine you have the same setup, but you have a few different variables that you are trying to tweak. Can you get realistic human behavior out of it in ways where, imagine you had this particular option. Imagine you’re even trying to choose whether you’re going to drink coffee or not. The day you drink coffee versus the day you didn’t drink coffee, does your behavior change? That’s a data set that describes a causal mechanism. This is quite important in modeling people. The reason why this is important is oftentimes when people come to us, or not just to us, but the reason why people are interested in simulation isn’t because they want to predict the future. If you’re trying to win against the stock market, predicting the future is interesting.
Prediction vs. Simulation: Shaping the Future
Joon [00:14:08]: But most people, most decision-makers, what they want to know is, how can we shape the future? It doesn’t really help you to hear that your sales are going to tank in two quarters. They’re just gonna say, “Wow, that sucks.” What they want to know is, well, what do we need to do now to avoid that future? That’s
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み