最先端モデルファクトリーの構築法
TLDR AI は、最先端のAIモデルを効率的に開発・運用するための「ファクトリー」構築手法について、48分にわたる詳細な解説記事を公開した。
AIニュース価値スコアβ
主要ニュースAI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 50
- 新規性
- 75
- 検索具体性
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
記事は特定の最新モデル(Laguna S 2.1)と独自のアーキテクチャ(思考モード/非思考モード切り替えなど)という具体的な新事実を含んでおり、新規性が高い。ただし、日本企業や日本固有の規制・価格情報などは含まれていないため、日本の関連性は低めとなる。
キーポイント
ファクトリーの定義と必要性
単なる実験環境ではなく、再現性、スケーラビリティ、コスト効率を担保するための体系的な開発・運用基盤の構築が現代のAI開発に不可欠であると説く。
データパイプラインの最適化
データの収集、クリーニング、アノテーションからトレーニングセットへの統合に至るまでの自動化されたワークフローの重要性を強調している。
インフラと計算リソース管理
大規模モデル訓練におけるGPUクラスタの効率的な利用、コスト削減のためのハイブリッド戦略、およびトレーニング中の監視・トラブルシューティング手法を詳述する。
重要な引用
Building a frontier model factory is not just about buying more GPUs; it's about creating a repeatable, scalable system for innovation.
The bottleneck in modern AI development has shifted from compute to data quality and pipeline efficiency.
影響分析・編集コメントを表示
影響分析
この記事は、AI開発が単なる実験の域を超え、大規模な工業化プロセスへと移行していることを示唆しており、開発チームや組織にとって実務的な指針となる重要な内容です。特にデータパイプラインとインフラ管理の最適化に焦点を当てている点は、コスト増が懸念される現状において、競争力を維持するための具体的な解決策を提供しています。
編集コメント
「ファクトリー」という概念を提示し、AI開発の工業化プロセスを体系化した点は非常に示唆に富んでいます。特に、計算リソースの獲得競争だけでなく、データとパイプラインの質が次のボトルネックとなっているという指摘は、現場の開発者にとって即座に実践可能な洞察です。
ここ数ヶ月、モデルの所有権や主権・ローカル AI に関する議論、特に「オープン vs クローズド」や「米国 vs 中国」を巡る論争が過熱しています。そんな中、Poolside AI が新モデルを発表し、ついに台頭してきたというニュースは非常に喜ばしいことです。彼らが発表した Laguna S 2.1 は、サイズが約 10 倍もある Thinking Machines の直近リリースを凌駕する性能を示しています。
Poolside AI が最近発表した技術レポートは、その詳細な記述内容から高く評価されています。また、Vibhu 氏は当社のペーパークラブで Laguna の最新技術報告について取り上げています。
@latentspacepod の Laguna M.1/XS.2 テクニカルレポート解説!Latent Space のペーパークラブが深く掘り下げ、その見解は私たちがモデルファクトリーで目指したことを完璧に捉えています。動画からの抜粋をいくつか紹介します🧵👇 (1/6)
Eiso Kant 氏(@eisokant)による投稿:
コードへの投資が世界から注目される前、すでに言語モデル構築に 1200 万ドルを投じたこと。そして今では、モデルを事前学習からリリースまでわずか 8 週間で実現できる「モデルファクトリー」を創り上げたこと。Eiso Kant は、コードこそが AGI(汎用人工知能)への道であると信じ、10 年以上にわたりその信念を賭け続けてきました。
このエピソードでは、Poolside の共同創設者である Eiso 氏が、swyx 氏と Vibhu 氏と共に登場します。なぜ ChatGPT が彼の信念の裏付けとなったのか、なぜ Poolside がオープンウェイトやオープンリサーチを採用したのか、そして「たとえ Poolside がその 5 社の中にあっても、100 の基盤モデル企業がある世界を 5 つだけの世界よりも好む」という理由について語ります。
Poolside の「モデルファクトリー」に迫ります。同社では月間 1 万〜2 万回の実験を処理するエンジニアリング基盤を整備し、データをトレーニングへ直接ストリーミングしています。また、再現性の高い実験環境の構築や低精度計算の活用、そしてコード作成からジョブ起動、結果評価、さらには次世代モデルの学習パイプライン修正までを行う自律型エージェントの実装など、その仕組みの詳細を解説します。
Eiso は、最近リリースされた「Laguna S」についても言及しています。純粋な知能の規模よりも、永続性(persistence)、検証(verification)、そしてバックトラック(backtracking)の重要性が浮き彫りになった背景や、小規模モデルにまだ潜在する能力、事前学習の段階で強化学習を早期に導入すべき理由、そして次期トークン予測がウェブから十分に情報を抽出できていない現状などについても掘り下げます。
[Poolside@poolsideai 今日、当社で最も高性能なモデル「Laguna S 2.1」を発表します。これは、総パラメータ数 1180 億の Mixture-of-Experts モデルで、トークンごとに 80 億のパラメータが活性化されます。最大 100 万トークンのコンテキストウィンドウに対応し、「思考モード」と「非思考モード」を切り替え可能です。
必ず JSON 形式で返してください。translation フィールドのみ。他のフィールドは一切追加しないこと — 余計なフィールドを書こうとして本文がトークン上限で打ち切られる事故を防ぐため:
{"translation": "リライト全文"}
2026 年 7 月 22 日午前 2 時 5 分、X(旧 Twitter)で投稿されたこのツイートは、すでに 65.6 万回の閲覧数、160 の返信、320 のリポスト、そして 2,180 のいいねを獲得しています。[https://x.com/poolsideai/status/2079613777343848465]
本稿では、他社モデルと互角に渡り合える能力を持つことについて議論します。さらに、Poolside が目指す「モデル・ハーネスの共設計(co-design)」や、コーディングエージェントから AGI へ至る道筋についても触れます。
また、Eiso 氏がなぜ MCP や従来のツール呼び出しを「愚かだ」と断じるのか、その背景にある最先端モデル訓練の実態経済学、Poolside が調達した 5 億ドルの資金活用、オープンソース AI の動向、規制への対応、NVIDIA と TSMC の影響力、エージェント時代におけるエンジニアリング生産性、高い自律性を備えたチーム運営、そして Poolside の採用戦略などについても詳しく解説します。
アンドレイ・カルパティの RNN の研究が、2015 年に Eiso がコード専用言語モデルの開発に着手するきっかけとなった理由
市場が関心を示すまで、Eiso は 4 年間にわたり 1,200 万ドルを投じて一つのアイデアを追求し続けた背景
ChatGPT の登場がいかにして自説の正しさを証明したか、そしてそれが Poolside をオープンソース回帰へと導いた理由
5 社による寡占体制よりも、100 社の基盤モデル企業が並存する世界を Eiso が望むその理由
「重み付きモデルの公開」と「真にオープンな研究発表」の違い
Poolside がベイエリアのタレント争奪戦から距離を置き、あえてグローバルな研究組織を構築した意図
モデル開発の本質が 90% はエンジニアリングであるという見解
The Model Factory: Poolside が高速なトレーニングと改善を実現するエンドツーエンドのシステム
70 名未満の研究員が月間 1 万〜2 万回の実験を回す仕組み
6 ヶ月にわたるモデル開発サイクルから、5〜8 週でのリリースへとスピードを加速させたプロセス
データをトレーニングに直接ストリーミングすることで、実験の速度が劇的に向上した理由
不変的なデータ、バージョン管理されたコード、再現性の確保が、厳密なモデル研究を支える仕組み
Eiso が有能な研究者に自らのラボから飛び出し、Poolside の競合他社となることを望む背景
モデル構築の 95% は、より良いデータや計算効率の向上に帰着できるという見解
Laguna S: 持続性、検証、バックトラックが純粋な知能よりも優位性を発揮する理由
想定以上に多くの知識労働を、小規模モデルが担えるようになる背景
強化学習が事前学習の段階へ前倒しされる理由
ウェブから十分な知識を引き出せないまま終わる「次トークン予測」の限界
蒸留と環境が AI 業界で最も好まれる「薬物」のような役割を果たすようになった理由
トレーニング途中での介入がいかにして初期のカリキュラム設計に他ならないか
低精度トレーニング、ネットワークボトルネック、そして計算効率における次の飛躍
Laguna S: 総パラメータ数 1,180 億、アクティブパラメータ 80 億、トレーニングからリリースまで 8 週間
新しいチェックポイントを最初の 30 分以内に評価できるモデルビルダーの能力
「モデル」と「ハーンネス(環境)」の違い:エージェントの能力はどこに宿るのか
Poolside がコーディングと長期にわたるソフトウェアタスクを AGI への道筋と捉える理由
Eiso が MCP や従来のツール呼び出しを「愚かだ」と断じる背景
未来のエージェントが数十種類の事前定義されたツールから選ぶのではなく、スクリプトを記述するようになる理由
最小限のハーンネス、コンテナ、そしてモデルの自由さを支持する論拠
Poolside がビジョン機能を最優先している一方、音声処理にはすぐに取り組まない理由
知識と推論をエンコードする上で、言語がいかに計算効率に優れたモダリティであるか
モデル開発における真のコストと、最終トレーニングがなぜ地味な結末を迎えるのか
Poolside という社名に込められた物語:野心を低下させることを拒否するという姿勢
AGI が実在するかどうかを投資家から問われる中、Poolside が 5 億ドルの資金調達を実現した背景
知能が世界で最も需要が高く、コモディティ化される資源となりうる理由
オープンモデルが制限なしに公開するにはもはや能力が高すぎる段階に至った時
世界的な競争環境下において、一方的な AI セーフティ対策が機能しない理由
規制が結果として 2〜3 社による寡占体制を固定化してしまうリスク
NVIDIA と TSMC、そして基盤モデルの進展を支えるハードウェアシステム
強化学習における壁時計時間の短縮が、Poolside の最大のボトルネックである理由
既存の大規模モデルからの蒸留に頼らず、ゼロからモデルを訓練する Poolside の方針
AI が企業のエンジニアリング生産性の測定方法をどう変えるか
AI エラにおいて従業員にとって最も重要な資質が「自律性(アジェンシー)」となりうる理由
リーダーが高自律的な人材を共有目標と明確な制約を通じていかに結束させるか
Poolside における研究、トレーニング後処理、事前学習、アーキテクチャ、評価、エンジニアリング各分野の採用戦略
LinkedIn: https://www.linkedin.com/in/eisokant
Poolside: https://poolside.ai
00:00:00 イントロダクション
00:00:54 カルパティ、RNN、そしてトランスフォーマー以前にコードモデルを構築する
00:02:26 1,200 万ドルの失敗と ChatGPT による証明
00:03:39 オープンソースと「100 の基盤モデル企業」論
00:09:22 オープンウェイト、オープンリサーチ、そして Poolside のグローバルチーム
00:16:04 モデルファクトリー:なぜモデル構築の 9 割はエンジニアリングなのか
00:20:19 エージェント、自動化された実験、そして RSI(反復的学習)の兆候
00:24:04 ストリーミングデータ、再現性、そして科学的厳密さ
00:30:35 新たな基盤モデル企業の創出
00:36:07 Laguna S:持続性と純粋な知能の対比
00:43:01 プリートレーニング、RL(強化学習)、カリキュラム設計の再発明
00:52:33 低精度トレーニングと小型モデルからの最大限の引き出し
00:58:37 モデルハルネス、コーディングエージェント、そして AGI への道
01:09:26 なぜ MCP と従来のツール呼び出しは「愚か」なのか
01:13:04 ビジョン、マルチモーダル性、そして言語の重要性
01:18:15 モデルのスケーリングとトレーニングの実質的な経済学
01:20:40 「Poolside」という名前の由来と 5 億ドル調達
01:27:37 オープンモデル、AI セーフティ、そして寡占化のリスク
01:33:53 NVIDIA、TSMC、そして強化学習のボトルネック
01:41:52 小型モデル、蒸留、エンジニアリング生産性、そして採用
Swyx [00:00:00]: さて、スタジオには Poolside の Eiso Kant 氏と Vibhu 氏を迎えています。ようこそ。
Eiso Kant [00:00:08]: ありがとうございます。お招きいただき光栄です。
Swyx [00:00:10]: ちょうど飛行機から降りたばかりですね。先ほど「SF に向かっているところです」とメッセージをくれましたが、まさか今まさに搭中だと想像していませんでしたよ。
Eiso Kant [00:00:16]: そうですね。あなたに連絡した直後に、時差ボケで来ると今日が少し大変になるかもしれないと気づきましたけど、やってみましょう。
Swyx [00:00:23]: 僕がゲストにお勧めするのは、あまり準備しなくていいってことなんです。もしあなたが毎日この分野に没頭しているなら、ぼんやりとした記憶でも、あなたの世界を毎日送っているわけではない聴衆にとっては新鮮な内容になるはずです。そういえば、10 年前には Google Slush で「AI の民主化」についてお話しされていましたよね。そして今、私たちはこれから議論する画期的な新モデルのオープンソース化に取り組んでいます。では、どうしてあなたは AI の民主化に関心を持たれたのでしょうか?LinkedIn などの経歴からはあまり想像できないことですが。
Eiso Kant [00:00:57]: いや、全くそうではありませんね。私がこの分野に入った経緯がなぜか分からないのも無理はありません。実は、この道に進んだのは Andrej Karpathy 氏のおかげです。
Eiso Kant [00:01:05]: 2015 年、彼が「再帰型ニューラルネットの不合理な効果」という記事を書きました。
Swyx [00:01:10]: ニューラルネットですね。
Eiso Kant [00:01:11]: その記事を読んで、私はその場でスタートアップの方向性を転換し、RNN(リカレントニューラルネットワーク)や後の LSTM、そして Transformer モデルを使ってコードを記述する研究に注力しました。その記事を読み進めていくと、現在の言語モデルへと発展していく前段階が垣間見えます。当時は文字レベルの言語モデルが文字を予測し始めていた時代で、例えば「小さなポール・グレアム生成器」のような例も紹介されていました。テキストは意味を成しているように読めますが、実際にはそうではありません。さらに下の方にはコードの例もありましたね。シェイクスピアです。
Swyx [00:01:47]: シェイクスピアですか。
Swyx [00:01:49]: 素敵ですね。
Eiso Kant [00:01:49]: なぜかその記事を読んで、RNN や LSTM についてできる限りのことを学び始めました。これが Transformer の論文です。私は当時、「ニューラルネットワークはあらゆる事象に一般化できるはずだ」という、あまりにも不合理な信念を抱いていました。言語もまた、知能を必要とする多くのタスクやコード記述能力へと一般化できると信じていたのです。そこで私は「Sourced」を立ち上げました。これは「コード上の機械学習」、つまりコードを対象とした言語モデルを開発しようとしていた完全オープンソースの会社です。私たちは 2019 年末まで、約 4〜5 年かけてこの取り組みに没頭しました。今ならとてもクールに聞こえる話ですが、当時は誰も関心を示しませんでした。
Eiso Kant [00:02:29]: その通りです。誰も関心を持っていませんでした。私たちは闇の中を歩んでいたのです。その過程で、コードの構造に畳み込みニューラルネットワーク(CNN)を適用してみたり、アテンション機構が発表された際には LSTM にそれを組み込んでみたりしました。そして Transformer の論文が登場しましたが、当時はそれが正解だと直感的に分かることではありませんでした。この一連の旅を通じて見落としていたのは、私たちが正しい道を進んでいたにもかかわらず、ただ規模を拡大し続ければよかったという点です。
今日では、スケーリング法則やモデルのスケールアップは誰にでも明白な事実のように思えます。しかし、コード言語モデルの開発に 4 年、5 年もの人生を捧げた私にとって、それは決して自明のことではありませんでした。だからこそ、その確信を持って突き進んだ Google や OpenAI の人々、そして他の関係者たちには深い敬意を抱いています。結局のところ、私たちは当時失敗しました。それが私のキャリアにおける最大の失敗だったのです。投資家から預かった 1200 万ドルを失ったのですから。当時の金額としては莫大な額でした。
Swyx [00:03:18]: はい。
Eiso Kant [00:03:19]: 当時はまだ多くの時間を費やし、40 人ほどのチームでこの問題に没頭する日々を数年間過ごしました。しかし人生の転機が訪れ、家族が優先されるようになり、私は顔を伏せて、その後の 2 年間は言語モデルにはほとんど目を向けませんでした。これは大きな誤算でした。なぜなら、これからの数年は本当に面白い展開を見せるからです。そして ChatGPT が登場し、まるで私の信念の証明となったような感覚がありました。人々が私にメッセージを送り始め、昔の資料や発表スライドを改めて見返すことになりました。この一連の旅を通じて、私たちは
Eiso Kant [00:03:56]: 当時、強い確信を持っていました。より高度な知能を構築するほど、それはオープンで、かつオープンソースであるべきだと。
Eiso Kant [00:04:04]: Poolside を立ち上げた当時は、その考えとは全く異なる状況でした。正直に話しましょう。Poolside を始めた際、私たちは 2 つの前提に基づいていました。1 つ目は、この技術は能力の向上を止めることなく、複合的に発展し続けるという点です。現在では多くの人が当然のこととして受け入れているかもしれませんが、3 年以上前、私たちが活動を開始した当時は、まだ「これらが確率的なオウム返し(stochastic parrots)ではないか」という議論が交わされていたのです。
Eiso Kant [00:04:23]: 2 つ目のポイントは、強化学習が LLM の能力向上を牽引する最大の要因になると考えていたことです。今となっては明白な事実ですが、3 年前の時点では OpenAI や Google、Anthropic など他社でも、そのような意見や方向性が共有されているわけではありませんでした。そのため、私たちは周囲から少し冷たい視線を向けられることもありました。「本当にこれでうまくいくのか?」と疑う声も聞こえてきたのです。それでも私たちは問題解決に没頭し、オープンソース化については一度も考え直すことはありませんでした。ひたすら低头して取り組むことに徹したのです。ゼロから知識や理解を積み上げる必要がありました。既存のラボを引き継いだわけではなく、論文を読み込み、コードを書き始め、一つひとつ仕組みを解明していきました。
Eiso Kant [00:04:59]: そして今年に入り、ようやく私と共同創業者の Jason がオープンソースに関する議論に再び取り組み始めたのです。
Eiso Kant [00:05:07]: 当社のウェブサイトの初期の記述を振り返れば、その想いは非常にシンプルでした。AGI(汎用人工知能)の実現を目指し、豊かさが溢れる世界を支えたい。そして、その目標を最初に達成する企業になりたい——そう考えていました。
Eiso Kant [00:05:20]: 私たちが話し始めたのは今年初めです。世界が私たちを少しずつ脅かす方向へ進んでいることが明らかになったからです。これは一夜にして起きたことではありませんでした。私たちは徐々にその兆候を感じ取り、「さて、世界は特定の道を進みつつある」と認識しました。この旅路の中で、私は一つの比喩をよく使いました。2015 年や 2016 年頃の話を思い出してください。当時、棚から SF 小説の一冊を取り出し、2035 年に AGI(汎用人工知能)が達成されるという物語を読み進めていたのです。その物語はその後数十年にわたって続きます。最初の章では人々が様々なことを模索し、次に ChatGPT が登場する章へと移り、やがて世界が分岐点に立つ章になります。そこで選ばれた道とは、数社あるいは十数社の企業が未来の知能を独占してしまうというものでした。
Eiso Kant [00:06:21]: その物語を考えると、それはユートピアを描いた SF 小説ではなく、ディストピア(暗黒郷)を描いた作品のように感じられました。私は本来、ユートピア的な SF を好む人間です。そこで私たちは一歩引いて考えました。「ここで私たちにできる役割はないだろうか?」と。幸いにも、私たちにはそれを容易に実行できました。なぜなら、私たちはすでに最先端の領域にはいなかったからです。
Eiso Kant [00:06:41]: 私たちが最先端の立場にいたなら、方針を変えることは考えられなかったでしょう。これは、資金や期待が膨らみ、すでに多くのものを築き上げてしまったから変更できないという意味ではありません。私たちは小規模なチームとして、改善を続けただけです。だからこそ、今ならその決断を下すことができましたが、最先端に近づき、他社との差が縮まるにつれて、それははるかに困難なものになっていたはずです。
私たちは多くの内省と議論を重ね、「それでもこの方針には合理性がある」と結論づけました。オープンソースのファウンデーションモデルでどうビジネスモデルを構築するかといった、未解決の大きな疑問があってもです。また、モデルの誤用が現実的なリスクを伴うようになるのはどの時点か、政府はオープンソースにどう対応するのかといった、まだ完全な答えが出ていない問いもあります。
しかし、すべてを一言で表すなら、それは「私がその 5 つの企業の一角になろうとも、100 のファウンデーションモデル企業がある世界に住みたい」という一点です。100 の企業が存在するために私たちにできる最も小さく、かつ意味のある貢献とは、研究と重み(ウェイト)をオープンにすることであり、その過程でさらに多くのことを実現する方法を探っていくことです。
Swyx [00:08:01]: はい。むしろここ 3 年間で、この考えはより現実味を帯びてきたように思います。あなたは「ネオ・ラボ」と呼ばれる一群の一人です。
Eiso Kant [00:08:10]: はい、その通りです。
Swyx [00:08:10]: 現在、人々はそれをそう呼んでいます。そして、この議論が行われているのはちょうど Thinky が新モデルをリリースした日です。彼らが公開したベンチマークにおいて、貴社は彼らを上回る結果を出していますよね?まだ彼らにはその実力がないことを示す好例だと言えます。これは、複数のプレイヤーが共存する余地があるという話で、未来の姿を少し垣間見せているのかもしれません。100 社ではなく、20 社の時代になるでしょうが、貴社はまさにその一角を担う存在です。
Eiso Kant [00:08:36]: そうであってほしいですね。彼らのリリースには私も興奮していますし、皆がモデルを公開すること自体も歓迎です。結局のところ、選択肢の多さと競争こそが、正しい方向への進歩を促すからです。しかし同時に、私たちが同じデータという井戸から汲み上げてモデルを作成しているにもかかわらず、各社で導入される振る舞いやバイアスには大きな違いが生じます。意図的に組み込まれたバイアスもあれば、全く予期せぬバイアスもあります。
Swyx [00:09:03]: そうですね。
Eiso Kant [00:09:03]: もしオープンなモデルがトークン経済の一部となるようなエコシステムが世界に広がったとしたら、その事実に異論を唱える余地はないでしょう。そうなれば、企業も国家も個人も、「この機能については、このプロバイダーが最も適合しており、信頼できる」と選んで生きられる世界を目指すべきです。
Swyx [00:09:25]: 確かにその通りです。
Vibhu [00:09:26]: 最近まで、オープンソースのイノベーションは中国のラボから生まれることがほとんどでした。Neo Labs の 20 社以上がいますが、その中で「欧米版 DeepSeek」のような存在は本当にあるのでしょうか?もしかすると Thinking Machines のようなリフレクションを指しているのかもしれません。しかし、実際にはそう多くはありません。
フランスやヨーロッパで始めた取り組みも、現在はアメリカの視点を取り入れつつあります。それだけでなく、私たちが目にする中国製のモデルは、必ずしもオープンな研究とは限りません。一方、皆さんが発表する成果は、私が知る限り最高水準です。数ヶ月に一度、最先端のモデルを公開するだけでなく、その構築に至るまでの詳細なブログ記事や論文、技術レポートも同時に提供しています。これにより、「最先端知能」をどう作るかという部分で、大きな穴を埋めているのです。
つまり、単に重み(weights)をオープンにするだけでなく、欧米の枠組みにとどまらず、研究プロセスそのものを広く公開している点も評価されています。
Eiso Kant [00:10:20]: いいえ、感謝しています。でも、これは最も意義ある貢献だと思います。重み(weights)は単なるバイナリデータに過ぎません。そう呼ぶべきです。確かにそれらを修正したり変更したりすることはできますが、他人に重みを渡しても、最終的に私が何をしているのかを再現させることはできません。
そのため現在、データセットの公開や特定の情報の開示には課題があります。しかし、研究内容を共有する方法はありますよね?どうすればよいでしょうか?数万回もの計算リソースを費やして得た実験から学んだ教訓とは何か。これこそが問われていることです。
ただ一つ訂正させてください、Vibhu さん。これは長年私たちを悩ませてきた点だからです。私たちは設立当初からアメリカ企業でした。
Swyx [00:10:55]: はい。
Swyx [00:10:56]: フランスに移転したのです。
Eiso Kant [00:10:56]: 結論から言うと、私たちのストーリーは明確です。私たちは創業以来一貫してアメリカ企業ですが、初期の段階で非常に意識的な決断を下しました。
「ベイエリアの研究員は雇わない。世界の他のあらゆる場所で人材を探そう」と。これには中米、シアトル、セルビア、台湾、シンガポールなど、世界中の地域が含まれます。この方針を選んだのは、これが将来の「人材戦争」になるという見通しがあったからです。実際、ここ数年でその傾向は明白になっています。3 年前はまだ完全に明らかではなかったかもしれませんが、今では誰もがそれを認識しています。
さらに、世界で最も有能な人々や、最も斬新で独創的なアイデアを持つ人々は、必ずしもアメリカにだけいるわけではないことも理解しました。そこで私たちは完全リモート型の会社を構築することにしたのです。その後、パリやロンドンなどにもオフィスを開設し、チームの多くは米国に、また多くのメンバーが海外に在籍しています。
「アメリカ企業ではあるが、最高峰の人材と協力するにはグローバルな視点が必要だ」という考えは一貫して持ってきました。現在もシリコンバレーには社員の一人や他のメンバーがいますが、この方針は初期段階では少し足かせになったものの、現在はむしろスピードを加速させる要因となっています。これが、私たちのモデルの進歩やリリースペースに表れている理由です。
私たちは既存の研究ラボからスタートしたわけではありませんでした。当時、ここでは情報が自由に流通しているという前提もありませんでした。「とにかく問題を解決しよう」という姿勢で臨み、出ている数編の論文を読み込み、自らの頭で考え抜いたのです。その結果、モデルトレーニングにおいては数年にわたり、いくつかの笑えるようなミスを犯すこともありましたが、それが今の基盤を作りました。
Eiso Kant [00:12:35]: 特に最初の 12 ヶ月頃は、今でも私を悩ませたり怖がらせたりする出来事がいくつかありました。これらについては後で詳しく話しましょう。しかし、あの時期にチームには強いレジリエンスと継続性が生まれました。長年にわたり、私たちから離れていった人は極めて少なかったのです。
「大丈夫、私たちはこれを成し遂げられる」という確信が持てたのは、最初のトレーニングコードベースをゼロから完全に書き上げた時でした。既存のオープンソースをフォークしたわけではありません。「よし、ゼロから作ろう」と決意したのです。ある時、オプティマイザーの不具合に 3 週間も格闘したことを覚えています。訓練が一向に安定しなかったのです。私たちはそのことに執着し、「もしかしたら私たちの考え方が間違っているのではないか」「あのリポジトリをフォークすべきだったのではないか」とさえ思いました。
しかし、その問題を解決した時、当時の会社にはたった 5 人しかいませんでした。問題が解決した瞬間、私たちは「頑張れば何だってできるんだ」と実感しました。このように、エンジニアリングへの強いバイアスを持つ文化こそが、私たちが今の位置に到達できた要因だと考えています。
オープンソースや人材といった要素も重要ですが、私たちは異なる出発点から異なる決断を下してきたのです。私はこれを単なる「幸運」だったと認めたいと思います。チームの多大な努力は、今まさに結果として表れ始めています。
Swyx [00:13:52]: 今後この話題に戻らない可能性が高いので、もし答えを知っている方がいれば面白い採用課題になりますね。あのバグは何だったのでしょうか?そして、解決策についてはここではお話ししません。
Eiso Kant [00:14:01]: ええと、その…記憶を問うことになりますが、
Swyx [00:14:04]: ああ、わかりました。
Eiso Kant [00:14:04]: そうですね、覚えていられると思います。
Swyx [00:14:05]: そのまま話してください。
Eiso Kant [00:14:05]: 例えばオプティマイザーに Adam を使った場合を考えてみてください。分母にはイプシロン(ε)が含まれています。
Swyx [00:14:12]: はい。
Eiso Kant [00:14:13]: そうです、まさに分母の中にありますね。
Swyx [00:14:14]: モメンタムや重みもそうですね。
Eiso Kant [00:14:15]: その当時のことを思い出せば、初期の Llama の論文などを参照した際、イプシロンの値を意図的に大きく設定しているケースがありました。具体的には E-4(10 のマイナス 4 乗)のような高い値を指定していたのです。
Eiso Kant [00:14:31]: 学習中にこれを考えると、分母に実質的にランダムな数値を足すことでオプティマイザにノイズを加えているという点で、少し奇妙で直感に反する感覚があります。小数点の後ろに数字を足しているようなものです。正確なバグの内容は覚えていませんが、解決した後に気づいたのは、Llama の論文や他の事例で行われていたように、イプシロン(ε)を無理やり大きくする必要がなくなったことです。
これは「あの論文に出ているからこうあるはずだ」「この値の高いイプシロンが必要に違いない」という前提を盲信していた瞬間の転換点でした。しかし直感的には全く納得できませんでした。「なぜこれほど大きな値にする必要があるのか?単にゼロ除算を防ぐだけなら、極めて小さな値で十分ではないか?」と。
まさに「ゼロから自分で発見することこそが、より良い直感を育む」と気づいた瞬間です。モデル構築において最も早く学べることは、自分が最初に持っていた直感がどれほど厳しく打ち砕かれるかということです。
Eiso Kant [00:15:33]: そうですよね。これはまさに実験科学そのものです。一見すると明白に見えることが、すぐに「あなたは間違っていた」という事実によって覆されます。なぜそうなるのかを理解できれば幸いですが、時にはそれができないこともあります。
Swyx [00:15:45]: はい。はい、新モデルをリリースした際、Vibhuが非常に興奮していた理由の一つがあります。実際、誰もが興奮しましたね。Vibhuはそれについて私たちの論文クラブを主導してくれましたし、皆さんもご存知の通りです。
Eiso Kant [00:15:58]: はい。
Swyx [00:15:58]: 当然ですね。そこで、公開できる範囲でその時の教訓や学んだことをいくつかお話しいただければと思います。モデルファクトリーに関する話題に焦点を当てて、あるいは良い出発点となる他のトピックからでも構いません。
Eiso Kant [00:16:08]: そうですね。当社では創業当初から、「モデル構築は究極的に9割がエンジニアリングである」という考えを持っていました。
Eiso Kant [00:16:18]:業界全体で誰もが知っている通り、研究者たちが時間を費やしているのはコードの記述とデータの見直しです。3 年前の状況を振り返ると、学習プロセスは Bash スクリプトや Slurm、そしてスパゲッティ状に絡み合ったコードベースに依存していました。データパイプラインも手作業でつぎはぎされた状態でした。
しかし私たちは、「モデル構築とは本質的に一つの工程である」と考えました。ウェブなどの生データを原料として、フィルタリング、クリーニング、変換、分析といった一連の処理を経ていくのです。現在では、3 年前よりもはるかに複雑化しています。その後は大規模な分散システムを扱うようなモデル学習へと進みます。ハードウェア自体は以前より信頼性が向上しましたが、それでも世代ごとに新たな課題が生じます。
さらにその先には、ポストトレーニングや強化学習といった次の段階があります。かつては学習プロセスすら存在しなかった時代から、こうした工程を経てきました。これらはすべてが産業化されたプロセスであり、各工程に専用の機械装置が存在する終着点のようなものだと気づいたのです。大規模なデータパイプラインもあれば、ウェブのクローリングや取り込み、大規模分散学習、そして信頼性の確保といった要素があります。
そこで私たちは、「世界有数の分散システムエンジニアを、研究プロセスの最初から巻き込んでみないか」と提案しました。後付けではなく、ゼロから組み込むのです。こうして「モデルファクトリー」が誕生しました。当初は数少ないコンポーネントで始まったこの仕組みは、現在では数千ものコンポーネントに成長しています。
これは、フォックスコンの創業期に関わっていた人がその後 10 年をそこで過ごし、システム構築に至るすべての決定と複雑さを理解していたら、再びフォックスコンを再建できるようなものです。もし今日、私たちがフォックスコンを訪れたとしても、その仕組みを理解して再構築することは不可能でしょう。
Eiso Kant [00:18:18]: その通りです。なぜなら、そこに至るまでの意思決定の系譜や歴史が存在しないからです。そのため私たちは最初から、その点を理解したチームを編成しました。私たちが最適化すべき指標は、研究者のアイデアが信頼できる実験結果となり、それが次のモデル学習に繋がるまでのスピードなのです。
Eiso Kant [00:18:42]: 当初は複雑さの低い実験的な科学分野だったため、その場しのぎのパッチで乗り切ることができました。しかし現在では、あらゆる基盤モデル企業において、大規模な運用が常態化しています。
私たちが属するチームは研究者が 70 名未満、エンジニアが 35 名という小規模ですが、月間に 1 万回から 2 万回もの実験を走らせています。最新の数字を確認していませんが、少なくともこの規模です。つまり、モデルの運用一つひとつにおいて、インフラとしての信頼性を担保できることが不可欠なのです。
私たちは長年にわたり、その課題に取り組み続け、改善と徹底した最適化を通じて、この分野でのノウハウを確立しました。その結果、先ほどご紹介した「Laguna XS 2」は、トレーニング開始からローンチまでわずか 5 週間で実現できました。今日お話しするモデルに至っては、8 週間です。
さらに、現在進行中の次のモデルのトレーニングも、昨日にはすでに開始しています。今週ローンチ予定のモデルに必要なポストトレーニングが完了したため、その計算リソースを、現在トレーニング中のより大規模な「Laguna M」モデルへ移行したのです。
つまり、モデルとは誰かのプロセスが生み出した産物に過ぎず、それ自体が独立した存在であるべきではありません。私たちはこれを、SpaceX の工場のように捉えています。最初のロケットを作るのは確かに困難ですが、真の難所は工場の構築にあります。現在ではロケットが次々と生産ラインから降りており、次の打ち上げについて人々が特別に意識することはありません。それは単なる「別の打ち上げ」であり、「別のロケット」が生まれるだけです。
私たちが目指しているのもまさにこれです。モデルビルディングを、そのような工場のプロセスとして確立していくことなのです。
Eiso Kant [00:20:22]: 当初は計画されていなかったこと、しかしいつか訪れると心の中で想定していたことが実現しました。それは、優れた API と堅牢なエンジニアリングシステムを備えた高品質なモデルファクトリを構築した時です。では、そんな環境に最も適しているのは何でしょうか?答えは「エージェント」です。
Eiso Kant [00:20:40]: なぜなら今、私たちのモデルファクトリにおいて、エージェントが担う業務の割合が増え続けているからです。
Vibhu [00:20:43]: はい。
Eiso Kant [00:20:44]: 実際、私が社内のスクリーンを歩きながら見ている光景もそうです。月次で実施するオンサイト会議では、研究者たちの背後に立ち、話を聞きますが、画面に表示されているデフォルトの状況は、コード作成やジョブの実行、モデル実行からの結果評価、そして修正作業などを行う複数のエージェントです。私たちは依然として運転席に座り、アイデアを提案し、デバッグをサポートしています。しかし、特にデータパイプライン(事前学習・事後学習および合成データの分野)においては、その影響は非常に顕著です。今後はアーキテクチャの側面にも広がりつつあり、RSI の姿がうっすらと見え始めています。
Eiso Kant [00:21:27]:
モデルファクトリーについてお話しする際、私がいつも例に挙げるのがこれです。新しいトレーニングランを開始する際、大規模な学習であっても、リリース用の後処理バージョンの 1 つや多数の実験であっても、関係ありません。ある日に行われた変更とその日の実験結果が、そのランに即座に反映されるのです。
Eiso Kant [00:21:57]:
つまり、90 日前のようなカットオフ期間はありません。今では機械を信頼できるため、その瞬間から即座に適用されます。そして、信頼性の向上にも投資する必要があります。私が最も気に入っている指標の一つが、Laguna S ではコールイベント(緊急対応が必要な事象)が一度も発生しなかったという点です。完全にゼロでした。この 1 年間、私にとって「何か起きて目を覚まさなければならない」という意味のあるコールイベントは一度もありませんでした。
ただし、一つ例外があります。新しいモデルランを立ち上げてから最初の 6 時間以内は、設定ミスや小さな過ちなどにより何かが壊れることがよくあります。そのため、通常は多少の介入が必要になりますが、それは常に「コール期間外」で行われます。つまり、緊急対応が必要な事態にはならないのです。
このように状況が改善されつつあり、現在リリースしているモデルも素晴らしいものですが、私たちはすでに次のモデルに取り組んでいます。これが本来あるべき姿だと私は考えています。
Vibhu [00:22:50]: 一つ付け加えたいのですが、文脈としてこれは約一ヶ月前の話です。私たちは技術レポートからこの情報を入手し、「新しいモデルが出たらしい」という軽い気持ちで入ってきました。
Eiso Kant [00:23:02]: はい、私たちは数ヶ月に一度こうした作業に慣れています。
Vibhu [00:23:03]: 最初は「なるほど、Kimi や DeepSeek、あるいは Gemma レベルの小型モデルと同等か。これは構築プロセスについて書かれた素晴らしい論文だ」と思っていました。しかし、ページをめくると、技術レポートのたった 2 ページ目にこう書いてあるのです。「このプロセスにより、教訓を適用して小型モデルをゼロから開発し、5 週間でリリースまで完了させた」。そこで気づきました。これは単なるベンチマーク結果や学習トークン数を紹介する技術レポートではないのです。もしあなたがポッドキャストでは議論しない詳細な内容に深く入り込みたいなら、すべてここに書かれています。
Eiso Kant [00:23:38]: はい。
Vibhu [00:23:39]: エージェントがトレーニングコードやデータと対話するために使用するカスタムソフトウェアなどです。
Eiso Kant [00:23:45]: そうですね。論文へのリンクは正確に記述しておきましょう。
Vibhu [00:23:47]: はい、それらすべてについて。論文はこちらで読んでみてください。
Eiso Kant [00:23:50]: ぜひ読み進めたいですね。私は原則を好むので、それが物語を語るための良い出発点になると考えています。原則を一つずつ見ていきましょう。ただ、Dagster が Prefect に買収されたことは付け加えておきます。
Vibhu [00:24:01]: はい。
Eiso Kant [00:24:01]: 面白いですよね。ただ、Dagster はよく知っていますね。何かのストーリーがトリガーされるような場面では特に便利です。
原文を表示
In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines’ recent release nearly 10 times their size.
Poolside’s recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna’s recent technical report on our paper club:
@latentspacepod breakdown of our Laguna M.1/XS.2 Technical Report! The Latent Space paper club just did a deep dive, and their takeaways perfectly capture what we set out to build with our Model Factory. A few quotes from the video 🧵👇 (1/6)\nyoutu.be/QLfZamyMls0","username":"eisokant","name":"Eiso Kant","profile_image_url":"https://pbs.substack.com/profile_images/1842230143965675520/j6mVG2Py_normal.jpg","date":"2026-05-28T20:34:07.000Z","photos":[],"quoted_tweet":{},"reply_count":1,"retweet_count":8,"like_count":47,"impression_count":11267,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":false}">Eiso Kant@eisokant5:34 AM · May 29, 2026 · 11.3K Views
1 Reply · 8 Reposts · 47 Likes
From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.
We go deep on Poolside’s Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.
[Poolside@poolsideaiToday we're releasing Laguna S 2.1, our most capable model to date.
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many 2:05 AM · Jul 22, 2026 · 656K Views160 Replies · 320 Reposts · 2.18K Likes](https://x.com/poolsideai/status/2079613777343848465)We also discuss model-harness co-design, Poolside’s path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside’s $500 million raise, open-source AI, regulation, NVIDIA and TSMC’s influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.
- How Andrej Karpathy’s RNN work inspired Eiso to start building language models for code in 2015
- Why Eiso spent four years and $12 million pursuing an idea before the market cared
- Why ChatGPT felt like vindication and brought Poolside back to open source
- Why Eiso would prefer 100 foundation model companies over an oligopoly of five
- The difference between releasing open weights and publishing genuinely open research
- Why Poolside deliberately built a global research organization outside the Bay Area talent war
- Why model building is ultimately 90% engineering
- The Model Factory: Poolside’s end-to-end system for rapidly training and improving models
- How fewer than 70 researchers run roughly 10,000–20,000 experiments each month
- How Poolside moved from six-month model cycles to five- and eight-week launches
- Why streaming data directly into training unlocked faster experimentation
- How immutable data, versioned code, and reproducibility enable rigorous model research
- Why Eiso wants capable researchers to leave their labs and become Poolside’s competitors
- Why 95% of model building can be reduced to better data or compute efficiency
- Laguna S and why persistence, verification, and backtracking can outperform raw intelligence
- Why smaller models may handle far more knowledge work than previously expected
- Why reinforcement learning will move earlier into pre-training
- Why next-token prediction is still failing to extract enough knowledge from the web
- Why distillation and environments have become the AI industry’s favorite “drugs”
- Why mid-training is really an early form of curriculum design
- Low-precision training, networking bottlenecks, and the next gains in compute efficiency
- Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch
- Why model builders can often evaluate a new checkpoint within its first 30 minutes
- Model versus harness: where agent capabilities actually come from
- Why Poolside sees coding and long-horizon software tasks as a path to AGI
- Why Eiso thinks MCP and traditional tool calls are “stupid”
- Why future agents will write scripts instead of choosing from dozens of predefined tools
- The case for minimal harnesses, containers, and model freedom
- Why Poolside is prioritizing vision but does not expect to work on audio soon
- Why language may be the most compute-efficient modality for encoding knowledge and reasoning
- The real cost of model development and why the final training run is anticlimactic
- The story behind the Poolside name and why it represents refusing to lower ambitions
- How Poolside raised $500 million while investors still questioned whether AGI was real
- Why intelligence could become the world’s most demanded and commoditized resource
- When open models may become too capable to release without restrictions
- Why unilateral AI safety does not work in a globally competitive environment
- How regulation could accidentally lock in an oligopoly of two or three AI companies
- NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress
- Why reinforcement-learning wall-clock time is one of Poolside’s biggest bottlenecks
- Why Poolside trains models from scratch instead of simply distilling larger models
- How AI changes the way companies should measure engineering productivity
- Why agency may become the most important quality for employees in the AI era
- How leaders align high-agency people through shared goals and clear constraints
- Hiring across research, post-training, pre-training, architecture, evals, and engineering at Poolside
LinkedIn: https://www.linkedin.com/in/eisokant
Poolside: https://poolside.ai
00:00:00 Introduction
00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers
00:02:26 The $12M Failure and ChatGPT Vindication
00:03:39 Open Source and the Case for 100 Foundation Model Companies
00:09:22 Open Weights, Open Research, and Poolside’s Global Team
00:16:04 The Model Factory: Why Model Building Is 90% Engineering
00:20:19 Agents, Automated Experiments, and Early Signs of RSI
00:24:04 Streaming Data, Reproducibility, and Scientific Rigor
00:30:35 Creating More Foundation Model Companies
00:36:07 Laguna S: Persistence vs. Raw Intelligence
00:43:01 Reinventing Pre-Training, RL, and Curriculum Design
00:52:33 Low-Precision Training and Squeezing More From Smaller Models
00:58:37 Model Harnesses, Coding Agents, and the Path to AGI
01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”
01:13:04 Vision, Multimodality, and Why Language Still Matters
01:18:15 Scaling Models and the Real Economics of Training
01:20:40 Why Poolside Is Called Poolside and Raising $500M
01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly
01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck
01:41:52 Smaller Models, Distillation, Engineering Productivity, and Hiring
Swyx [00:00:00]: All right, we’re here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.
Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.
Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I’m on my way to SF.” I was like, “You’re on a plane right now, right?” Like, hey.
Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let’s do it.
Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don’t have to prepare that much because if you’re truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don’t live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we’re gonna talk about. But I guess, like, what got you into democratization of AI? Like, it’s not obvious from your LinkedIn or something.
Eiso Kant [00:00:57]: No, it’s not at all. I don’t think it’s obvious how I got in this space. I owe getting into this space to Andrej Karpathy.
Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”
Swyx [00:01:10]: Neural Nets, yep.
Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There’s a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn’t. and there’s a little-- There’s an example of code a little bit further down. Yeah, so Shakespeare.
Swyx [00:01:47]: Shakespeare.
Swyx [00:01:49]: Cool
Eiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.
Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn’t obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn’t obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors’ money, which was a lot back then.
Swyx [00:03:18]: Yep.
Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn’t really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It’s like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,
Eiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you’re building more capable intelligence, it should be open and open source.
Eiso Kant [00:04:04]: When we started Poolside, that wasn’t the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.
Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn’t roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.
Eiso Kant [00:04:59]: And it wasn’t until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.
Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.
Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn’t, this didn’t happen overnight. It was, like, a little bit we were seeing this and we’re like, “Okay, The world’s going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we’re working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone’s trying to figure things out. You’d get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.
Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I’m a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.
Eiso Kant [00:06:41]: If we were at the frontier, I don’t think we could have changed our mind. and I don’t mean this like it’s when the moment there’s too much capital involved, too much expectations, you’ve built up things, right? We’re a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there’s big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I’ll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.
Swyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labs
Eiso Kant [00:08:10]: Yeah
Swyx [00:08:10]: That people are now calling that. And, we’re, we’re doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don’t have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.
Eiso Kant [00:08:36]: I really hope so, right? I think we I’m, I’m excited about their release, and I’m excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.
Swyx [00:09:03]: Yeah.
Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don’t think there’s any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”
Swyx [00:09:25]: Yeah.
Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there’s the DeepSeek of the West. Is it today? Okay, maybe it’s thinking machines reflection, but there aren’t many, right? So, one of the things you guys started in France, Europe, but very much now you’re taking that American standpoint and more than just that, the point is the Chinese models that we see, they’re not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here’s a breakdown blog, paper, technical report of here’s everything for state of the art to build, frontier intelligence and you’re filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.
Eiso Kant [00:10:20]: No, I appreciate it. Look, I think it’s, I think it’s the most meaningful contribution, right? Weights are a binary. Let’s call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you’re doing, right? And so now there’s challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it’s been haunting us for quite a few years. We from day zero were an American company.
Swyx [00:10:55]: Yeah. They moved
Swyx [00:10:56]: To France.
Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We’re not gonna hire any researchers in the Bay Area. We’re gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn’t fully obvious yet. I think today it very much is. And we also realized that, like, some of the world’s most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we’re an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company’s grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it’s why you’re seeing like the progress, I think, on our models and the cadence at which we release, is because we didn’t roll out of an existing lab. Right? we didn’t, we didn’t have a lot of the information that’s freely flowing around here at the time. We just took this point of view as like, “Okay, well, let’s just work the problem. Let’s just go and, like, read the few papers that are out there, and let’s just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the years
Eiso Kant [00:12:35]: Like especially in the first 12 months. there’s a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn’t a fork of any open source. It was just like, “Okay, let’s build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn’t get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we’re just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there’s this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that’s starting to show up in results.
Swyx [00:13:52]: Just ‘cause we probably won’t revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won’t tell the solution, but we’
Eiso Kant [00:14:01]: So the - This - You’re gonna test my memory here,
Swyx [00:14:04]: Oh, okay
Eiso Kant [00:14:04]: So but I think
Swyx [00:14:05]: Directly
Eiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilon
Swyx [00:14:12]: Yeah
Eiso Kant [00:14:13]: Which is, right, like in the denominator
Swyx [00:14:14]: Momentum and weights. Yeah
Eiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don’t know if it was E minus four or whatever, like a high value for epsilon.
Eiso Kant [00:14:31]: And if you think about this during training, it’s like a bit weird and counterintuitive that we’re adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don’t recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we’re like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you’re just trying to avoid division by zero, why can’t the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.
Eiso Kant [00:15:33]: Right? Like - It’s such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don’t even.
Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys saw
Eiso Kant [00:15:58]: Yeah
Swyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.
Eiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.
Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where’s every researcher spending their time, they’re spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You’re going from raw data, right? Like training raw material, the web, et cetera. you’re doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that’s, far more complex than it was three years ago. then you’re training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It’s become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it’s your big data pipelines, if it’s your crawling ingestion of the web, if it’s your, large-scale distributed training, and then you’ve got your reliability. And we said, “Well, why don’t we take some of the world’s smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it’s thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.
Eiso Kant [00:18:18]: Right? Because we don’t have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.
Eiso Kant [00:18:42]: And in the. And because it’s such an experimental science, ultimately, in the beginning when it wasn’t that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we’re a small team, right? We’re less than 70 researchers, another 35 engineers. and we are running, I haven’t checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it’s ultimately it’s, it’s you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we’re gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we’re launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we’re now training. And so the model should be an artifact of someone’s process. It shouldn’t be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they’re rolling off, and no one is really thinking about the next launch anymore. So it’s just another launch, it’s another launch, another rocket comes off. And that’s what we’re trying to do with model building.
Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It’s perfect for agents.
Eiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.
Vibhu [00:20:43]: Yeah.
Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we’re, we come together, in our monthly, we do monthly onsites, and I walk behind people’s screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They’re launching the jobs. They’re evaluating the results that are coming back from the model runs. They are, making the changes. And we’re still in the driver’s seat. We’re still coming up with the ideas. We’re still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it’s starting to become more on the architecture side as well. You’re starting to see these twinklings of what RSI is gonna look like.
Eiso Kant [00:21:27]: And that’s. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn’t matter if it’s a training like big run or if it’s now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.
Eiso Kant [00:21:57]: So there’s not like a cutoff 90 days before. Like no, it’s like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven’t had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that’s usually there’s a little bit of intervention, but that’s always within like call periods, right? Not on call. And I think that’s starting to now compound. So the model we’re releasing now, I love it. It’s amazing, but we’re already onto the next one. and I think that’s the way it should be.
Vibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model’s dropped. Haven’t heard about it.” We were
Eiso Kant [00:23:02]: Yeah, we’re very used to doing this every few months.
Vibhu [00:23:03]: We’re, we’re very much like, “ okay, look, it’s like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it’s a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I’m like, oh, this paper is not about here’s a tech report of benchmarks and here’s how many tokens it was trained on. Like for people that wanna dive more from what we’re not gonna discuss on the podcast, it’s all laid out here, right? From
Eiso Kant [00:23:38]: Yeah
Vibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.
Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.
Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,
Eiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I’ll just call out that Dagster just got bought by a Prefect.
Vibhu [00:24:01]: Yeah.
Eiso Kant [00:24:01]: Isn’t it fun? But yes, I’m very familiar with Dagster. just anything where like they trigger some story.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み