Poolside、コーディング特化モデル「Laguna S 2.1」公開
Poolside は、118B パラメータのオープンウェイトモデル「Laguna S 2.1」を公開し、SWE-Bench Multilingual でトップクラスの結果を出しながら単一の NVIDIA DGX Spark で動作可能な高効率なエージェントコーディングモデルとして業界に衝撃を与えた。
キーポイント
圧倒的なコストパフォーマンスと性能の両立
118B パラメータのうち活性化するのは約 6.8%(8B)のみという MoE アーキテクチャにより、大規模モデルに匹敵する性能を維持しつつ、単一の DGX Spark で動作可能な低コストを実現している。
SWE-Bench Multilingual でのトップスコア
同モデルは SWE-Bench Multilingual で 78.5% のスコアを記録し、DeepSeek-V4-Pro-Max や NVIDIA Nemotron 3 Ultra などの大規模クローズドモデルを含む競合群の中でトップの成績を収めた。
1M トークンのコンテキストと思考モード
思考(thinking)モードと非思考モードをサポートし、最大 100 万トークンのコンテキストウィンドウを処理可能で、複雑な長期的コーディングタスクにも対応できる。
迅速な開発と多様なフォーマット公開
2026 年 5 月のトレーニング開始からわずか 9 週間でリリースされ、BF16, FP8, INT4, NVFP4 など複数の精度形式や GGUF/MLX 変換が Hugging Face で公開された。
Max モードによる性能向上とコスト
デフォルトの「max thinking」モードにより Terminal-Bench 2.1 で 70.2%、DeepSWE で 40.4% と大幅なスコア向上を達成したが、その分トークン使用量は約 2.5 倍に増加する。
単一ハードウェアでの動作と最適化
118B パラメータモデルが 4-bit 量化で NVIDIA DGX Spark 1 台、FP8 でも同機または H200 1 台で動作可能となり、NVIDIA との連携による TRT-LLM 最適化が実現されている。
実証された自律的なコーディング能力
人間介入なしで HTML/CSS ブラウザエンジンの構築や自己エージェントハッチの最適化、および知識カットオフ以前の数学問題の独立した再発見など、長時間にわたる複雑なタスクを完遂する。
重要な引用
Laguna S 2.1 holds its own against models several times its size
That sparsity is why a mid-size model can behave like a larger one while staying cheap to serve.
On SWE-Bench Multilingual it scores 78.5%, topping the published table outright.
Poolside team reports coherent, productive reasoning running for hours and hundreds of thousands of tokens.
That result is an independent rediscovery; GPT-5.2 Pro solved the same problem in January 2026, and Laguna's knowledge cutoff is November 2025.
an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual
影響分析・編集コメントを表示
影響分析
このリリースは、オープンソースコミュニティにおける「高性能かつ低コスト」の基準を劇的に引き上げるものであり、企業や個人開発者が大規模モデルに匹敵するコーディング支援ツールをローカル環境で利用可能にする可能性を開いた。特に、SWE-Bench Multilingual でのトップスコアは、多言語対応コード生成におけるオープンソースモデルの成熟度を証明し、クローズドモデルへの依存を減らす重要な転換点となる。
編集コメント
Poolside の Laguna S 2.1 は、単なるパラメータ数の競争から「いかに少ない計算資源で最大の性能を引き出すか」という次のフェーズへの明確なシフトを示しています。特に、FP8 精度での強化学習や NVFP4 対応といった技術的細部が実用化されている点は、今後のオープンソースモデル開発の標準となるべき重要な指標です。
Poolside は、エージェント型コーディングに特化したオープンウェイトモデル「Laguna S 2.1」をリリースしました。このモデルは Mixture-of-Experts (MoE) アーキテクチャを採用し、トークンあたりで活性化されるパラメータ数は 8B です。思考モードと非思考モードの両方で最大 100 万トークンのコンテキストウィンドウをサポートしており、重みは Hugging Face で OpenMDW-1.1 ライセンスの下に公開されています。単一の NVIDIA DGX Spark でも動作可能なサイズです。
長期ホライズンコーディングベンチマークでは、Laguna S 2.1 は DeepSeek-V4-Pro-Max や NVIDIA の Nemotron 3 Ultra、Thinking Machines の Inkling など、自身よりも数倍大きいモデルたちと互角に渡り合っています。Laguna S 2.1 は Laguna XS ファミリーのスケールアップ版であり、XS 2.1 と同じ事前学習データでトレーニングされています。
Laguna S 2.1 とは何か
このモデルは、任意のトークンに対して全パラメータのうち約 6.8% を活性化します。118B の全パラメータはメモリ上に保持されますが、1 ステップあたり実際にネットワークを通過するのは約 8B です。このスパース性が、ミドルサイズのモデルでありながら大規模モデルのような振る舞いを可能にしつつ、運用コストを抑える理由となっています。
Poolside チームは、BF16、FP8、INT4、NVFP4 の重みに加え、公式の GGUF および MLX 変換版、DFlash ドラフトモデルも公開しました。トレーニング開始からリリースまでにかかった期間は 9 週間未満です。事前学習は 2026 年 5 月 22 日に NVIDIA H200 GPU を 4,096 台使用して開始されました。これは、強化学習が FP8 精度で実行された Poolside 初のモデルとなります。
(function(){var f=document.getElementById("mtp-x1c-p");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.__mtpH){f.style.height=e.data.__mtpH+"px";}});})();
パフォーマンス
思考機能を有効にした状態での Terminal-Bench 2.1 では、Laguna S 2.1 が 70.2% のスコアを記録しました。これは Poolside がまとめたリーダーボードにおいて、サイズが明記されたオープンモデルの中でトップの成績です。ただし、より大規模なモデルやクローズドシステムには及びません。また、SWE-Bench Multilingual では 78.5% を達成し、公開されている表で首位に立ちました。
Poolside が発表した完全な比較データは以下の通りです。
Benchmark | Laguna S 2.1 (118B-A8B) | Tencent Hy3 (295B-A21B) | Inkling (975B-A41B) | Nemotron 3 Ultra (550B-A55B) | DeepSeek-V4-Pro-Max (1.6T-A49B) | Kimi K3 (2.8T-A50B) | Qwen 3.7 Max | Muse Spark 1.1 | Claude Fable 5
---|---|---|---|---|---|---|---|---|---
Terminal-Bench 2.1 | 70.2 | 71.7 | 63.8 | 56.4 | 64.0 | 88.3 | 74.5 | 80.0 | 88.0
SWE-Bench Multilingual | 78.5 | 75.8 | – | 67.7 | 76.2 | – | 78.3 | – | –
SWE-Bench Pro (Public) | 59.4 | 57.9 | 54.3 | – | 55.4 | – | 60.6 | 61.5 | 80.3
DeepSWE v1.1 | 40.4 | – | – | – | 9.0 | 69.0 | – | 53.3 | 70.0
SWE Atlas (Codebase QnA) | 46.2 | – | – | – | 27.2 | – | – | 42.2 | –
Toolathlon Verified | 49.7 | – | 45.5 | 34.3 | 55.9 | – | – | 75.6 | –
最も明確な示唆を与えているのが DeepSWE v1.1 の結果です。ここにはまだ大きな伸びしろがあります。Laguna S 2.1 は 40.4% を記録しましたが、DeepSeek-V4-Pro-Max はわずか 9.0% でした。しかも、アクティブパラメータ数は約 6 分の 1 です。
Claude Fable 5 や Kimi K3 といったクローズドな最先端モデルは、いくつかのベンチマークで依然として首位を維持しています。Poolside が主張したいのは、絶対的なトップであることではなく、この重さ(パラメータ数)のクラスにおける優位性です。
最終評価セットからの軌跡データは、trajectories.poolside.ai で公開されています。
2 つの思考モードとスコアの由来
Laguna S 2.1 は「off」と「max」の 2 つのモードを持ち、デフォルトでは max モードが有効になっています。この max モードでは、モデル自身がテスト時の計算リソース(compute budget)を自動で割り当てます。現時点では、ユーザーが低・中・高の作業負荷レベルを手動で切り替える機能は提供されていません。
max モードへの切り替えにより、Terminal-Bench 2.1 のスコアは 60.4% から 70.2% に向上しました。また、DeepSWE では 16.5% から 40.4% へと大幅な伸びを見せました。ただし、この性能向上にはトークン数の増加が伴います。具体的には、DeepSWE の思考モードにおける推論プロセスで消費される完了トークンは約 249k に達するのに対し、通常モードでは 99k です。Poolside チームによると、モデルは数時間にわたり数百数千単位のトークンを消費しながらも、一貫性があり生産的な推論を継続できるとのことです。
公開された 3 つの実行事例
Poolside チームはスコアではなく実際の振る舞いを示すため、編集を加えていない 3 つの推論軌跡(trajectories)を公開しました。1 つ目の例では、モデルは空のフォルダから動作する HTML/CSS ブラウザエンジンを構築しました。この実行には人間の手を介さず 50 分間、181 ステップを要しましたが、最終的にはヘッドレス Chromium を用いて出力を検証しています。
2 つ目の例では、モデルは Poolside 自社のエージェントハッチ(harness)の最適化を行いました。その結果、実行速度が 5.2% 向上し、メモリ割り当て量が約 71% 削減されました。これは O(n²) の文字列結合をバッファ処理に置き換えることで実現されています。
3 つ目の例では、サンドボックス環境に Python が存在しないため、Perl を用いて offline で「Erdős Problem #397」の再導出を行いました。このプロセスには 68 分かかりましたが、結果は独立した再発見として認められます。なお、GPT-5.2 Pro は 2026 年 1 月に同問題を解決していますが、Laguna の知識カットオフ日は 2025 年 11 月です。
実際にデプロイできるのか?
Sizing uses the full 118B parameters, not the 8B active count, because every expert stays in memory. At 4-bit (NVFP4 or INT4) the weights need about 59 GB, which fits comfortably on a single DGX Spark's 128 GB of unified memory. At FP8 they need about 118 GB, still within a single Spark or a single H200. At BF16 they need about 236 GB, which calls for two linked Sparks or a multi-GPU datacenter node.
(function(){var f=document.getElementById("mtp-x2c-p");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.__mtpH){f.style.height=e.data.__mtpH+"px";}});})();
Poolside worked with NVIDIA to optimize inference from TRT-LLM serving to NVFP4 on Blackwell, down to a single DGX Spark. It shipped day-one support for vLLM, SGLang, and Ollama. Hosted access runs through OpenRouter, free at 256K context and paid at the full 1M context for $0.10 / $0.20 / $0.01 per 1M input / output / cache-read tokens. The model is also on Baseten, Kilo, Prime Intellect Prime Lab, and ZML.
Key Takeaways
Laguna S 2.1 is a 118B-total / 8B-active MoE coding model with a 1M-token context, open under OpenMDW-1.1.
It scores 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-Bench Multilingual, leading open disclosed-size models.
At 4-bit it runs on a single NVIDIA DGX Spark; FP8 fits one Spark or H200, BF16 needs two.
Default 'max thinking' drives most of the score, at a real token cost (DeepSWE: 16.5% → 40.4%).
Poolside は、NVIDIA と協力して Blackwell アーキテクチャ上の TRT-LLM サービングから NVFP4 への推論を最適化し、単一の DGX Spark で動作可能にしました。vLLM、SGLang、Ollama についてはリリース当日からサポートしています。
ホスト版のアクセスは OpenRouter を経由して提供されており、コンテキスト長 256K は無料です。一方、最大 1M のコンテキストを利用する場合は有料プランとなり、入力・出力・キャッシュ読み取りそれぞれ 100 万トークンあたり 0.10 ドル、0.20 ドル、0.01 ドルとなります。
このモデルはまた、Baseten、Kilo、Prime Intellect の Prime Lab、ZML などのプラットフォームでも利用可能です。
このモデルは、4,096 台の H200 GPU を用いてわずか 9 週間未満で訓練されました。Poolside が 3 ヶ月以内に発表したモデルとしては 3 作目となります。
出典:Poolside Technical(Laguna S 2.1 の紹介)、Robert McHardy氏の X 投稿、Poolside公式の X と Hugging Face モデルカード。
本記事は MarkTechPost に掲載された「Poolside が SWE-Bench Multilingual で期待以上の成果を収めたオープンウェイトのエージェント型コーディングモデル『Laguna S 2.1』を発表」という内容に基づいています。
原文を表示
Poolside has released Laguna S 2.1, a 118B-parameter open-weight model built for agentic coding. It is a Mixture-of-Experts (MoE) model with 8B activated parameters per token. It supports a context window of up to 1M tokens in both thinking and no-thinking modes. The weights are on Hugging Face under an OpenMDW-1.1 license, and the model is small enough to run on a single NVIDIA DGX Spark.
On long-horizon coding benchmarks, Laguna S 2.1 holds its own against models several times its size, including DeepSeek-V4-Pro-Max, NVIDIA’s Nemotron 3 Ultra, and Thinking Machines’ Inkling. Laguna S 2.1 is a scale-up of the Laguna XS family, trained on the same pre-training data as XS 2.1.
What is Laguna S 2.1
The model activates roughly 6.8% of its parameters on any given token. All 118B parameters remain resident in memory, but only ~8B route through the network per step. That sparsity is why a mid-size model can behave like a larger one while staying cheap to serve.
Poolside team publishes weights in BF16, FP8, INT4, and NVFP4, along with official GGUF and MLX conversions and DFlash draft models. It went from the start of training to launch in under nine weeks. Pre-training began on 22 May 2026 on 4,096 NVIDIA H200 GPUs. It is the first Poolside model where reinforcement learning ran in FP8 precision.
(function(){var f=document.getElementById("mtp-x1c-p");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.__mtpH){f.style.height=e.data.__mtpH+"px";}});})();
Performance
Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1 with thinking enabled. That places it first among open, disclosed-size models on Poolside’s compiled leaderboard, behind only larger or closed systems. On SWE-Bench Multilingual it scores 78.5%, topping the published table outright. The full comparison Poolside released is below.
BenchmarkLaguna S 2.1 (118B-A8B)Tencent Hy3 (295B-A21B)Inkling (975B-A41B)Nemotron 3 Ultra (550B-A55B)DeepSeek-V4-Pro-Max (1.6T-A49B)Kimi K3 (2.8T-A50B)Qwen 3.7 MaxMuse Spark 1.1Claude Fable 5
Terminal-Bench 2.170.271.763.856.464.088.374.580.088.0
SWE-Bench Multilingual78.575.8–67.776.2–78.3––
SWE-Bench Pro (Public)59.457.954.3–55.4–60.661.580.3
DeepSWE v1.140.4–––9.069.0–53.370.0
SWE Atlas (Codebase QnA)46.2–––27.2––42.2–
Toolathlon Verified49.7–45.534.355.9––75.6–
The clearest signal is DeepSWE v1.1, which still has real headroom. There, Laguna S 2.1 scores 40.4% against DeepSeek-V4-Pro-Max’s 9.0%, with roughly one-sixth the active parameters. Closed frontier models such as Claude Fable 5 and Kimi K3 still lead on several benchmarks. Poolside’s claim is about the weight class, not the outright top. Trajectories from the final evaluation set is published at trajectories.poolside.ai.
Two thinking modes, and where the score comes from
Laguna S 2.1 has two modes: off and max, with max enabled by default. In max mode the model sets its own test-time compute budget. Poolside is shipping without user-configurable low/medium/high effort control for now.
Max thinking lifts Terminal-Bench 2.1 from 60.4% to 70.2%. It lifts DeepSWE from 16.5% to 40.4%. Those gains cost tokens: DeepSWE trajectories run about 249k completion tokens in thinking mode against 99k without. Poolside team reports coherent, productive reasoning running for hours and hundreds of thousands of tokens.
Three published runs
Poolside team shared three unedited trajectories to show behavior rather than scores. In one, the model built a working HTML/CSS browser engine from an empty folder. That run took 181 steps across a 50-minute session, with no human intervention, and it validated its own output against headless Chromium.
In a second, the model optimized Poolside’s own agent harness. It made the harness 5.2% faster with roughly 71% lower memory allocation, replacing O(n²) string concatenation with buffers. In a third, it re-derived Erdős Problem #397 offline in Perl over 68 minutes, since the sandbox had no Python. That result is an independent rediscovery; GPT-5.2 Pro solved the same problem in January 2026, and Laguna’s knowledge cutoff is November 2025.
Can you actually deploy it?
Sizing uses the full 118B parameters, not the 8B active count, because every expert stays in memory. At 4-bit (NVFP4 or INT4) the weights need about 59 GB, which fits comfortably on a single DGX Spark’s 128 GB of unified memory. At FP8 they need about 118 GB, still within a single Spark or a single H200. At BF16 they need about 236 GB, which calls for two linked Sparks or a multi-GPU datacenter node.
(function(){var f=document.getElementById("mtp-x2c-p");window.addEventListener("message",function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.__mtpH){f.style.height=e.data.__mtpH+"px";}});})();
Poolside worked with NVIDIA to optimize inference from TRT-LLM serving to NVFP4 on Blackwell, down to a single DGX Spark. It shipped day-one support for vLLM, SGLang, and Ollama. Hosted access runs through OpenRouter, free at 256K context and paid at the full 1M context for $0.10 / $0.20 / $0.01 per 1M input / output / cache-read tokens. The model is also on Baseten, Kilo, Prime Intellect Prime Lab, and ZML.
Key Takeaways
Laguna S 2.1 is a 118B-total / 8B-active MoE coding model with a 1M-token context, open under OpenMDW-1.1.
It scores 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-Bench Multilingual, leading open disclosed-size models.
At 4-bit it runs on a single NVIDIA DGX Spark; FP8 fits one Spark or H200, BF16 needs two.
Default ‘max thinking’ drives most of the score, at a real token cost (DeepSWE: 16.5% → 40.4%).
Trained in under nine weeks on 4,096 H200 GPUs, it is Poolside’s third model in under three months.
Sources: Poolside Technical— Introducing Laguna S 2.1, Robert McHardy on X, Poolside on X and Hugging Face model card.
The post Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual appeared first on MarkTechPost.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み