神話の物理学(25 分読み)
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
ラファ・シュウィンガーは、Claude の Mythos と Fable を逆解析し、競争優位性の源泉がアーキテクチャではなく環境基盤であると論じた。テキストや計算資源が不再重要となる中、検証可能な報酬が新たな決定的要素となっている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ラファ・シュウィンガーは、モート(参入障壁)がアーキテクチャではなく環境ファウンドリにあると主張し、Claude のミソスとフェーブルを逆解析します。その能力分解モデルでは、基盤となる基礎の上に抽出可能な信号 gradeable signal を乗じたものが能力となり、テキストや生計算資源がもはや希少でなくなった今、検証可能な報酬 verifiable reward が決定的な希少入力となっています。
このレシピは、高密度事前学習 dense pretraining、報酬ハッキングの健全性が実際の制約となる GRPO スタイルの検証者強化学習 verifier RL、32K のアクティブコンテキストで百万トークンウィンドウに勝る学習された文脈折りたたみ context-folding を備えた長期ホライズンのプロセス報酬 long-horizon process rewards、そして試行時の計算リソース test-time compute を努力度合いのダイヤルとして露出させる Best-of-N 戦略を積み重ねたものです。
原文を表示
Rafa Schwinger reverse-engineers Claude Mythos and Fable by arguing the moat is not architecture but the environment foundry, with capability decomposing as base foundation times gradeable signal extracted on top, and verifiable reward becoming the scarce decisive input now that text and raw compute no longer are. The recipe stacks dense pretraining, GRPO-style verifier RL where reward-hacking soundness is the actual binding constraint, long-horizon process rewards with learned context-folding that beats million-token windows at 32K active, plus best-of-N test-time compute exposed as an effort dial.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み