Manning から LLM の後学習と RLHF を解説する新刊が発売
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Interconnects
著者が長年の経験に基づき執筆したポストトレーニング専門書『Reinforcement Learning from Human Feedback』が Manning より刊行され、オンライン教材や割引コードも併せて発表された。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月10日 22:20
AI深層分析
キーポイント
書籍の刊行と内容
著者の長年の教訓をまとめた新書『Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs』が Manning より刊行され、LLM のポストトレーニングにおける直感的な理解や歴史的経緯を解説している。
独自性の高いトピック
リジェクトサンプリング、アウトカム報酬モデル、キャラクタートレーニングなど、既存のオンライン資料が不足していた分野について、基礎的な直観に基づいて詳述している点が特徴である。
対象読者と難易度
本書は初学者向けではなく、コンピュータサイエンスの学士号を取得した者が前提知識として必要とするレベルであり、熟練者でも参照する実用的なリソースであると著者は述べている。
付随する学習リソース
書籍にはオンラインで無料公開された 12 時間のコース(スライドと YouTube 動画)、簡易コードベース、モデル比較が含まれており、8 月 19 日まで Manning で半額割引が適用される。
RLアルゴリズムの直感的理解と分類
本書はPPOやGSPOなどのアルゴリズムが、アドバンテージの正負やポリシー比率に基づいて6つの領域に分解されるという直観を教える。この理解により、新しいアルゴリズムが本物か偽物かを判断する基準となる。
重要な引用
This book was my attempt to explain in simple terms why post-training works, what trade-offs people need to make to get it right, and what misconceptions people often get stuck on.
It's not a beginner book — it's more tailored to someone who has already finished a bachelor's degree in CS — but if you master it you will be far ahead in your post-training worldview.
no new algorithm will be proven right out of the gates
Most of modern RL is a systems problem balancing a few problems — how off-policy the data is, training-inference mismatch, and throughput.
編集コメントを表示
編集コメント
ポストトレーニング技術の学習リソースとして、専門的な書籍と付随するコースがセットで提供される点は実務家にとって有益である。ただし対象読者が学士号取得者を想定しているため、初学者は前提知識の確認が必要となる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
おまけ:今回は音声解説なしの「リリース」記事です。近々、またエッセイをお届けします!
オープンモデルのトレーニングから得た教訓を記録する時間を確保するのに数年かかった末に、ついにポストトレーニングに関する書籍が完成しました!この本は Manning 社より出版され、『Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs』というタイトルで刊行されています。
本書の成り立ちを語ることは、なぜあなたがコピーを手元に置くべきかの理由を説明する有効な手段となります。
本書は当初、ウェブサイトの形式でした。ポストトレーニングにおける重要な手法について、オンライン上に解説資料が存在しない、あるいは存在しても見つからないケースが多々あることを記録したかったのです。拒絶サンプリング(rejection sampling)、アウトカム報酬モデル(outcome reward models)、キャラクター学習(character training)などがその好例です。2024 年(ドメインを取得した時期)にはすでにポストトレーニングが人気を博していたにもかかわらず、想定以上に多くのトピックでこうした情報不足が続いています。
このサイトは現在でも、これらのトピックを基礎的かつ直感的なレベルで解説している数少ない場所の一つとして、かなりの人気を集めています。
それ以外、本書の大部分は直感と歴史の解説に充てられています。LLM業界を定義する中核技術はここ数年で大きく変わっていません。本書は私が、ポストトレーニングがなぜ機能するのか、それを正しく行うために必要なトレードオフ、そして人々がしばしば陥る誤解について、平易な言葉で説明しようとした試みです。
その一環として、Interconnects の古い基礎的なブログ記事の一部を再構成し、主要な数学的トピックの背後にある物語をつなぎ合わせました。そのため、解説文は一般的な教科書よりも、著者の声(ボイス)が強く反映されたものになっています。
これは私が数年前にこの分野を始めた頃に読みたかった本です!いまだに多くの人が基本的なポストトレーニングの質問をしており、その数は加速度的に増えていると見られます。そのため、本書は非常に高く評価されるでしょう。私も現在も頻繁に参照しており、業界全体から活躍する研究者たちからも同様に利用していると聞いています。
したがって、これは初心者向けの入門書ではありません。コンピュータサイエンスの学士号をすでに取得した人向けに特化しています。しかし、この本をマスターすれば、ポストトレーニングの世界観において他を圧倒する優位性を手にできるでしょう。
読者限定の割引情報!
本書はオンライン上で無料で利用可能で、12 時間分の講義(スライドと YouTube の動画)、推奨演習付きのシンプルなコードベース、モデル完成度の比較データも付属しています。8 月 19 日まで Manning で「PBLambert」というコードを使用すると半額で購入できます。
Manning から本書を購入する
Amazon で本書を購入する
はい、本書のタイトルは少し古く感じるかもしれませんが、出版プロセスでの教訓や葛藤を経てきた結果です。しかし、中身は非常に新鮮です。私は研究活動でこの本を頻繁に参照しており、友人たちも同様に活用しているとの声を聞いています。最後の最後まで粘って、「オンポリシー蒸留」に関する章を追加しました!本書は現在、Manning と米国アマゾンから発送されており、英国版は10月に発売予定です。
- RL アルゴリズムがモデルの出力をどう変化させるかという直感的理解
単語数やページ数で計算すると、本書の約25% が強化学習(RL)に関する内容です。これは適切な配分と言えるでしょう。本書が最も力を入れているのは、さまざまな RL アルゴリズムをどう捉えればよいかという「考え方」を伝えることです。この直感的理解こそが、新しいアルゴリズムが単なる偽物なのか、それとも有望なものなのかを見極めるために不可欠です。どんな新アルゴリズムも、発表された瞬間から完全に正しいと証明されるわけではありません。
以下は、本書を読み終えた後に身につけておきたい直感的な理解の一例です。
例えば、PPO のクリッピング(clipping)の理解を深めるために工夫した図があります。結局のところ、PPO の代理目的関数は、トークンのアドバンテージが正か負か、また現在のポリシー比率の値によって、6 つの領域に単純化されます。
これは、2 つの勾配として捉えることができます。個々のサンプル(完成されたテキスト内の1つ)は、このプロット上のどこかに位置します。バッチ内の最初の勾配ステップであれば、x 軸の 1 から始まります(勾配は常に流れます)。その後、RL ポリシーの挙動を変更した後の比率更新の仕方次第で、勾配は同じ値を維持するか、0 になります。このクリッピング引数が果たす役割です。

この直感的な理解は、勾配や数値的な問題(数値不安定さ)を管理するためにシステムがどのように設計されているかに深く反映されています。数学に焦点を当てた政策勾配のセクションは非常に詳しく、過去 3 年間で耳にしたことのあるすべてのアルゴリズムを網羅しています。
- Policy Gradient の導出
- Vanilla Policy Gradient(バニラ・ポリシー勾配)
- REINFORCE
- REINFORCE Leave One Out (RLOO)
- Proximal Policy Optimization (PPO)
- PPO 目的関数の理解
- Value Functions and PPO(価値関数と PPO)
- Group Relative Policy Optimization (GRPO)
- Group Sequence Policy Optimization (GSPO)
クリップド・インポータンス・サンプリング・ポリシー・オプティマイゼーション (CISPO)
アルゴリズムの比較
- 新しい RL システムとアルゴリズムが直面する重要な要因の理解
現代の強化学習 (RL) の多くは、いくつかの課題をどうバランスさせるかというシステム設計の問題です。具体的には、オフポリシーなデータの度合い、トレーニングと推論のミスマッチ、そしてスループットです。学習者(勾配ステップを実行する GPU)とアクター(環境でロールアウトを生成する GPU)を別々の GPU で担当する非同期 RL のコアシステム設計は、ここ数年ほとんど変わっていません。

アジェンティックなタスクは、これらの基本の上にさらにインフラを追加するだけですが、本書は LLM の RL 知識がほぼゼロの状態から、システムをいじる準備ができるまでを最短でカバーするように設計されています。まずは強化学習アルゴリズムを実装する一般的な形式の解説から始めます。
pg_loss = -advantages * ratio
その後、ロス集約について学びます。これは DAPO や Dr. GRPO といった、初期の GRPO バリアントとして先駆的な役割を果たしたアイデアです。さらに、早期の RL 実験で PPO を機能させるために使われた技術である「截断インポータンス・サンプリング」についても解説します。
ポリシー勾配の基礎
ロス集約のトレードオフ
非同期 RL システム
截断インポータンス・サンプリング
例:PPO
例:GRPO
シェア
- 現代のポストトレーニングに至る歴史
分野がどのように形成されてきたかを知ることは、私にとって常に魅力的なテーマです。ある分野でエキスパートになるほど、その分野の歴史を誰よりも深く理解していることが、最良の予測を立てるための鍵となります(ビル・ガーリーも最近の著書で同様の助言をしています)。
深層学習の基盤が構築されつつあった時代、トランスフォーマーが登場した頃、現代のポストトレーニングの中核となる概念は、アライメントの分野においてすでに誕生していました。本書では、研究者たちが一般論としての RL(強化学習)を人間の嗜好データに適用する方法を学んだ 2018 年までの第 1 期、2019 年から 2022 年にかけて言語モデルへの応用方法を模索した第 2 期、そして 2023 年以降は ChatGPT が示した事例を踏まえて発展させた第 3 期の 3 つの時代をたどります。
スケーリング LLM(大規模言語モデル)の黎明期に参入した人々と同様に、10 年前にこの分野を切り開いた人々への多大な功績は、現代の技術進歩がどのように展開されたかを考えれば、決して過小評価されるべきではありません。
本書の第 2 章ではこれらを集中的に解説していますが、本書全体を通じてこのような歴史的視点に基づく考察が随所に散りばめられています。

- 「蒸留」の魔法を解く
AI 政策に関する議論が活発化する中、むしろ「ディストillation(知識蒸留)」という地味なトピックを教科書として扱う意義は大きいものです。本稿の第12章では、LLM の出力をどう活用して下流モデルを訓練するかという、業界標準的な手法について解説しています。
この技術が時折「悪意のある手段」や「地政学的競争の道具」といったネガティブな文脈で語られる一方で、300ページにわたる教科書を執筆し、攻撃対象となっている用語の広範な意味を丁寧に説明することは、むしろ緊張緩和につながる行為だと言えます。
これにより、第10章から第12章までの一連の記述は、データ業界におけるこれまで不明瞭だった慣行を、読者にとってより明確に理解できるものへと昇華させることを目的としています。
ディストillationに関する本章では、前述したテーマを引き継ぎます。具体的には、2015年頃の初期の知識蒸留研究から、Xiaomi MiMo-V2-Flash や DeepSeek V4 といったモデルに見られる「多教師オンポリシー蒸留(MOPD)」へと至るまでに必要とされた重要な転換点や、2〜3 の核心的な洞察について解説します。
- 正しくポストトレーニングを行う際に遭遇する、その他の小さな頭痛の種に関する調査
本書の後半では、過剰最適化(over-optimization)、正則化(regularization)、評価、そしてキャラクター学習について解説します。ここでは、ポストトレーニングがどのように失敗しうるか、またその兆候をどう見極めるべきかを多角的に扱います。これが、単なるコード演習や数式の羅列に終始する他の書籍との決定的な違いです。
具体的には、「なぜ SFT(Supervised Fine-Tuning)では忘却が進むのに RL(Reinforcement Learning)では一般化が可能なのか」という数学的な理由や、最先端の研究所がモデルの個性を形成するために採用している手法、そしてそれらがなぜ行き過ぎる傾向があるのかといった実態に迫ります。本書はあなたが実際に活用するツールを紹介すると同時に、それらを実装する際に直面することになる数々の課題への扉を開くものです。
本書のまとめを書く過程で、私が以前 AI 時代における企業構築について語った助言を思い出しました。「今こそ研究に目を向けるべきだ」という内容です。AI の進歩スピードは凄まじく、研究成果が最先端モデルに実装されるまでには従来より短縮され、現在はわずか3〜9ヶ月となっています。
かつてビッグテック企業では、新技術への対応に数年を要したため、研究に対して手を離す(hands off)アプローチでも問題ありませんでした。しかし、LLM の特定のニッチ分野で最先端を維持することに依存する企業の経営者にとって、現在のこのダイナミクスは企業の命運を分ける要因となり得ます。
本書が役立つ理由は、どの研究が重要かを判断する力を養うためです。つまり、「研究の味覚(research taste)」を磨く訓練になるのです。
もしこれらの内容に共感できないなら、それでも私の本を買ってください。そうすれば私が喜び、そして私がこれらに取り組むために費やした苦労が報われるからです。
この写真は本当に気に入っています。

原文を表示
Housekeeping: No voiceover on another quick “launch” post. More essays soon!
After a few long years of finding time to document my lessons from training open models, my post-training book is done! It’s published by Manning, under the title Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs.
Telling the story of the book is a useful way to explain why you may want a copy.
The book started as a website where I wanted to document key methods of post-training that had potentially no online material explaining them. If there was something, I couldn’t find it. This existed for more topics than you would expect, given post-training was already popular in 2024 (when I bought the domain), and continues to this day. Topics like rejection sampling, outcome reward models, and character training are prime examples. This has helped make the website fairly popular, as it’s still one of the few places discussing these topics at a foundational, intuitive way.
Otherwise, most of the book is about communicating intuitions and history. Much of the LLM industry is defined by core techniques that haven’t changed much in the last few years. This book was my attempt to explain in simple terms why post-training works, what trade-offs people need to make to get it right, and what misconceptions people often get stuck on. To do this, some of the older, foundational blog posts on Interconnects were reworked to stitch together the story behind key mathematical topics. For this reason, a lot of the explanatory text is likely higher voice than your average textbook.
This is the book I wanted to read when I was getting started a few years ago! With how many people still ask me basic post-training questions, in fact a population that’s accelerating in size, I suspect this book will be very well received. I still use the book regularly and hear from established researchers all over the industry that they do too. So, it’s not a beginner book — it’s more tailored to someone who has already finished a bachelor’s degree in CS — but if you master it you will be far ahead in your post-training worldview.
Share
A discount to readers!
The book is also freely available online and comes with a full 12 hour course (slides + video on YouTube), a simple code-base with suggested exercises, and model completion comparisons. It’s 50% off until August 19th on Manning with the code PBLambert.
Buy my book from Manning
Buy my book on Amazon
Yes, the title of this book is a little outdated — I’ve learned some lessons and fought some battles with the publishing process — but I can guarantee the content is very fresh. I regularly reference the book for my research work, and hear from friends that do the same. I added a section on on-policy distillation at the last possible moment! The book is shipping from Manning and Amazon US now, and from Amazon UK in October.
- Intuitions for how RL algorithms change the outputs of models
By word or page count, the book is about 25% RL. This seems appropriate. If there’s one thing the book is doing it’s teaching people how to think about various RL algorithms. This intuition, from the policy-gradient theorem to PPO to modern versions like GSPO and CISPO, are crucial to understanding if a new algorithm is fake or has potential (no new algorithm will be proven right out of the gates).
Below is an example intuition you should be able to follow after reading.
For example, here’s a fun figure that we’ve iterated on for the PPO clipping understanding. At the end of the day, PPO’s surrogate objective reduces to six regions. These can be seen as two gradients, when the advantage for a token is positive or negative, depending on the current value of the policy ratio. For an individual sample in a completion, it lives somewhere on this plot. If it was the first gradient step in the batch, it starts at 1 on the x axis (gradient always flows), then depending how the ratio updates after changing the RL policy behavior, the gradient is either the same or becomes 0 (which is what the clipping arguments are for).

This intuition filters very closely into how the systems are designed, in order to manage the gradients and numerical issues they tend to cause. The math-focused policy-gradient section is pretty thorough, covering all the algorithms you’ve likely heard about in the last 3 years:
Deriving the Policy Gradient
Vanilla Policy Gradient
REINFORCE
REINFORCE Leave One Out (RLOO)
Proximal Policy Optimization (PPO)
Understanding the PPO Objective
Value Functions and PPO
Group Relative Policy Optimization (GRPO)
Group Sequence Policy Optimization (GSPO)
Clipped Importance Sampling Policy Optimization (CISPO)
Comparing Algorithms
- An understanding of the crucial factors facing new RL systems and algorithms
Most of modern RL is a systems problem balancing a few problems — how off-policy the data is, training-inference mismatch, and throughput. The core systems design, asynchronous RL with separate GPUs for the learners (the GPUs which take gradient steps) and actors (the GPUs which generate the rollouts in the environment), has been similar for a few years.

Agentic tasks are only adding more infrastructure on top of these fundamentals. The book is designed to be the simplest resource to start from roughly 0 LLM RL knowledge and be ready to tinker with the systems. It starts with basics, such as explaining the general form of implementing an RL algorithm:
pg_loss = -advantages * ratio
It continues with teaching you about loss aggregation — the idea that spawned DAPO and Dr. GRPO as some of the seminal, early GRPO variants — and truncated importance sampling — the technique used to make PPO work in early RL experiments.
Policy-Gradient Basics
Loss Aggregation Tradeoffs
Asynchronous RL Systems
Truncated Importance Sampling
Example: PPO
Example: GRPO
Share
- The histories that lead to modern post-training
Knowing how a field came to be has always been fascinating to me. As you become an expert, knowing the history of your field better than anyone is what lets you make the best predictions (Bill Gurley gives similar advice in his recent book). In a time when the foundations of deep learning were being built, like the transformer, the core of modern post-training was also born in the alignment field. The book will walk you through 3 eras, when researchers learned to do RL on preferences generally until ~2018, spent a few years learning how to apply it to language models from 2019 to 2022, and from 2023 on exploited the examples set by ChatGPT. Much like those early to scaling LLMs, the people who created this field a decade ago deserve incredible credit for how modern progress has unfolded.
Chapter 2 of the book is a crash course on this, but the book is littered with this type of thinking.

- Dispelling the magic of “distillation”
It’s really nice to have a boring textbook chapter on distillation given the broader AI policy discussions ongoing. This one, chapter 12, explains the various industry-standard ways that outputs from an LLM are used to train downstream models. When the technique is often described in nefarious ways, and as a tool of geopolitical competition, it’s a deescalatory action to break out a 300 page textbook to explain to someone how broad the term they’re attacking is.
With this, chapters 10 through 12 are all about making some opaque practices of the data industries clearer to readers.
The distillation chapter continues the theme I outlined above, explaining the key changes that needed to be made to transition the early knowledge distillation literature of 2015 to 2-3 key insights that got us to the multi-teacher on-policy distillation (MOPD) of models like Xiaomi MiMo-V2-Flash and DeepSeek V4.
- A survey of all the other little headaches you encounter when trying to do post-training right
The second half of this book goes into a tour of over-optimization, regularization, evaluation, and character training, which is all about the many ways post-training can go wrong and what you need to stay on top of. This is what differentiates the book most from those that are just a list of code exercises and equations, but it explains things like why RL generalizes when SFT forgets (at the math level) or which techniques frontier labs use to shape the personalities of the models and why those often go too far. The book presents the tools you will use and then opens up the floodgates of all the challenges you’re going to face when you actually try to put them to use.
As I wrote the takeaways for the book, I was reminded of an old piece of advice I gave for the AI era of building companies, and how people do need to care about research right now. With the pace of progress in AI, the time it takes for a research paper to land in a frontier model is 3-9months. Previously in big tech, that would be years, so it was fine to take a hands off approach to new research. For people with companies relying on being at the frontier in specific niches of LLMs, the dynamic today can make or break the company. This book is useful because it trains you at understanding which research matters — it helps you develop research taste.
If none of this resonates with you, you should buy my book because it makes me happy and I work very hard on all of this.
I really like this photo.

関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み