AI 教科書執筆経験者が AI の優越性を問う
本文の状態
日本語全文を表示中
詳細モードで約17分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Interconnects
著者は長文非小説文における LLM の停滞が懸念される現状を指摘し、知識の整理能力が欠如しているため科学分野での革命的洞察には至らず、人間によるガイドが必要であると論じる。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 22:24
AI深層分析
キーポイント
創造的執筆と非小説文の対比
LLM は高ボイスな創作では劣るが、非小説文やコピーライティングでは有用であり、その欠陥は知能不足やトレーニング問題に起因すると考えられてきた。
科学分野での長文記述の限界
現在のモデルは確立された科学知識を組織化し説得力のある形で提示する能力に欠けており、これが自律的な科学問題解決への自然な前提条件である。
知識整理とエントロピーの増加
知識の整理は圧縮であり洞察を生むが、現在の LLM は長文非小説文においてエントロピーを増加させる傾向があり、これを無限に積み重ねることは不可能である。
狭義な科学分野での進展と一般化
数学などの特定領域では顕著な進歩が見られるが、これらの基礎的な知識問題における観察から、モデルの汎用性には驚くべき欠如があることが示唆される。
非-fiction 執筆におけるモデルの停滞
コーディングや数学などの他のタスクで超人的な進歩が見られる一方で、文章作成能力は著しく停滞しており、長文技術書では構成が混乱し誤った概念を提示する傾向がある。
重要な引用
Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future.
Organizing knowledge is a compression. This compression is needed to make insight.
Today's LLMs increase entropy in long-form non-fiction writing, and I don't see how that can be stacked on top of itself endlessly.
Writing well is a very hard task!
編集コメントを表示
編集コメント
著者は Anthropic の Claude がリーマン予想への進展を示した直後に、AI の科学分野における本質的な限界を指摘しており、技術楽観主義に対する冷静な視点を提供している。この分析は、LLM を科学研究の主力として位置づける際の過信を防ぐ重要な示唆を含んでいる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI による文章生成への批判は数多く存在しますが、その多くはこのブログのような、創造性が高く、著者の個性が強く反映された文体を対象としています。私自身もそう主張してきましたが、良質な文章とは「高い声(=強い個性)」を持ち、明確な視点や深い人間性の表現があり、また選んだ言葉を通じて読者に思考のプロセスを垣間見せるものだと考えられています。
しかし、LLM が対話型アシスタントから、より洗練されたツールへと進化していくにつれて、むしろ「モデルがインスピレーションを与える文章を生み出す」という目標は後退しているように思えます。
一方でノンフィクションライティングの領域では事情が異なります。以前、サーム・アルトマン氏も最近のポッドキャストで言及していた通り、「埋め合わせ用のコピーテキスト」を作成する能力こそが、LLM が持つ真に有用な機能の一つでした。当時は、この分野での欠陥はモデル自体の知能不足や訓練上の問題によるものだと考えられており、非フィクションや解説文も、技術進歩の速いペースの中でいずれ克服されるだろうと楽観視されていました。
しかし、ここ数年、私は執筆アシスタントとしてこれらのモデルを実際に使い続けてきました。確かに性能は向上しましたが、いったい何がボトルネックになっているのか、今一度振り返る価値があります。
長文のノンフィクション分野でモデルの進化が停滞していることは、近い将来に AI が自律的に大規模な科学課題を解決するはずだと期待している人々にとって、警鐘を鳴らすべき事態です。現在のモデルは、自らの専門領域にある確立された科学的事象さえも整理し、説得力を持って提示することに苦戦しています。
これは、広範で開かれた問題を自律的に解決できるようになる前に、まずマスターすべき自然な前提条件のように思えます。この課題が解決されるまで、科学分野における LLM の進歩は、容易に収穫できる果実(ロー・ハンギング・フルーツ)の獲得や、異なる分野間の遠いつながりの統合にとどまり、革命的な洞察には程遠いものとなるでしょう。
これは、AI の進展に対して非常に楽観的な立場にある者にとって、やや物議をかもす見解かもしれません。特に Anthropic が Claude によるリーマン予想への進捗を発表したブログ記事を書いたその日にこのように記している点も、皮肉な状況です。科学問題には広大な範囲があり、現在の AI モデルが多くの人が考えるほど広くカバーできているとは私は思えません。
知識の整理とは、本質的に圧縮行為です。そして、洞察を得るためにはこの圧縮が不可欠です。しかし、今日の LLM は長文のノンフィクションにおいてエントロピー(無秩序さ)を増大させる方向に働いており、それが無限に積み重ねられるなどあり得ないと私は考えます。将来的には、AI は人間をガイド役として頼らざるを得なくなるでしょう。
私は、数学における極端な進歩のような特定の科学分野からの翻訳が、一貫した広範な進展へとつながる可能性について、まだ非常に楽観的です。LLM は科学者たちがこれまで使った中で最も強力なアシスタントです。しかし、こうした確固たる低レベルの知識問題に対してモデルがどのように動作するかを文章化して観察すると、驚くべき一般化能力の欠如に気づかされます。
より文脈を理解してもらうために、私は最近、ポストトレーニング用の教科書『Reinforcement Learning from Human Feedback』を完成させました(Manning または Amazon で購入可能)。この作成には LLM を多様な形で活用しました。数式の LaTeX 形式の調整や、徹底的な校閲、TikZ(LaTeX)や Python などのプログラミング言語のための図表作成などです。
なぜモデルの文章化能力が停滞しているのでしょうか?
非-fiction(ノンフィクション)の執筆におけるモデルの進歩については、もっと大きな成果を期待していました。2024 年の状況を見ると、2026 年にノンフィクションの本を出版するのは馬鹿げたことだとさえ思いました。
現在、執筆能力で最も著名なモデルは古参です。例えば OpenAI の GPT 4.5 や Moonshot の Kimi K2 が挙げられます。これらのリリース前後で、モデルはコーディングや数学などの他のタスクにおいて「まあまあのレベル」から「超人的なレベル」へと飛躍しました。
より近い、しかしまだ不完全な比較としては、検索やリサーチの分野での進歩が挙げられるでしょう。ここではモデルは「不可能だった状態」から「それなりにはできる状態」へ変化しました。他のスキルにおける進歩のペースは急峻ですが、「上手に書く」という能力は、それらのスキルとは直交しているように感じられます。
執筆を単に無視されているだけではないと思います。むしろ課題が難しく、それを直接改善するための良質なトレーニングデータが不足しているのです。
AI モデルの執筆能力を向上させるための「低 hanging fruit(容易な成果)」は確かに存在します。Claude Code といった専用フレームワークや、プロンプト、モデルが出力に対してより多くの推論トークンを費やすように設計された学習環境などがその例です。しかし、これらのアプローチが能力に乗数的な影響を与えるとは考えにくいです。
上手に書くことは非常に難しいタスクなのです! 知的活動の一大分野である執筆において、推論時間のスケーリング(inference-time scaling)が解き放たれていないのは残念なことです。いずれにせよ、モデルが得意とする領域と「文章を書くこと」は、根本的に異なる性質のもののように思われます。
現在、長文の技術記事作成において AI モデルは依然として不十分な結果を示しています。1 文程度であれば正確に記述できることもありますが、章全体を生成させようとすると、言葉遣いが混乱し、構成が不明瞭になり、全体的におかしい印象を与えてしまいます。必要のないところで可愛げを出そうとするあまり、かえってランダムな概念的誤りを生じさせているのです。
近い将来のモデルは、特に規模が大きくなることで世界知識をより多く保持できるようになるため、細かいミスを修正する能力は向上すると考えられます。しかし、その知識をどう活用するかという点で劇的な変革が起きるとは思えません。
例えば、GPT モデルは長年にわたり、誤字や軽微な問題の発見に優れていました。私は本書のほぼ完成した原稿を PDF として GPT-5.5 Pro に渡しましたが、200〜300 ページにわたる原稿の中から、深く驚くべき細かい誤字を発見してくれました。
一方、Claude モデルは編集者としての役割においてより有用です。これらは優れた審美眼を持ち、タスクのメンタルモデルをよりよく理解しており、執筆の行き詰まりを打破するための興味深い提案を数多く提示してくれます。
上記の例に共通するテーマがあります。モデルは、文や数式、図といったコンテンツの各単位をすべてチェックする方法を知っており、特定のセクションでユーザーが迷子になるのを防ぐことができます。しかし、これらのスキルがあるにもかかわらず、モデルはコンポーネントを再確認してつなぎ合わせる作業にはあまり得意ではありません。多くの追加処理を重ねるうちに、エラーが蓄積していくように感じられます。
かつて私たちは数学やコードの分野でこうしたエラーに対処してきましたが、振り返ってみると、RLVR はそれを大幅に削減する真に魔法のような解決策でした。
Interconnects AI は読者支援型の出版物です。購読をご検討ください。
現在のモデルをライターとして活用する方法
私の著書には、AI モデルから生まれた技術的な説明文が数行含まれています。全体の 1% に満たない量ですが、私がそれらを採用したのは、本心からその文章の質に感動したからです。真の専門家である自分が、「これは読者にとって必要な一文だ」と感じた場合、それを書籍に含めることは不正行為には当たらないと考えました。
特に編集プロセスでは、細部まで目を光らせながら進めていましたが、私の多忙な状況の中で本書が完成するかどうか不安を抱えていました。そんな中で AI を活用した道は、非常に価値のある前進の手段となりました。
例えば、編集者からの質問リストを LaTeX ファイルに埋め込み、特定の区切り文字(\editor{})で囲んでおきました。Claude Code に各コメントへ移動させ、前後の文脈を出力させてもらい、単純なタイプミスなのか、それともより微妙な修正が必要なのかを確認してもらいました。私はその場で回答テキスト(挿入する文章)を書き込むか、あるいは Claude に提案を求めてから修正を行いました。知的には非常に集中力を要する編集プロセスでしたが、本を改善するための楽しい方法でもありました。時には Claude の提案したフレーズがそのまま本書に採用されることもありました。
これは明らかにスリップリー・スロープ(傾斜路)です。私が AI の提案をいくつか受け入れたのは、手稿の二回目の完全レビュー中のある時点でした。感情的にはプロジェクトは完了したように感じましたが、実際にはまだやるべき仕事が残っていました。教科書執筆のプロセスを終えて今、私は Interconnects における私の執筆ルールを深く感謝しています。「コンテンツに AI の出力を一切使用しない」という明確で厳格なルールです。自分自身の声が高く響き、プロセスそのものが極めて価値あるものとして尊重されるような書き方をする方がはるかに楽しいものです。ただし、標準的なリファレンス書物の作成が「楽しい活動」であるとは通常言われません。多くの人が AI ツールを足場(クランチ)にしてしまう理由も理解できます。彼らの書くことの多くが、単にスペースを埋めるための出力であり、目的を達成するための手段ではないからです。
私は学び、感じ、表現するために、量を書き続けることに動機を感じています。
私も科学論文の執筆において、同様のバランス感覚を求め続けています。AI モデルは、すでに頭に入っている関連研究や背景説明などの定型部分を起草する際には非常に役立ちます。しかし、要約、序論、実験結果、結論といった部分に頼りきるのはもったいないことです。これらは論文の物語と魂が伝わるべき場所であり、私たちの研究が本当に何についてのものであるかを理解できる重要な箇所だからです。
AI モデルを活用して非小説系の執筆作業を生成・検証できたおかげで、私はより大きなネットバリューを生み出せたと確信しています。数式の作成は容易になり、リポジトリのリファクタリングや言語間の移植など、多くのタスクを支援してくれます。最初は非常に楽しかったのですが、出版プロセスの長さや分野の急速な変化に疲れを感じるようになると、その喜びも薄れていきました。
なぜこのケースで AI が不可欠だったのか、具体例を挙げます。私はウェブ版に対する読者のフィードバックと、Manning の編集チームによるフォークされたコピーのレビューという二つの異なるプロセスに対応するため、Markdown 版と LaTeX 版の両方を同時に維持する必要がありました。AI エージェントがなければ、この二つのバージョン間の同期作業は少なくとも五倍の時間がかかっていたでしょう(実際にもすでに数十時間を要するタスクです)。
この話と深く結びついているのが、エージェントを使っても理解のスピードは上がらないという事実です。将来価値を持つのは、直感や審美眼、勘といった「理解」そのものです。非-fiction(ノンフィクション)執筆に AI を使うことは、こうした成長を阻害します。さらに、すでに専門家でなければ、AI の生成物にある欠陥にも気づけません。
私の場合、脳内の知識をページ上に書き出す緊急性が非常に高かったため、AI モデルを使うことが有益な手段となる場面もありました。本書の動機の一つは、ウェブ上にほとんど存在しない重要なポストトレーニング手法(リジェクト・サンプリングやキャラクター学習など)に関する単一の参照資料を作ることでした。
この教科書はコミュニティへの還元が主目的であり、どのような形であれ完成させたことが大きな喜びでしたので、許容できる範囲だと感じました。もちろん、より多くの人的労力を投入すれば、私自身ももっと学べたでしょうし、製品もわずかに改善できたかもしれません。しかし決定的だったのは、出版される頃には本書の内容が陳腐化してしまうという懸念です。一方では AI モデルの能力への不安があり、他方ではこの分野の進化速度の速さがありました。
しかし、これは大きな誤りでした。結果に非常に満足しており、2024 年の開始時よりも、現在では本書の永続性に対する自信を持っています。それは、AI モデルがノンフィクション執筆において期待されたほどの成果を出せていないからです。
Share
技術文書写作の今後
これらのモデルは驚異的なツールであり、知識をさまざまな形式で表現することを可能にします。クリエイティブな埋め込み素材や背景資料の作成には特に優れており、例えば教師が講義の話題として利用するためのスライドの初稿などがその一例です。これにより、あらゆる知識を一つの媒体から別の媒体へ変換することが可能になります。
非-fiction や参考書的な教科書の執筆における初期段階には、このブログ記事のように「高品質な」投稿を書く感覚に近い部分があります。組織化を進め、新しい知識の骨格となる核心部分を提示する過程でこそ、新たな知見が生まれます。ここが洞察を要する部分であり、LLM はこれを代替できる点においてまだ遠く及んでいません。
上記の段落とそれ以前のセクションの核心は、世界の専門家たちが AI モデルを活用して教科書のわずかな部分を執筆し、より多くの知識を世界に共有することを望んでいるという点です。しかし問題は、現状では AI モデルを使って節約できる労力が 10〜20% に過ぎず、この割合がすぐに多数派になることはないだろうと見ていることです。
また社会的な圧力も存在します。人々は LLM が最高の個別最適化された教育者だと期待しており、そのため書籍や教育コンテンツの作成は無意味だと考えてしまうのです。しかしこれらの意見は古くなりつつあります。なぜなら、最高品質の教育コンテンツには常に深刻な不足があり、それは昔から変わらない事実だからです。AI は既存のコンテンツを学生の好みに合わせて加工する点では優れていますが、ゼロからコンテンツを生み出すことについてはまだ至っていません。
その一方で、私たちは行き詰まりを感じています。AI モデルは非-fiction 執筆における平均的な作業量を全体的に減らす方向に進むでしょうが、同時に素晴らしい表現を可能にする可能性も秘めています。しかし、プロジェクトを開始し、最後までやり遂げる人の数は減るかもしれません。
したがって、今後 2〜5 年の間において、最良の教科書は依然として人間の手によって丹念に作られると予想しています。その先については確信が持てませんが、これらのモデルが持つ膨大な知識量や、それを流暢に出力する構造的な傾向を考慮すると、この予測は多くの人が想像していたよりも長い期間です。
能力に関する結論としては、AI モデルは以下の 2 つの文脈において特に優れています。1) 真に検証可能な分野、および 2) 大量のコンテキストを与えられ、小さな修正(バグの発見や非常に特定の数学問題の解決、フィードバックの提供など)を行う場合です。これは、自由な形式で文章を生成する行為とは対照的です。長文執筆は創造的執筆よりも先に限界に達するでしょうが、モデルが未定義の問題において知識の全容を表現できないという事実は、重要な示唆となります。私たちが AI モデルを「データセンター内の天才」として、大規模な科学的問題を解決する存在へと進化させようとする試みにおいて、これは根本的な制約のように思われます。
LLM が優れた編集者でありながら、同時に劣ったライターである理由について、素晴らしい記事がありました。
この問題の一部は、人々がモデルをどう活用しているかにかかっています。Claude Code(またはチャットアプリ)で「金魚について素晴らしい詩を書いて」とClaude 3.5 Sonnet に依頼すると、モデルはすぐに回答を吐き出します。私はそのプロセスについて尋ねてみました。「何か良いものを返す前に、下書きや推敲を行うメモ帳のようなものを使っていますか?」と。しかし答えはノーでした。モデルは推論トークンの中で最小限の計画を立てるだけで、自己回帰的に詩を生成しているのです。推論時のスケーリングを活用しようともせず、このタスクを難問として扱おうともしていません。
おそらく多くの人が、文章作成においてこのような使い方をしているでしょう。そのため、結果が平凡なのは驚きではありません。最高の成果を引き出すには、モデルに対して非常に詳細なプロンプトを与え、回答する前に徹底的に作業を行わせ、他の評価モデルの意見も参照させた上でテキストを返させる必要があります。現在、モデルには確かな能力が備わっているのですから、長文作成を改善するための簡単な方法が存在します。
原文を表示
There are a lot of criticisms of AI writing, but most of them are focused on more creative, high-voice writing like this blog. Those — including my own piece — often argue that it is because good writing is high-voice, has a point of view, has a deep human expression that needs to come across, and or a process of thinking that you peek into with the chosen words. As LLMs get more refined as tools, rather than conversational assistants, I think we are actually going backwards on our goals of having models produce inspiring writing.
On the other side of things is non-fiction writing. Filler, copy text was one of the genuinely useful abilities of an LLM (Sam Altman said so much about the early business of GPT-3 on a recent podcast). It has seemed like any flaws here were mostly down to a general lack of intelligence in the models, or some other training issue, and all non-fiction and explanatory text would get obliterated by the rapid pace of progress eventually. Having worked with the models as a writing assistant over the last few years, they’ve gotten a bit better, but it’s worth reflecting on what’s holding them back.
Models being stagnant in long-form, non-fiction writing should be alarming to those reliant on models autonomously solving grand, open science problems in the near future. The models today struggle to organize and compellingly present some of the most established science in their area. This seems like a natural prerequisite that we should expect the models to master before they can solve broad, open-ended problems on their own. Until this is solved, the progress of LLMs for science will look closer to solving low-hanging fruit and merging distant connections across fields, rather than any sort of revolutionary insight.
This is a somewhat controversial take for someone who is very optimistic about AI’s progress, especially writing it on the day that Anthropic published a blog post on Claude making some progress on the famous Riemann Hypothesis. Scientific problems have a vast breadth, and I don’t think current AI models have as much coverage as many think.
Organizing knowledge is a compression. This compression is needed to make insight. Today’s LLMs increase entropy in long-form non-fiction writing, and I don’t see how that can be stacked on top of itself endlessly. They’ll be reliant on humans acting as sort of guides.
I am still very optimistic about translation from these narrow forms of science, like the extreme advancements we’ve seen in math, into consistent, broader progress — LLMs are the most powerful assistants scientists have ever used. I first need to explain how observing the models work on such grounded, low-level knowledge problems in writing makes me see a surprising lack of generalization.
For more context, I just finished writing a post-training textbook, Reinforcement Learning from Human Feedback (buy on Manning or Amazon). I used LLMs in many ways to support this, from helping wrangle LaTeX formatting for equations, doing extensive copyediting, and creating diagrams for programming languages like TikZ (in LaTeX) or Python.
Share
Why have models stagnated in writing quality?
I would’ve expected way more progress on non-fiction writing from the models. I almost thought I would look dumb publishing a non-fiction book in 2026, given how things looked in 2024. Today, some of the most famous models on writing ability are pretty old, examples include OpenAI’s big GPT 4.5 and Moonshot’s Kimi K2. In and around these releases, the models have gone from okay to superhuman at other tasks like coding and mathematics. Maybe a closer, but still imperfect, comparison is how the models went from incapable to decent at search and research tasks. The pace of progress on most other skills is steep, but writing well feels orthogonal to most of them. I do not think writing is just ignored, but rather it’s challenging and lacks good training data to specifically intervene on it.
There is certainly some low-hanging fruit for making AI models better at writing — such as specialized harnesses like Claude Code, prompts, and training environments that make models spend a lot more inference tokens on the output, but I don’t think these will have a multiplicative impact on ability. Writing well is a very hard task! It’s a shame that we haven’t unlocked inference-time scaling for one of the great intellectual pursuits. Regardless, writing seems very different than what the models are good at.1
Today, the models seem genuinely horrible at long-form technical writing. They can get a sentence right, but if you try and get them to write an entire chapter it’ll be a mix of sprinkled with confusing wording, muddled in its organization, and generally a bit off. They try to be too cute where they don’t need to be and in the process make random conceptual errors. The models in the near future will get much better at the small errors, especially as models get bigger — which allows them to hold more world knowledge — but I do not expect their ability to utilize it to transform.
For example, the GPT models have been incredible at finding typos and minor issues for a long time. I passed a near-final draft of my book as a PDF to GPT 5.5 Pro and it found deep, surprising minor typos across the manuscript that is 200-300 pages.
On the other hand, the Claude models have been much more useful as an editor. They have a lot more taste, tend to understand the mental model of the task better, and have more interesting suggestions to unstick the different forms of writer’s block.
The examples I’ve given above all have a sort of consistent theme. The models know how to check every unit of content, in this case usually a sentence or equation or figure, or make one, specific section where you are caught. With these skills, they don’t do a good job revisiting components and stringing them together as they make many additions on top of each other. It feels like a sort of irreducible compounding errors. We used to deal with these errors in math and code, but reflecting on it, RLVR has been a truly magical solution in reducing them.
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
Getting value out of current models as a writer
I’m willing to share that there are a few technical explanation sentences in my book that came from an AI model — well less than 1% — they’re there because I really loved them. I let myself consider including some AI tokens in the book, as it didn’t feel like cheating if I, as a true expert, felt that the sentence was what the reader needed. Especially in the editing process, where I had a very close eye on things and plenty of concern on if my book would ever be done with all the things I have going on, it was an extremely valuable path forward.
For example, I had a list of questions from my editor interspersed in a LaTeX file with a specific delimiter like \editor{}. I would have Claude Code navigate to each comment, print the context before and after, and let me know if it was an easy typo fix or something more nuanced. I would write a response — the text to insert — or ask Claude for suggestions before fixing it. Intellectually it is a very focusing process of editing, it was a fun way to improve the book. Sometimes phrases from Claude’s suggestions are what made it into the book.
It is definitely a slippery slope and when I accepted a few AI suggestions it was at the point where I was going through my second full-manuscript review. Emotionally the project felt completed but I had more work to do. Coming out of the textbook-writing process I so deeply appreciate the cut and dry rule I have for my writing on Interconnects to never use AI outputs in the content. It is way more fun to write in a way that is only you — high voice, valued so deeply for the process — but writing a standard reference is not really an activity known for being fun. I see why people turn AI tools into a crutch when most of their writing is just an output to fill space, rather than a means to an end. I am motivated to write voluminously to learn, to feel, and to express.
I am working through similar balances in my scientific work too. AI models are great for repetitive pieces of the paper, like drafting a related work or background section that you know by heart, but using them for the abstract, introduction, experiments, or conclusion is a shame. Those are where the story and soul of the work is communicated — it’s where you learn what your research is really about.
I am confident I created a lot more net value by being able to have AI models create and check my non-fiction writing work. They make writing equations trivial, can help refactor the repository, port between languages, and many other things. At the beginning, it was very fun, until I was a bit worn down by the length of the publishing process, watching the field move on.
For an example of why AI was crucial in this case, I had to maintain Markdown and LaTeX versions of my book simultaneously in two spots, as readers gave feedback on the web version and my Manning editorial team reviewed a forked copy. Without AI agents, syncing between the two of them would’ve easily taken me five times as long (and this task took tens of hours already).
Something intertwined with this story, which I stumbled upon when thinking about agents, is how your pace of understanding won’t increase by using agents. That understanding, in the form of intuition, taste, instinct, etc. is what will be valuable in the future. Using AI for non-fiction writing takes away from that progression. Doubly, if you weren’t already an expert you won’t be able to catch its flaws.
In my case, I felt such an urgency to dump the knowledge out of my brain onto the page that there were times that using the AI models was a worthy tool. Much of the motivation of my book was to have a single reference for important post-training methods like rejection sampling or character training, where very little exists on the web.
This textbook was so much of giving back to the community, that it was just such a win to complete it in any form, that I felt it was okay. I would’ve learned more and the product could’ve been marginally improved with more human effort, I am sure. The determining factor was that I felt like the book was going to be aged out by the time it was published, a fear of AI model’s capabilities on one side and how fast the field moves on the other.
This turned out to be really wrong? I’m very happy with the result and I’m more confident in its staying power now than when I started in 2024, as the models have so failed to live up to the hype in non-fiction writing.
Share
Where technical writing goes from here
The models are incredible tools, they let you express knowledge in different forms. They’re wonderful for creating creative filler or background material — e.g. the first draft of slides whose real value is being a talking point for the teacher to lecture over — that let any knowledge be transformed from one medium to another.
There’s some subtle, early phase of writing a non-fiction or reference textbook that feels a bit closer to writing a high-voice blog post like this. When pushing through the early organization and the presentation of the core skeleton new knowledge is created. This is the part that takes insight, and the LLMs are far behind in being able to replace it.
The crux of the above paragraph and preceding section is that I would be happy if more of the world’s experts used AI models to write a tiny bit of their books in order to get more of their knowledge shared with the world. The problem is that you can only use AI models to save 10-20% of the effort today, and I don’t see that percentage becoming the majority anytime soon.
There’s also the social pressure, where people expect LLMs to be the best, personalized educators out there, so they think working on a book or educational content is pointless. I think some of these opinions are aging out, as there’s a massive dearth in the highest quality educational work — and there always has been. AI is great at manipulating said content into the form that suits the student, not creating the content from scratch.
In the meantime I feel that we are stuck in a frustrating local minimum, where AI models are going to on net reduce the average effort spent on non-fiction writing, but they could enable great expression. Fewer people will start and push through.
So, in 2-5 years I still expect the best textbooks to be heavily crafted by the human hand. I’m not sure after then, but that’s longer than many would’ve predicted, given just how much knowledge these models have and their structural propensity to stream it.
As for a conclusion on capabilities, the models are great in two contexts: 1) any truly verifiable domain and 2) when given a ton of context and making a small edit — like finding a bug or solving a very specific math problem or giving feedback — not generating prose in an open-ended manner. Long-form writing will definitely fall before creative writing, but it’s a strong tell that the models are not able to express the full extent of their knowledge in underspecified problems. As we try to push the models to be something like “geniuses in a datacenter” solving grand scientific problems, this seems like a fairly fundamental limitation.
had a great piece on why LLMs make good editors, while being bad writers too.
1Part of this is in how people use the models. If you ask Claude Fable 5 in Claude Code (or the chat app, I’m sure too): “write me a great poem about a goldfish,” the model will quickly spew out an answer. I asked the model how it did this, and if it had a sort of scratchpad it wrote to and iteratively updated before returning something good, and it said no. It made a minimal plan in its reasoning tokens and then autoregressively generated a poem. It’s taking no advantage of inference-time scaling or approaching it like a hard task.
This is how most people surely use models for writing, and it’s no surprise the results are mediocre. The way to get the best results out of them would be to prompt the models very heavily, get them to work extensively before answering you, and reference other judge models’ opinions before returning you the text. There’s a simple way to make the long form better, given the models have some genuine skills right now.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み