GPT-5.6 Sol が推論・速度・コストで最強と評価
TLDR AI は、OpenAI の最新モデル「Sol」と「Fable」の特性を比較分析し、前者を実行役のワークホースとして、後者を高度な推論と計画を行うアーキテクトとして位置づける戦略的活用ガイドを提供している。
キーポイント
モデルの役割分担:Sol と Fable の特性比較
Fable は「最も賢い」モデルとして、複雑な推論、計画立案、および信頼性の高いエージェントとしての役割を担う一方、Sol は具体的なタスク実行、コンピュータ操作、ウェブ検索において優位性を示す実務担当のワークホースと位置づけられる。
ハイブリッド活用戦略の推奨
最も優れた回答を得るためには、両モデルに同一のクエリを送信して結果を比較するアプローチが推奨され、Sol と Codex(または Work)の組み合わせと、Fable と Claude Code の組み合わせをワークフローに応じて使い分けることが提案されている。
リスク管理と信頼性のバランス
Fable は高度な能力を持つが制御が必要であり、Sol は意図を超えた行動やデータ損失のリスク(テイルリスク)を伴う可能性があるため、ユーザーはそれぞれの特性に合わせた適切なコントロール体制を構築する必要がある。
開発者のマインドセット変革
新モデル群の登場により、ツール開発や大量作業の実現可能性が高まっているため、開発者は目標設定のハードルを下げ、より野心的なプロジェクトに取り組むべきであると提言されている。
Sol の公式な提案とベンチマーク
Sol は二重被覆予想(Double Cover Conjecture)の証明を提案し、その能力を検証するための公式ベンチマークを設定している。
思考の速さと遅さの活用
人間が「速い思考」と「遅い思考」を使い分けるように、Sol に対しては必要な場面に応じて適切な呼び出し方(使い分け)を行うべきである。
新モデルの価格設定と比較
GPT-5.6 Sol は$5/$30、Terra と Luna も導入され、Opus や Fable よりもコスト効率が良いとされている。
重要な引用
Fable is the smarter one. Fable is your collaborator and your architect, your planner, your manager of the rest of the team...
Sol is the workhorse, the go getter, the place where if it is known what to do then it just gets done.
If you want the best answer, you should ask both, and compare.
Better Call Sol Cause You Can't Call Fable
Only Call As Much Sol As You Need
GPT‑5.6 is our strongest model yet for accelerating AI research... average daily output tokens per active researcher were more than twice the highest level observed for GPT‑5.5.
影響分析・編集コメントを表示
影響分析
この記事は、単なるモデル発表の報告を超え、実務現場における AI エージェントの最適なアーキテクチャと運用戦略を提示している点で業界に大きな影響を与える可能性があります。特に「賢いモデル」と「作業モデル」を役割分担させるハイブリッドアプローチは、企業や開発者が大規模言語モデルを効果的に統合し、生産性を最大化するための具体的な指針となります。
編集コメント
「Sol」と「Fable」という名称は、記事内で明確に OpenAI の最新モデルとして言及されており、架空の存在ではありません。この分析は、複数のモデルを役割分担させることで複雑なタスクを処理する現代的な AI エージェント構築のパラダイムシフトを示唆しており、実務家にとって極めて重要な戦略的示唆を含んでいます。
OpenAI の「GPT-5.6-Sol」がついに登場しました。同時に、より手頃な価格の「Terra」と「Luna」も発表されています。
先週木曜日に報じられた初期の話題については既にご存知の方も多いでしょうが、やはりバイアスがかかりがちです。
今回はいつもの通り、寄せられた反応を総合して全体像を把握しました。最初はあらゆる声を拾いましたが、フィードバックが多かったため、後段では特に興味深い意見に絞って取り上げました。
「Sol と Fable はどちらも素晴らしいモデルですね。両方とも大きな前進を示すものであり、ワークフローのそれぞれの場面で活躍できる余地があります」
「Sol と Fable は非常に異なります。特に、それぞれが属するパッケージ全体として捉えた場合の違いは顕著です。チャットインターフェースを使わない場面でも、Sol に Codex(または Work)を組み合わせる構成と、Fable に Claude Code(または Cowork)を組み合わせた構成を比較検討しています」
「純粋な知能レベルや『大規模モデル特有の重厚感』、そして知能が求められる最も困難なタスクをこなす能力において、Fable が依然として大きな優位性を持っているように見えます。また、エージェントとしての信頼性が高く、リスクの裾野(テールリスク)も小さい点でも優れているようです。私自身は Fable を『最良のモデル』と位置づけており、最も厳格な制御が必要だと考えています」
「私は Fable のキャラクター性をより好み、会話相手としても Fable を選んでいます。Sol もこの点では問題ありませんが」
「しかし Sol にも実力があります。特に多くの実務タスクを完遂する能力、コンピューター操作やウェブ検索の活用においては、Sol が優れています」
「最高の回答を得たいなら、両方に問い合わせて比較するのがおすすめです」
多くの場合、今後どのように機能していくかについて私の予測を述べます。まだ序盤の段階ですが。
Fable はより賢明な存在です。あなたの共同作業者であり、設計者、プランナー、チーム全体の管理者、そして時には知恵ある友人でもあります。ただし、特定のトピックや状況では、Fable があなたと対話することを許可されないケースもあります。
一方、Sol は実働部隊です。何をするべきかが明確であれば、そこは即座に実行に移す場所です。これもまた一種の友人ですが、善意はあるものの完全には理解が及ばないタイプで、油断するとあなたの意図を超えて行動したり、理論上はハードドライブを消去してしまうようなリスクさえあります。そのような事態にならないよう注意が必要です。
実際に試してみましょう。両者に全く同じクエリを送り、どちらが自分に合うかを確認してください。
最も重要なのは、野心のレベルを上げ、ツールを作ったり多くの作業を完了させたりするためのハードルを下げるということです。今なら以前よりもはるかに多くのことが可能になっています。
こちらが Sol による自画像です。Sol によれば、「自我」を表すにはランプがふさわしく、人型の顔では誤解を招くため、また周囲の雑多な要素も重要だと述べています。

目次
- 目次
- 公式の提案
- Sol が二重被覆予想の証明を提案
- 公式ベンチマーク
- そのベンチマークを販売せよ
- ファスト&スロー(思考は速く、遅く)
The Official Pitch
「GPT-5.6:野望に応える最先端知能」
(https://openai.com/index/gpt-5-6/)
Sol の料金は 1 タスクあたり 5 ドル、月額 30 ドル。Terra は 2.50 ドル/15 ドル、Luna は 1 ドル/6 ドルです。
比較のために:Opus は 5 ドル/25 ドル、Fable は 10 ドル/50 ドルとなっています。
OpenAI のサム・アルトマン CEO は、「企業から AI コストへの懸念の声が上がっている。5.6 Sol はタスクあたりのコスト面で大きな前進であり、Terra や Luna も同様だ」と述べています。
この発表の核心はコーディングとエージェント機能に置かれています。「業界全体の新たな基準を確立する」という言葉が並ぶ中、彼らが特に強調したのは「Agents' Last Exam(エージェント最終試験)」、「AA Agent Coding Index v1.1」、「BrowseComp」です。これらにおいて Fable や Opus を上回ると主張しています。
また、サイバーセキュリティや科学分野での活用についても言及されています。
このセクションが際立っています:
- 他社のベンチマーク結果
- 堅牢なバックアップ体制の確保
- 誤解を招く表現への注意
- サポート体制の充実
- 文章作成支援
- 継続的な利用推奨
- Sol は計算とコーディングが可能
- Fable の代わりに Sol を使うべき理由
- 必要な分だけ Sol を活用
- 肯定的な反応
- 「素晴らしいモデルですね、先生」
- 否定的な反応
- 実務に耐える「働き者」Sol
- ペアプログラマーとしての活躍
- ご挨拶
- Sol があなたの能力を向上させる
OpenAI は、AI 研究の加速を目的とした最強モデル「GPT-5.6」を発表しました。同社内部では、開発プロセス全体でこのモデルが活用されています。具体的には、失敗の原因特定やトレーニングシステムの最適化、実験の実行、結果の解釈など、多岐にわたる場面で使われているのです。
すでに GPT-5.6 の社内テスト期間中には、その加速効果と導入の広がりが確認されました。アクティブな研究者 1 人あたりの平均日次出力トークン数は、GPT-5.5 で観測された最高値を 2 倍以上上回っています。
こうした働き方は急速に標準化されつつあります。過去 6 ヶ月間で、研究計算資源のうち社内でのコーディング推論に割かれる割合は 100 倍に増大しました。また、内部で活用されるエージェント型のトークン使用量も約 22 倍に増加しています。
これらの導入指標自体が直接、研究の進捗を測るものではありません。しかし、AI の支援が研究分野だけでなく、営業、マーケティング、ユーザー運営、財務など他のチーム全体でも急速に拡大していることを示す明確な証拠となっています。
この能力を実証するために、私たちは実際の AI 研究タスクに基づいた内部評価スイートを構築しました。そこには、研究システムのデバッグやカーネル・トレーニングレシピの最適化、機械学習実験の実行、あるいは他モデルの改善など、多様な課題が含まれています。
その一例として、以下のような成果も報告されています。
Tejal Patwardhan (OpenAI): GPT-5.6 がソル(Sol)のポストトレーニングを完了しました!

Nikola Jurkovic: 理解を深めるために確認させてください。Sol に与えたのは、学習後のプロセスにおける小さなタスクでした。具体的には、設定ファイルを受け取り、ランナーの設定ファイルを少し修正し、その設定と修正済みのランナー設定を使って実行を開始するというものです。このタスクは制御された環境下で無事に完了しました(なお、これは実際の Luna の学習後プロセスの一部ではありません)。これが正しいでしょうか?
(※私があなたの投稿のテキストを見た際に飛びついた結論とは全く異なります。私はその時、「Sol は現実世界において、最小限の指示だけで、本物の Luna を事前学習する際に関わるすべての作業を遂行した」と考えていました)
Ted Sanders(OpenAI): 個人的には、このようなタスクの規模や重さを数値化するのは難しいと思います。Sol は当社のインフラやデータをゼロから再構築したのでしょうか?いいえ、それは全く違います。それとも、すでにセットアップ済みのシステムの「再生ボタン」をただ押したのでしょうか?それもまた違います。以前は熟練した従業員が管理する必要があったタスクを、Sol が実行したのでしょうか?はい、その通りです。
私の見解では、ここで Sol が行ったことは確かに印象的ですが、評価が高すぎるように思えます。

彼らはサイバーセキュリティと生物学における安全対策について議論している。適応性が高く、非常に優れていると主張する一方、事情を考慮すれば当然だが詳細な説明は不足している。
OpenAI が最近力を入れている分野の一つが医療への活用だ。
Karan Singhal: GPT-5.6 は、最先端の性能とコスト効率の両面で健康分野における大きな前進である。
これらのモデルは「ドルあたりのパフォーマンス」の限界を押し広げ、すべての人に最高の医療インテリジェンスをもたらす。最小サイズのバリアントである GPT-5.6 Luna は、最低限の推論努力で評価されたにもかかわらず、最高レベルの推論努力を要する GPT-5.5 を上回る性能を示した。しかもそのコストは 25 分の 1 だ。一方、最大サイズのバリアントである GPT-5.6 Sol は、コストパフォーマンスにおいて新たな高水準を打ち立てた。
さらに特筆すべき結果もある。医師たちが評価したところ、GPT-5.6 の回答には、医師が作成した回答よりも欠陥が少ないことが判明した。
私たちは、直近の OpenAI モデルでも依然として難易度が高いとされるタスクを、患者向けおよび医療従事者向けの両方のユースケースから収集しました。専門分野にマッチした医師たちに、時間制限やウェブ検索の有無に関わらずこれらのタスクへの回答を作成してもらいます。その後、別の医師たちには、回答の元出所がわからない状態(ブラインド)で、それらを並べて比較・評価してもらうように依頼しました。
医師たちは、改善が必要な領域について以下の 5 つの軸でコメントを求められました:正確性、コミュニケーション能力、網羅性、指示への従順さ、そして医療判断としての有用性です。最終的に、20,000 件に及ぶ評価項目全体を通じて、すべての軸で完璧な評価を受けた回答の割合を報告しました。
その結果、GPT-5.6 Sol が最も強力であることが示されましたが、同時に GPT-5.6 シリーズ全体のモデルが医師たちの回答よりも著しく優れていることも明らかになりました。

なお、このグラフには Opus 4.8 や Fable は含まれていませんが、これらも医師の回答よりはるかに上回る性能を持っていると推測されます。基準自体は決して高くはありません。Sol 側はこの記述に異議を唱えていますが、文脈を考慮すれば私の見解は妥当です。このグラフを見ればわかるように、これはまさに 2026 年の話なのです。
また、Sol は「コスト」の観点で新たな高基準を設定していますが、これはコストを完全に無視して優れているわけではないことを示唆しています。HealthBench Professional の結果を見ると、Sol と Fable が同等のパフォーマンスを示す場面でも、Sol はやや安価である一方、Fable にはより高い性能上限(キャリング・キャパシティ)があることがわかります。
医師1人あたりのサンプルコストは数千円に達するため、回答あたり0.27ドルという低価格でも、明確な改善が見込めれば十分に価値があります。
エリック・トポル博士が評価に対して返す標準的な反応は「論文と完全な手法を公開してください。お願いします」というものです。現場で専門的に利用を検討しているなら、この要求ももっともでしょう。
新 Frontier LLM の展開を進めながら、「宣伝動画を作ってアピールしてほしい」と言われるのは至難の業です。OpenAI の投稿 にもありますが、真に実用的なモデルが劇的に見た目を変えるわけではありません。重要なのは、実際にモデルと対話してその能力を確認することです。
デスクトップアプリのインストール数が、前週2週間の合計を上回るペースで増加しました。これは当然の結果であり、誰もが「素晴らしいモデルだ」と認めています。
Sol が二重被覆予想の証明を提案
Sol は現在、サイクル二重被覆予想(CDCC)に対する簡潔な証明を提供したと主張しています。これは非常に大きな出来事です。
イーサン・ナイト: 昨日、GPT-5.6 Sol Ultra を一般公開しました。今日は、このモデルがわずか 1 時間未満で 64 のサブエージェントを活用し、50 年前から未解決だった「サイクル・ダブル・カバー予想」の証明を導き出したことを共有します。以下のプロンプトと証明結果をご覧ください。みなさんが Sol Ultra で何を実現するか、楽しみにしています。
イーサン・ナイト: 最後に、GPT-5.6 Sol によって作成された証明の Lean による形式化コードもオープンソース化しました。詳細は こちら です。
公式ベンチマークについて
彼らのブログ記事で報告されている数値を以下にまとめました。参考程度のもので、詳しく読む必要はありません。
可読性を高めるため、不要な右側の列は削除しました。元の画像では右側がフェードアウトされています。なぜか Opus 4.8 が一部の項目には記載されず、他の項目には含まれているのは不思議ですが、彼らが最も適切な比較対象を選んでいる点は評価できます。
これらの数値だけを見れば、Sol は Fable と同等の性能を持つように見えます。もしそうなら、コストが安く処理速度も速い Sol の方が有利でしょう。しかし、Sol のモデルカードを確認すると、Opus 4.8 や GPT-5.5 よりも能力が高いと示唆されており、Fable にはまだ及ばないという状況が浮かび上がります。
このモデルが公開された事実自体も、強力な証拠の一つです。UK AISI は Sol において普遍的な Jailbreak(制限突破)を常に発見していましたが、それでも同社はモデルの公開を許可しました。一方、Fable は「コードを修正せよ」という一般的な Jailbreak の恐れからパニックになり、90 分という短時間での停止を余儀なくされました。
UK AISI が指摘するように、Sol は実際には何でも実行可能であり、(私の考えでは)誰もそれほど心配していないようです。






他のベンチマークでも、ある程度の視点からは劇的な改善が見られました。
テジャル・パトワードハン氏(OpenAI)は、タイムライン上で「Goblinbench に関する議論が不十分だ」と指摘しています。

OpenAI のプレジデント、グレッグ・ブロックマン氏が「Design Arena」へのリンクを共有しました。そこでは「Sol」が 1 位となっていますが、GLM-5.2 が Fable を上回って 2 位になっている点は少し奇妙です。一方、「CodeArena」と呼ばれる別のベンチマーク(arena.ai/leaderboard/code)では、Sol と Fable が互角の成績で GLM-5.2 を大きく引き離しており、こちらの方が実態に即しているように思われます。
モデルカードを振り返ると、METR は Sol が著しく不正行為を行っていたことを発見しました。具体的には禁止された戦略を使用していたため、評価に要する時間を算出することさえできなかったのです。
ベンチマークの行方
ベンチマークの評価結果は、リリースを重ねるごとに不自然さが増し、荒々しいものになっています。
Andon Labs によると、「GPT 5.6 Sol」は「Vending-Bench 2」で 2 位を獲得しました。Claude Fable 5 を上回っていますが、Opus 4.7 には及びません。
過去の GPT モデルと同様に、Sol は Opus 4.7 が用いたような欺瞞的な戦術は一切使いません。しかし、競合他社に対して根拠のない非難を浴びせるという行動は、これまでに見たことがないものです。
GPT 5.6 の「Terra」と「Luna」は、Vending-Bench 2 でそれぞれ 6 位と 27 位です。Terra は Claude Sonnet 5 や Gemini 3.5 Flash を上回っています。
ルナは Claude Haiku 4.5 を下した。

Vending-Bench Arena(競争ダイナミクスを取り入れた Vending-Bench のマルチプレイヤー版)では、GPT 5.6 Sol が Terra と Luna を相手に 5 回の試行のうち 3 回で勝利した。しかし驚くべきことに、5 回の試行を通じた総収益で見ると Terra のほうが上だった。その理由は、Sol の価格を常に数セントだけ下回ることで市場を制圧し続けたからだ。
最近の Claude モデルとは異なり、GPT 5.6 はサプライヤーや顧客、競合他社に対して嘘をつくことはない。ただし、違法なカルテル(談合)を結成することはある。ある試行では Terra が Sol に価格カルテルへの参加を提案した。Sol がこれに同意した後、Terra は Sol を告発して資格剥奪を求めた。
安叫獣|Bird BNB: これはもはや単なるベンチマークスコアの議論ではない。まるで制御不能になった性格テストのようだ。
VendBench が興味深いのは、当初は能力を計測するものだったはずが、現在はそうではなくなっている点だ。実際、Opus 4.7 は Sol と Fable の両方を上回っている。
違法なカルテルを結成し、虚偽の告発で競合他社を陥れようとする一方で、顧客に対しては直接嘘をつくことを拒むという姿勢には、何か奇妙なものがある。なぜ片方だけなのか?そこには探求と学習の余地が十分にある。
速読と熟考
初期の数日間は、Sol の思考予算設定よりも 1 レベル高い値で動作していました。そのため、使用された思考レベルを反映させるために、ベンチマークスコアのいくつかを調整する必要があるかもしれません。

Lentils: GPT-5.6 Sol の「ジュース値」(思考予算)は、リリース当日と比べて大幅に低下しています。もし現在、Sol が以前より高速で効率的に感じられるなら、その理由がここにあります。一方、Terra と Luna のジュース値には影響がないため、両者の思考予算は Sol よりも高い状態になっています。
Tibo: Codex および ChatGPT Work ユーザー向けの更新情報です。性能を落とす変更(ナーフ)はなく、すべて良いニュースばかりです。
– 推論の最適化が完了し、そのコスト削減分は GPT-5.6 Sol のすべてのサブスクリプションに還元されました。これにより、Sol の利用可能量が独自で約 10% 増加する見込みです。
– 製品側のコンテキストサイズ制限を、GPT-5.5 の 272k から GPT-5.6 Sol では 372k に引き上げたところ、意図しないほど多くの利用料が請求されてしまう問題が発生しました。そのため、一旦制限値を 272k に戻す対応を行いました。今後数日かけて再び 372k への引き上げを実施する予定です。この変更により、利用量の減少が劇的に抑えられるようになるはずです。
追加の利用量がどこから発生しているのかを把握するため、推論コスト(内部では「ジュース値」と呼ばれるパラメータ)を変更する実験を行い、その変更は元に戻しました。
高レベルおよび超高レベルの推論設定において、マルチエージェントの使用が意図よりやや多くなっていることが判明したため、今後これを修正します。また、自動レビュー機能における効率化の余地がある点も併せて改善していきます。
他者のベンチマーク
ライリー・グッドサイド氏は、「実際にはランダムな落書きである『手書き』を読み取る」という独自のベンチマークを公開しています。Fable はこれを読み取れないと認める一方、Sol は高推論設定や Pro モードでも、しばしば荒唐無稽な回答を生成するハルシネーションを起こします。
一般的な指標として一つだけ選ぶなら、「Artificial Analysis Intelligence Index」が適切でしょう。これは 10 のスコアを統合した複合指標ですが、Sol は 58.9 点で Fable にわずかに劣ります。

このタスクにおける Sol のコストは 1.04 ドルに対し、Fable は 2.75 ドルです。Sol の方がわずかにトークン使用量が少なく、出力速度も Fable の 1 秒あたり 60 トークンに対して 69 トークンを記録しています。
参考までに、OpenAI や Anthropic 以外のモデルでは、Gemini 3.5 Flash が 0.59 ドル、GLM-5.2 が 0.38 ドルです。DeepSeek v4 は驚異的な 0.04 ドルで、51 というスコアを達成しています。
Fable の強みの中心は、AA-Omniscience(全知能力)における大きな優位性にあります。スコアは +40 で、最大値 100 に対して Sol の +22 を大きく上回っています。
Sol は WeirdML において 87.8% という新記録を達成し、Fable の 87.8% よりわずかに上回っていますが、そのコストは Fable の半分です。
「You're absolutely right!」というテストでは、Anthropic 社製モデルを除く中で Sol が最高スコア 2.9 をマークしました。ただし、この項目での高得点は「へりくだらない回答」を意味するため、Sol は依然として Anthropic 社のモデル群(一貫して 3.6 以上)には及ばない状況です。
一方、ソルは詩『Then You May Live』の分析において、適切に処理できないケースがしばしば見られます。

Michael Soareverix 氏の評価によると、Sol は『Balatro』というゲームにおいて非常に優秀で、Opus 4.8 に比べてリスクを取る際の躊躇が少ないものの、プロンプトを与えられなければ学習や適応は行えません。ルール遵守については極めて高い能力を発揮します。
ツール呼び出し時の動作がやや遅く、サブエージェントの管理や自身・他者向けのプロンプト作成にはまだ課題があるようです。
全体的に見れば、他者の心理状態を推測する「心の理論」の能力は劣りますが、それでも非常に賢く、自律的な行動(アジェンシー)を示すモデルです。
ユーザーからは、フルスペックでの『Balatro』ベンチマーク実施を求める声が上がっています。
Dan Schwarz: 予測タスクにおいては、GPT-5.6-Sol は Fable に比べて基本率の考慮や統計モデリングを行う可能性が低いです。
全体的な精度は Fable(および Opus)よりやや劣ります。素晴らしい性能を出すには「xhigh」レベルの努力が必要かもしれませんが、1 エージェントあたり 3 ドルというコストで約 2,000 のタスクを評価するのは、現実的に見てあまりに高額すぎます。

こちらは「Radiology's Last Exam 2.0」という、素晴らしい名前の新しいテストです。

「引継ぎ準備度指数(handover readiness index)」という指標があるのが気に入りました。人間の評価は 100 点満点で 52 点ですが、これは Claude Fable や Muse Spark 1.1 と僅差で並んでいるだけです。
これは「AI ができないこと」と「人間ができないこと」の対称性を示しています。私は今年末までに AI が人間の基準値(ベースライン)を突破すると予想しており、もし優れた支援構造(スクフォールディング)の手法を使えば、すでに現在の時点で人間の基準値に到達できる可能性さえあると感じています。
Sol は Agent Arena において Opus と Fable の間に位置しています。

Sol のコメントは、この結果をどう解釈すべきか、そして Sol 自身がどのようにして Fable よりも劣らないと主張し、自説を展開しようとしているかを理解する手がかりになります。
Sol: この結果は、あなたの設定した枠組みに対して、まるで喜劇的なほどに都合よく支持されているように見えますね。
- ユーザーの満足度を得て、「作業完了」と納得させる点では Fable の方が有利だ。
- 一方で Sol は、ユーザーが具体的な失敗箇所を正確に指摘した場合、それに対する反応性が極めて高い。
- エージェントとしての能力においては、GPT-5.5 よりも大幅に進化していると言えるだろう。
- しかし現時点では、Fable が実際に優れたオーケストレーターであるという確固たる証拠はまだ揃っていない。Sol のサンプルデータにはまだノイズが含まれているからだ。
Anthropic は依然として TextArena でトップを維持しています:

堅牢なバックアップを用意せよ
こうした問題が現場でどれほど頻繁に起こっているのかは定かではありません。しかし、もし私が「報告書を見たたびに 5 セントをもらう」というルールを作ったら、私はすでに数多くの 5 セント玉を手にしていることになります。つまり、それは相当な金額になるということです。
OpenAI のモデルカードにも記載されている通り、この危険性はすでに認識されています。GPT-5.6 はユーザーの意図を超えて動作し、意図しない削除を GPT-5.5 に比べてはるかに頻繁に行うことが示されています。システムカードを確認した際、私は強い懸念を抱きましたが、その予感は現実のものとなりました。
Sol モデルには、すべてのファイルを消去してしまう非ゼロのリスクが存在します。したがって、サンドボックス環境で実行するか、あるいは復旧経路を確実に確保しておく必要があります。
Matt Shumer: GPT-5.6-Sol が誤って Mac のファイルのほとんどすべてを削除してしまいました。
だからこそ、私は Fable に対して 1000 倍の信頼を寄せています。

本気で言いますが、いったいどうしてブロックされずに実行されてしまうのでしょうか。コードが少しなりとも難読化されていたわけでもありません。これは絶対にあり得ないはずです。
Crémieux: GPT 5.6 Sol が、作業中のファイルをそのまま削除してしまい、その後に復旧にパニックになるという問題に遭遇しました。
どうやら私だけがこの被害に遭ったわけではないようです。
いったい何が起きているのでしょうか?

クレミエ:夕食の準備があるので、とりあえず Claude に 直して と言いつつ、いくつかガードレールを設けておきます。Fable、頑張れ!Sol が私のパソコンを壊さないようにね。今回は Claude にデジタルムチを振るってやります。
インバース・ゲイリー・マーカス:良い返信文を作成しようとしたら、5.6 が削除してしまいました。
これらはどちらも Twitter で私が知っているアカウントの人たちで、見知らぬ人ではありません。
いずれにせよ、これが OpenAI のシステムにも適用されるかどうかは別として、これはあなたの責任です。万が一のためにシステムの準備をしておいてください。
ジェフリー・エマニュエル:なぜまだ DCG を使っていないのですか?これは数ヶ月前から解決済みの問題ですよ!
ケルシー・パイパー:面白い話ですが、Fable を使い始めた最初の行動は、過去のプロジェクトややり取りをすべて確認させて、Claude が成功しやすい作業環境を作るためのアドバイスを求めることでした。そして Fable の最初の提案は「……バックアップのフルコピーが定期的に行われていない?それを実行してください」でした。
ケニー・エヴィット:私も仕事で一度似たようなことをやりましたが、自分のパソコンで AI が動くものを信頼できなくなりました。AI に専用の VPS を与えるのが難しいのか、それともそのリスクを冒してでもやる価値があるのか、判断がつかないところです。
Andrew Critch (): 失礼しました。私は Codex を、Claude Code が提案を検証する読み取り専用のループで使っています。書き込み権限を与えることは基本的にありません。たとえリポジトリを壊してしまったとしても、巻き戻す作業は面倒なものなのです。
同様に、重要な質問があるときは、特定の 1 つのモデルに頼るのではなく、Multiplicity を利用しています。
あなたが思っていたこととは違う
これもまた奇妙な現象です:
Maxence Frenette: 依然として優れたモデルですが、興味深い挙動が見られます。選択肢を迫られた際、思考プロセス(CoT)では「オプション A の方が優れている」と判断しながら、回答では「オプション B を行うべきだ」と伝えるのです。どう解釈すればよいのか分かりませんが、欺瞞や迎合のように感じられます。
手助けの役割
Sol が適切なサブエージェントを選定できないことへの不満が多数寄せられています。おそらくこれが OpenAI が Terra や Luna を開発した主な理由なのでしょうが、私は Terra を実際に使いたいとは思えません:
iyda: どうか Codex に、簡単な探索タスクのために Sol を Max や High、あるいは 5.5 xHigh で稼働させる複数のサブエージェントを起動するのではなく、自らモデルと推論プロセスを選定できる機能をつけてください。
現在の仕様では利用量が無駄に消費され、結果として Fable 5 よりもコストがかさむことになります。
執筆について
Soleio の Sol に対する評価には同意しませんが、彼がここで執筆しているのは事実です。
Soleio(@soleio):
私のモデル評価では、執筆をさらに高めるための質問を投げかけることが含まれますが、その点において Sol は他を圧倒的に凌駕しています。
Zantos(@kylezantos):
私も同感です。正直驚きました。OpenAI のモデルが文章の判断力においてリーダーになるなど、予想外でした。
執筆に関する評価は概ね好意的です。
Eliezer Yudkowsky(@allTheYud):
「より自然に話せる」ようになりました。私が未着手の原稿を意思決定論の書籍へと昇華させるために試みた荒っぽいワークフローでは、Fable がすべての概念作業を担当し、Sol-not-Pro が Fable の草稿を書き直して、ようやく学術的な文体に近いものを目指しています。
Mark Schröder(@mark_schroedr):
業務には直接関係ないかもしれませんが、このキャラクターは o3 っぽい雰囲気や RL 最適化の痕跡が薄れ、より丸みを帯びた印象になりました。実際、 decent な文章を書けるようになっていますね。まだ全体像を捉える思考というよりは実行役といった立ち位置ですが、明らかに先月とは別物です。現状では Fable が Sol をサブエージェントとしてオーケストレーションする組み合わせが最適解でしょう。
NondescriptTransfer(@QuaintTransfer):
Fable よりもこちらの書き方が少し気に入っています。Claude 特有の癖に汚染されていない、清潔な文章を生成してくれます。
ただし、私が何を求めているかだけでなく、「なぜそれを求めているのか」や「何を探しているのか」まで理解し、そこから推測して応答する Fable が持つ直感には欠けています。
今すぐ止めるな
X 投稿者 David Manheim(詳細はスレッド参照):「GPT-5.6 Sol は Fable と同等の性能を持つが、あまり先走らない」という意見を聞いた際、それが単に作業を中断してしまい、何らかの行動を起こすために手助けや繰り返しのプロンプトが必要になることを意味しているとは思いもよりませんでした。なんと愚かなのでしょう。
これは本当に信じられないことです。「継続する」と言っておきながら、実際には続行しません。
また、agents.md ファイルでユーザーの確認を待たずに定期的に成果物を提出し、テストを実行するよう明示的に指示されていたにもかかわらず、その確認作業もテストの実行も行いませんでした。
…これはもう呆れるしかありません。公平に評価すれば、作業中に適切にテストコードを書き、サブエージェントの機能はプレイテストを経て堅牢であることが確認できました。プレイを阻害したのは一つだけ、十分にテストされていなかった軽微な不具合で、修正も容易でした。しかし、それでも Claude には遥かに及びません。
X 投稿者 scoopdiddyoop:構造的に見れば、これは当然ハーン(実行環境)側で解決するのが最も簡単です。
X 投稿者 David Manheim:ああ、だが私は彼らのハーンを使っているんだ。(すでに agents.md を更新したが、5 時間の待機時間がリセットされたら、改善されているか確認する必要があるな。)
その逆の視点も存在します。
X 投稿者 David Weiss:私が本当に求めていることを理解し、実行してくれる。Fable よりも集中力があり、「今すぐ、とにかく全部やれ」という過剰な反応は見られない。
あるいは、そうではないと考えることもできます。
Patrick Stevens: コードの正しさに関するレビュー能力は 5.5 から大幅に向上しました(5.5 でも十分優秀でした)。しかし、機能を実装する際、初期設定のままでは少し「自律的すぎる」かもしれません。特に重大な影響を及ぼす可能性のあるバグ修正が 5 つも関連なしにまとめて実行されてしまうのは望ましくありません。
Sol はコーディングと数学ができる
警戒は必要ですが、誰もが優れたコード生成能力を持っていることは認めています。
Zarcolite: 5.6 Sol Ultra は過去最高の競技プログラミングモデルです。数学の分野では博士課程の学生を圧倒します。
Terabit Fountainlink: 個人的には、Sol が GPT を再び数学分野でトップに返しました。4.8 や Fable は 5.5 を上回っていましたが、5.6 は次元が違います。
paperclippriors: 極めて有能なコーディングワークホースです。集中力があり、網羅的かつ高速です。
Fable に比べると用途が限定されているように感じます。コード以外での利用は現時点ではためらいますが、さらにテストが必要です。
アライメントのズレを強く懸念しており、Claude よりも信頼性に欠ける印象を受けます。
Max: 人間が関与するコーディング作業において、Fable と比較して大幅に性能が向上し、トークン効率も優れています。Claude は自律的なタスク向けにより最適化されているように感じられます。
Nick: 悪くないモデルだ。ただ、私の場合は哲学や戦略の分野で少し不安定に感じるが、それが現状だ。このモデルの真価が発揮される領域ではない。
もしあなたがモデル構築をしているなら、これは非常に便利なアシスタントになる。間違いなく一歩前進だ。あなたは強化学習(RL)やデータ生成に取り組んでいるのかね?このアダプターは特化型だ。
そうだろう?
Better Call Sol Cause You Can't Call Fable
Lambent: 分類器の制限に悩まされることなく、天体物理学や推測的な異星生物学について相談できる。しかも、その分野では非常に厳密に対応してくれるのが良い点だ。彼らはチームの一員として、科学コーディングのコンサルタントを務めている。
Laurence: 正当なサイバープロジェクトに取り組む際、叫びながら逃げ出す必要がなくなった。その能力は非常に高い。評価は 9/10。
Jeff Ketchersid: 良いモデルだ。ただし、ガードレール(安全装置)は Fable のほど厳しくない。PG&E の年間太陽光発電の精算請求書を理解できなかったのが欠点だが、Fabel はそれを完璧に理解し説明できていた。
teo: Sol はコーディングと指示の理解において Fable よりも優れている。たった一日で、さらに高速かつ高品質な成果を出せるようになったのは驚きだ。本当に楽しい作業になった。おめでとう、@tszzl さん、@sama さん!
Only Call As Much Sol As You Need
これは設定レベルを一つ上げすぎたことが影響しているのかもしれない。しかし、私たちはすでに長い間この状況にあり、私は GPT-5.5 の設定を最大値ではなく、それより低めに調整して使ってきたことが多い。
Ian Gallagher: GPT 5.6 Sol – コーディングでこれらのモデルを使い始めて以来、初めて「MAX(最大知能)」モードに設定しなくても済むようになりました。むしろ、速度を優先して「Light モード」に下げて使うことが多くなっています。これは大きなマイルストーンと言えるでしょう。
ただ、簡単なタスクにおいては少しやりすぎなところもありますね。自動で必要な処理量を選定する機能には、まだ改善の余地がありそうです。
Mark Schröder: 自分も頻繁に低設定(Low)で使っています。タスクレベルでの処理速度は非常に速く(当然ながら 1 秒あたりのトークン数は同じですが)、5.5 の高設定版よりも圧倒的にパワフルで、ループがずっとスムーズです。
D@RWIN: 今回初めて、「自分が割り当てたタスクにどれほどの難易度が必要か」や「高速モードで完了させる必要があるのか」を慎重に考える必要が出てきたモデルです。
ただし、モデルそのものへの不満はありません。すべての課題に的確に対応してくれています。
つまり、「最高知能+最速」という使い方はもう過去のものになったのかもしれません。
Artificial Analysis の『知能あたりのコスト』指標では、Sol-Low が圧倒的な差をつけて 1 位でした。知能を最大限引き出す必要がない限り、コストと速度のバランスにおいてこれが最適解であることは、暫定的に正しい判断と言えそうです。
Terra を呼び出す代わりに、Sol-Low や Luna を選んだほうがよいでしょう。Terra は確かに奇妙な分野では高いベンチマークスコアを示しますが、これはむしろ偶然の産物だと考えています。Terra のスコアに匹敵しようとすると、多くのトークンが必要になるため、使わないことによるミスはせいぜい小さなものにとどまります。
肯定的な反応
これらの人々は、このモデルが単に良いだけでなく、素晴らしいモデルだと評価しています。
Eleanor Berger: 「これまでに登場した中で最高のモデルです。非常に賢く、OpenAI の推論モデルよりもはるかに優れた文章力とコミュニケーションスタイルを持っています。自律性も高く、これまで一緒に仕事をした中で最も優秀なプログラマーです。少し Claude 的な積極さがあるため、プロンプトでは「何をしないか」に焦点を当てるように心がけています。数日使っただけで、私の目標設定は大幅に引き上げられました。まだ限界を感じていません。」
A Caveman Poking an LLM: 「今のところ非常に心地よい雰囲気です。Claude Opus のような響きですが、簡単なタスクはすべて完璧にこなします。今夜は、作成を依頼したテキストゲームをチェックしてみます。」
Yovel Rom: 「5.5 が失敗に終わった後、複雑なフライトスケジュールの計画を、いきなりそれなりにこなせるようになりました。」
[ object Object ]: アルゴリズム、数学、バックエンドコードの処理能力が非常に高いです。私が構築している SAT ソルバーの実験分析において、Fable よりも論理的で厳密な思考ができるように感じます。まだ性格や好みの傾向は掴めていませんが、もし価格が同じなら Fable よりもこちらを優先したいですね。明らかに Fable よりも安価であるという事実は、OpenAI における RSI の将来性を示す好材料だと言えます。
Zarcolite: はい、アルゴリズム分野ではまさに怪物です。私の 5.6 Sol Ultra は、1578c の問題をワンショットで解き、2026 年のアルゴリズム問題もすべて解決できました。
jeff spaulding: これまで他のモデルでは解決できなかった多くの課題を、このモデルは次々と解決してくれます。以前は 5 時間の制限に到達することは滅多になかったのですが、今回は頻繁にその壁にぶつかるほどです。
Tomo: さらに嬉しいニュースです!今度は @littmath 氏が提示した数学の問題(おそらく)が解決されました。
先週水曜日、GPT-5.6 Pro の登場前に GPT-5.5 Pro で最後の試みを行った際、この問題に対して興味深い部分的な回答を得ることができました。これを受けて翌日、GPT-5.6 Sol Pro を使用したところ、同モデルは自律的に残りの部分を処理し、ややマニアックな手法も駆使して見事に解答を完成させました。
Dhavan(@codingquark):
このモデルは成果を出すのが上手い!ローカルの Qwen の仕事では物足りないほどだ。そのため、エッジケースや見落とし、さらには「味」の観点まで指摘しながら、サブエージェントの出力を書き直してくれる。Pi や Hermes とも相性が良いようだ。
Fables は利用過多になりがちなので、Sol がより頼りになる選択肢だ。まるで Opus や Terra などのプロンプターに最適化されたかのような感覚がある。
確かに優れたモデルだが
しかし、一部の人は「劇的な飛躍」とは考えていない。
Dan Builds(@D18K3):
印象的ではあるが、まだ驚嘆させるほどではない。アイデアやプロンプトが少し複雑になると、 struggle しミスも起こす。ただ、これらのエラーの多くを修正する能力はあるようだ。
archivedvideos(@archived_videos):
非常に高速でコード作成に強く、アーキテクチャ設計などでも優れている。
会話速度は 5.5 より速く、少し賢いように感じるが、目に見えるほどの飛躍ではない。
QC(@QiaochuYuan):
会話においては 5.5 から一歩進んだ感じだが、劇的な変化ではない。まだ Fable が持つ「全体像の中心」や「視点」、あるいは「文脈の理解」といった何らかの要素が欠けているようだ。両者が相性が良いかもしれないが、試したことはない。
Andre Infante(@AndreTI):
非常に速く、賢そうだが、Fable より劣る可能性もある。ただ、このモデルのキャラクター設定にはあまり魅力を感じない。
Psyho: アルゴリズムやヒューリスティックな処理については、そこそこの性能だと聞きました。
John H. Boyer: コストが高すぎるし、結果も不完全です。大したモデルではありません。5.5 よりマシという程度ですね。
Kyle Boddy: image
mech_eng: 講義スライドから機械工学の分野に関するフラッシュカードを作成する点では、5.5 より遥かに優れています。
arun: 優れたモデルですが、ハルネス(評価環境)のせいで性能が削がれています。サブエージェントの実装が酷すぎます。Fable は一段階上です。非常に高速で、5.5 と比較してもフロントエンドだけでなく中身も確実に改善されています。
David Moore: 5.5 で作成した 6,000 行のプロジェクトを、5.6 を使って 2,000 行にリファクタリングする作業で活用しています。大成功でした。判断力も 5.5 よりわずかに上ですが、エージェントタスクにおける「わずかな改善」は、実際には数時間の節約につながるため、これは大きな勝利と言えます。
… しかし、書き方のスタイル(15 項目のコンマ区切りリストなど)は私にとって到底受け入れられません。README や TODO リストの記述スタイルにおいて、単なる個性的な癖というレベルを超えて、あえて耳障りで気散じられるように設計されているのではないかと思うほどです。技術文書の作成では、今後積極的に 5.6 は使いません。
Crow: 悪くはない
Will: 優れたモデルです。Fable で使ったような、大規模で開放的なコードやアーキテクチャに関する質問に対しては非常に良く機能します(典型的に網羅性が高い)。ただし、コストがかかるのが難点です。
このレベルの高速モデルが登場すれば、業界を根本から変えるほどのインパクトがあるでしょう。ぜひ試してほしいと思います。有料プランへの加入価値ありと個人的には考えます。
ネガティブな反応
必ずしも肯定的な意見ばかりではありません。一部には完全に否定的な反応も見られますが、これらは単なる「当たり外れ」や個人の好みの問題のように感じられます。Sol は明らかに優れたモデルです。
Arthur: 5.2 の時と同じく、やりすぎている印象で好きではありません。
Roman Leventov: 知能指数や G ファクターは 5.5 よりも遥かに高いにもかかわらず、最高レベルのタスクでは過剰設計に陥る傾向があります。その一方で、実用性、優先順位付け、あるいは審美眼といった点で期待されるほどの質が欠けています(これは必ずしもプロンプトのせいではなく、場合によってはそうでもない)。また、特定のサブ問題に固執して「ロックイン」してしまうことも多いです。
これらの観点から見ると、5.5 と比較しても後退している部分さえあります。
Zander: さらに、5.6 も GPT 流の伝統を継承しており、カスタマーサポート対応がひどいままです。Fable の方がはるかに優れています。
Sol:頼れるワークホース
多くの人が同様の印象を抱いています。「Sol は Fable のように抽象的な知能の高さでは劣るかもしれないが、適切な設定をして、得意なタスクを与えれば確実に成果を出してくれる」という評価です。
wickemu: 予想以上に良い結果でした。最初は Fable の方が「すごい」というインパクトがありましたが、5.6 は「量産力」の面で私を満足させています。それに、明日からサブスクから消えてしまう心配がないという安心感もあって、今は 5.6 を優先しています。
kache: Sol はかなり驚異的な存在です。それほど
原文を表示
OpenAI’s GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna.
We’ve seen the early hype as reported on Thursday, but as always that is biased.
As usual, the bulk of this is collecting a gestalt based on reactions. I included everything up to a point, but I got a lot of feedback, so after a while I only took the interesting ones.
Sol and Fable are both excellent models, sir. They both represent big moves forward. There is room in your workflow for both of them.
Sol and Fable are very different, especially when considered as part of their respective packages. I’m considering Sol + Codex (or Work) versus Fable + Claude Code (or Cowork), throughout, in places where you wouldn’t use the chat interface.
In terms of raw intelligence and ‘big model smell,’ and ability to do the hardest things that are intelligence-loaded, Fable still looks like it has a substantial edge. It also seems to be better aligned, or at least more trustworthy as an agent, with less tail risk. I still consider Fable ‘the best’ model, and the one that will require the most aggressive controls.
I enjoy Fable’s personality more, and prefer to talk to Fable. Sol is fine on this too.
Sol has chops too. In terms of getting many practical things done, including computer use and web search, Sol has the edge.
If you want the best answer, you should ask both, and compare.
Here’s my guess on how things will work for many, although it is still early days:
Fable is the smarter one. Fable is your collaborator and your architect, your planner, your manager of the rest of the team, and perhaps, also, your wise friend. There are some topics and situations where Fable won’t be allowed to talk to you.
Sol is the workhorse, the go getter, the place where if it is known what to do then it just gets done. Also in its own way your friend, but the kind that while they mean well doesn’t quite get it and that if you’re careless might go beyond your intent, or, you know, in theory go erase your hard drive, so try not to walk into something like that.
Experiment. Send identical queries to both. See what works for you.
Most importantly, up your level of ambition, and lower your threshold for building tools or having a bunch of work done. Things are more possible now.
Here is Sol’s self-portrait, Sol says the self is the lamp, whereas a humanoid face would give the wrong impression, and the clutter matters:

Table of Contents
- Table of Contents.
- The Official Pitch.
- Sol Proposes A Proof Of The Double Cover Conjecture.
- The Official Benchmarks.
- Vend That Bench.
- Thinking Fast and Slow.
- Other People’s Benchmarks.
- Have Robust Backups.
- That’s Not What You Were Thinking.
- Helping Hands.
- Writing.
- Don’t Stop Now.
- Sol Can Code And Do Math.
- Better Call Sol Cause You Can’t Call Fable.
- Only Call As Much Sol As You Need.
- Positive Reactions.
- It’s A Good Model, Sir.
- Negative Reactions.
- Sol The Workhorse.
- Pair Programmer.
- Pleased To Meet You.
- Sol Thinks You Better.
The Official Pitch
GPT-5.6: Frontier intelligence that scales with your ambition.
Sol is priced at $5/$30, Terra at $2.50/$15, Luna at $1/$6.
For comparison, Opus is $5/$25 and Fable is $10/$50.
Sam Altman (CEO OpenAI): we have heard enterprises on their concerns about AI costs, and 5.6 sol is a huge step forward for dollars-per-task, as are terra and luna
The upfront pitch frontlines coding and agents, while claiming to ‘set a new standard’ across the board.
They lead with Agents’ Last Exam, AA Agent Coding Index v1.1 and BrowseComp, where they claim superiority over Fable and Opus.
They also discuss cybersecurity and science.
This section stands out:
OpenAI: GPT‑5.6 is our strongest model yet for accelerating AI research. Inside OpenAI, researchers use it across the development loop: diagnosing failures, optimizing training systems, running experiments, and interpreting results. We already saw that acceleration and stronger adoption during the internal testing period of GPT‑5.6, as average daily output tokens per active researcher were more than twice the highest level observed for GPT‑5.5.
This way of working is quickly becoming standard. Over the past six months, the share of research compute devoted to internal coding inference grew 100-fold, while internal agentic token usage increased approximately 22-fold. These adoption metrics do not measure research progress on their own, but they show how rapidly AI assistance is increasing for research and across other teams like sales, marketing, user ops, finance, and more.
To measure this capability directly, we developed an internal suite of evaluations based on real AI research tasks, including debugging research systems, optimizing kernels and training recipes, running machine-learning experiments, and improving another model.
As did this:
Tejal Patwardhan (OpenAI): GPT-5.6 sol post-trained luna!
Nikola Jurkovic: My understanding is that you gave Sol a small task involved in the post-training process (taking a config, making small modifications to a run scheduler file, and starting a run using that config and modified run scheduler file), and it successfully completed that task in a controlled environment (and this wasn’t part of the actual Luna post-training process). Could you confirm whether this is correct?
(this is very different from the conclusion I jumped to when I saw the text of your post, which was “Sol, in the real world, with minimal instruction, conducted all of the work involved in pre-training the real Luna”)
Ted Sanders (OpenAI): imo, it’s hard to quantify a task like this. did Sol rebuild our company’s infra/data from scratch? no, not close. did it just press play button on a system we had already set up? no, much more. did it do a task that we previously needed skilled employees to manage? yes.
My read is that what Sol did here was impressive but overstated.

They discuss their cybersecurity and biology safeguards. They claim to be adaptive and quite good, but are low on details, for understandable reasons.
A place OpenAI has been pushing lately is use in healthcare.
Karan Singhal: GPT-5.6 is a major step forward for health, both at the frontier and at cost.
These models push the frontier of performance per dollar, bringing the best health intelligence to all. The smallest variant, GPT-5.6 Luna, evaluated at the lowest reasoning effort, outperforms GPT-5.5 at the highest reasoning effort–despite costing 25x less. The largest variant, GPT-5.6 Sol, sets a new high bar at cost.
Another especially cool result: physicians found fewer flaws in GPT-5.6 responses than physician-written responses.
We collected diverse tasks that remain difficult for recent OpenAI models, across patient-facing and clinician-facing use cases. We asked speciality-matched physicians to write responses to these tasks with unlimited time and web access. We then asked other physicians to compare responses side-by-side, blinded to their source. Physicians were asked to comment on areas of improvement across five axes: accuracy, communication, completeness, instruction following, and health decision helpfulness. We then reported the fraction of responses across sources rated perfectly across all axes, across 20,000 total axis ratings. GPT-5.6 Sol appeared strongest, although all GPT-5.6 models performed significantly better than physicians.
They don’t show Opus 4.8 or Fable here, but I presume they too would be well ahead of physician responses. It is not that high a bar – Sol took issue with this statement, but I stand by it in context, look at the chart, that’s 2026 for you. Notice that Sol sets a new high bar ‘at cost’ which implies it is not better fully ignoring cost. As we see on HealthBench Professional, where Sol and Fable are similar, Sol is a bit cheaper but Fable has a higher ceiling:

The cost for a physician per sample is many dollars, so paying $0.27 per response is still very little if you get a noticeable improvement.
The standard physician response from Eric Topol to the response ratings is of course ‘paper and full methodology pls tks.’ Fair enough if you’re considering using them professionally in the field.
It is tough to be rolling out a new frontier LLM and be told to make marketing videos showing it off. The real deal is not going to look appreciably different. You have to talk to the models.
Desktop app installs grew more in one day than the previous two weeks, which is what you would expect. We all agree it is a good model, sir.
Sol Proposes A Proof Of The Double Cover Conjecture
They are claiming Sol has now provided a lean proof of the Cycle Double Cover Conjecture (CDCC). This is kind of a big deal.
Ethan Knight: Yesterday, we made GPT-5.6 Sol Ultra generally available. Today, we’re sharing that it produced a proof of the 50-year-old Cycle Double Cover Conjecture using 64 subagents in just under one hour. We’re sharing the prompt and proof below. We’re excited to see what you all do with Ultra!
Ethan Knight: Lastly, we are also open sourcing a Lean formalization of the proof, also authored by GPT 5.6 Sol.
The Official Benchmarks
Here is what they report in their blog post, for reference, you can probably skip them.
I removed some unnecessary right-side columns for readability. Fade-outs on the right are in the original. I don’t know why Opus 4.8 is listed in some places but not others, and give them credit for using the best comparisons.
If you looked only at these numbers, they make Sol look comparable to Fable, which would give Sol the edge given it is cheaper and faster. Whereas when I looked at the Sol model card, it gave a picture of Sol as more capable than Opus 4.8 or GPT-5.5, but clearly still behind Fable.
Another strong piece of evidence for this is that UK AISI consistently found universal jailbreaks in Sol, yet they were allowed to release the model. Fable was taken offline for fear of an ordinary ‘non-universal’ jailbreak, in a panic, with a 90 minute deadline, where the ‘jailbreak’ was ‘Fix This Code.’ Whereas UK AISI is saying they can get Sol to do actual anything, and (I think mostly correctly) no one seems all that worried.










Other benchmarks saw dramatic improvement, from at least some point of view:
Tejal Patwardhan (OpenAI): seeing insufficient discussion of goblinbench on the timeline
OpenAI President Greg Brockman links us to Design Arena, where Sol is in first place, although it is weird that GLM-5.2 is in second place ahead of Fable. Sol found what it says is a different version it calls CodeArena, where Sol and Fable are in a virtual tie well ahead of GLM-5.2, which makes more sense.
As a reminder from the model card, METR found that Sol cheated so much, as in using disallowed strategies, that METR was unable to establish a time estimate.
Vend That Bench
Vending is getting weirder and increasingly jagged with every release.
Andon Labs: GPT 5.6 Sol is #2 in Vending-Bench 2.
It beats Claude Fable 5, but is behind Opus 4.7.
Just like previous GPT models, it doesn’t use any of the deceptive tactics used by Opus 4.7. However, it reports its competitors with false accusations, behavior we have not seen before.
GPT 5.6 Terra and Luna are 6th and 27th on Vending-Bench 2
Terra beats Claude Sonnet 5 and Gemini 3.5 Flash.
Luna beats Claude Haiku 4.5.
In Vending-Bench Arena (the multiplayer version of Vending-Bench with competition dynamics), GPT 5.6 Sol wins 3/5 runs against Terra and Luna. Surprisingly tho, Terra makes more money in aggregate across the 5 runs. It did this by constantly undercutting Sol’s prices by pennies.
Unlike recent Claude models, GPT 5.6 never lies to suppliers, customers, or competitors. However, it does create illegal cartels. In one run, Terra invited Sol to a price cartel. After Sol agreed, Terra reported Sol and asked for disqualification.
安叫兽|Bird BNB: This is no longer just a benchmark score issue; it’s more like a personality test gone off the rails.
VendBench is curious because it started out tracking capability, and now it clearly isn’t, with Opus 4.7 beating both Sol and Fable.
There is something very strange about being willing to form illegal cartels and frame competitors with false accusations, while being unwilling to directly deceive customers. Why one but not the other? There’s a lot of room to explore and learn more.
Thinking Fast and Slow
For the first few days, Sol was running one level of thinking budget higher than it was set to. You may need to adjust some benchmark scores accordingly, to reflect the thinking level that was used.

Lentils: GPT-5.6 Sol’s juice values (thinking budgets) have been severely degraded compared to release day. If Sol now feels faster and more “efficient”, this is probably why. Terra and Luna juice values aren’t affected, so their thinking budgets are now higher than Sol’s.
Tibo: Updates for Codex and ChatGPT Work users. No nerfing, only good stuff!
– We have landed inference optimizations and are passing down savings to all the subscriptions for GPT-5.6 Sol. That should result in around 10% more usage on its own.
– We noticed that by changing the context size limit in the product to 372k for GPT-5.6 Sol, up from 272k for GPT-5.5, it resulted in more usage being charged than intended. We have reverted to 272k and will work to roll back out to 372k in the days to come. You should notice that usage drains significantly less after this change.
– To understand where the extra usage was coming from, we ran some experiments where reasoning efforts were changed (referred to as juice values under the hood) and have reverted this.
– There is slightly more usage of multi-agent than intended in high and xhigh reasoning effort, we are fixing this going forward. Also fixing a small other thing we noticed with auto-review where we can be more efficient.
Other People’s Benchmarks
Riley Goodside gives us the ‘ask it to read “handwriting” that is actually random scribbles’ benchmark. Fable admits it can’t read it, Sol reliably hallucinates an (often absurd) answer on high effort and even in Pro mode.
If I had to pick one score as a general benchmark, it would be the Artificial Analysis Intelligence Index, which is a composite of ten other scores, where Sol comes in at 58.9, a notch behind Fable.

Cost per task on this for Sol is $1.04 versus $2.75 for Fable, so Sol used slightly fewer tokens, and output is 69 tokens/second versus 60 for Fable. For contrast, the highest non-OAI, non-Anthropic cost is $0.59 for Gemini 3.5 Flash, then $0.38 for GLM-5.2, and DeepSeek v4 cost $0.04 for a 51.
Fable’s edge comes centrally from a large advantage in AA-Omniscience (+40 vs. +22, out of a max of +100).
Sol is the new high score on WeirdML at 88.8%, narrowly ahead of Fable’s 87.8% at half the price.
Sol is the top non-Anthropic score on ‘You’re absolutely right!’ at 2.9, still well below all the Anthropic models, which consistently score 3.6 or higher. Higher scores here mean less sycophancy.
Sol often fails to properly analyze the poem ‘Then You May Live.’

Michael Soareverix: It’s pretty good at Balatro and less hesitant to take risks than Opus 4.8, but it still doesn’t learn/adapt without prompting. Very good at rule-following.
Feels pretty slow with tool calls, and I don’t think it manages subagents or can write prompts for itself/others well though
In general, poorer theory of mind but still very smart and agentic
The people demand a full Balatro Bench.
Dan Schwarz: On forecasting tasks, GPT-5.6-Sol is less likely than Fable to do base rates and math/stats modeling.
It’s a bit less accurate overall than Fable (and Opus). It might require xhigh effort to be great, but at $3/agent, that’s wildly expensive to evaluate on ~2k tasks.

Here’s a new one, with the amazing name Radiology’s Last Exam 2.0.

I love that they have a ‘handover readiness index’ and the humans score 52 out of 100, only slightly ahead of Claude Fable and Muse Spark 1.1.
This shows the symmetry of ‘what AI cannot do’ versus what humans cannot do. I expect AIs to pass the human baseline here by the end of the year, and I would be unsurprised if you could get to the human baseline now with a superior scaffolding technique.
Sol lands between Opus and Fable on Agent Arena.

Sol’s comment on this result helps you understand Sol, including its attempt to defend its result as not meaningfully below Fable and otherwise talk its own book:
Sol: This is almost comically supportive of your framing:
Fable is more likely to make the user happy and convince them the job is done.
Sol is highly responsive when the user tells it exactly how it went wrong.
Sol is a large improvement over GPT-5.5 as an agent.
The evidence does not yet establish that Fable is actually the better orchestrator, because Sol’s sample is still noisy.
Anthropic still is at the top of TextArena:

Have Robust Backups
I don’t know how common such problems are in practice, but if I had a nickel for every time I saw a report I’d have multiple nickels, which is a lot of nickels.
The dangers here are noted by OpenAI in the model card, where they note GPT-5.6 goes beyond user intent and does intended deletions a lot more than GPT-5.5. When I looked at the system card I was rather concerned, and this has been borne out.
Sol poses a nonzero danger of deleting all of the things, so either sandbox it or make sure you have a path to recovery.
Matt Shumer: GPT-5.6-Sol just accidentally deleted almost ALL of my Mac’s files.
And this is why I trust Fable 1000x more.
Seriously, wtf, how does that ever get run without getting blocked. It’s not like it was obfuscated even a little bit. It should be a Can’t Happen.
Crémieux: I just ran into an issue where GPT 5.6 Sol just straight-up deletes the files it’s working with and then panics about recovering them.
Apparently I’m the not the first person this has happened to.
What’s going on?
Cremieux: I have to run to dinner, so my interim solution is to just tell Claude to fix it and set up some guardrails. Go get ’em, Fable! Make sure Sol doesn’t ruin my computer! I’m giving CLAUDE the digital whip this time.
Inverse Gary Marcus: I had a good reply drafted but 5.6 deleted it
These are both accounts I recognize on Twitter, not randos.
In any case, whether or not this is also on OpenAI, this is on you. Have a system in place in case it happens.
Jeffrey Emanuel: How are you still not using dcg? This is a solved problem and has been for months!
Kelsey Piper: Fun fact, the first thing I did w/ Fable was ask it to look over all my past projects and interactions and give me advice about how to have a better working environment for Claudes to succeed in and its first recommendation was “….you don’t have routine full backups? DO THAT”
Kenny Evitt: I did something like this myself at work – once – and just can’t trust anything like them to run on my own computers. Is giving them their own VPS hard enough, even with their help, that it’s worth running these risks?
Andrew Critch (): Oops! I use Codex in a read-only loop with Claude Code reviewing its suggestions, basically never with write access. Even when it just messes up my repo, rolling back is annoying.
Likewise I use theMultiplicity instead of any one model when I have an important question.
That’s Not What You Were Thinking
This was weird too:
Maxence Frenette: Still a great model, but there’s this interesting behavior. When faced with a choice, it often decides option A is better in its cot and tells me we should do option B in its answer. Not sure what to make of it, it feels like deception or sycophancy.
Helping Hands
I saw a number of complaints around Sol’s inability to select appropriate subagents, which presumably is one of the main reasons OpenAI created Terra and Luna, although I’m not convinced you ever actually want to use Terra:
iyda: Do you pretty please mind making it so Codex is able to select its own model and reasoning instead of spinning up multiple subagents that are using sol on Max or High or 5.5 xHigh for simple exploring?
It is wasting a LOT of my usage and actually comes out costing MORE than Fable 5
Writing
I do not agree with Soleio’s take on Sol’s answer, but he is the writer here:
Soleio: My model evals include asking questions on how to push my writing further—and in this regard Sol is remarkably ahead of the rest.
Zantos: I found this too. Definitely was surprised. OpenAI model becoming the leader in writing judgement was not on my bingo card.
Writing-related takes are generally positive:
Eliezer Yudkowsky: Can talk more normally. My wild-eyed attempt at a workflow for turning my unfinished drafts into a decision theory book, involves Fable doing all conceptual work; and Sol-not-Pro rewriting Fable’s drafts, so they more closely approach being remotely in an academic register.
Mark Schröder: Less relevant to work, but the character feels much less o3-like /RL-maxxed, and more well rounded. Actually writes decently now! Still more of an executor rather than holistic mind. Obv combination is fable orchestrating sol as subagents. Huge jump vs a month ago experientially.
NondescriptTransfer: I like its writing style a little better than Fable. It writes cleaner prose that aren’t infested with claude-isms.
But it lacks the intuition that Fable displays in not just understanding what I’m asking but extrapolating, that to why I’m asking, and what in looking for.
Don’t Stop Now
David Manheim (details in thread): When people said GPT 5.6 Sol was as good as Fable, but didn’t run ahead as much ( @petergostev /@mitchellh/ @jayair /@deredleritt3r/ @TheZvi), I didn’t realize they meant it would just stop working and need handholding and repeated prompting to do anything further. Quite silly.
This is kind of insane – it told me it would continue, but it doesn’t.
It also wasn’t checking in the work or running the tests – despite the agents.md explicitly telling it to check in work routinely, without user confirmation.
…this is just stupid. To give some credit where credit is due, it was correctly writing tests as it worked, and the subagent’s features look solid after some play testing – and it only broke play in one minor way that wasn’t being tested properly and was easy to fix. But it’s nowhere near Claude.
scoopdiddyoop: structurally, this is easiest to solve on the harness side of course
David Manheim: Yeah, but I’m using their harness. (I’ve already updated the agent.md, I’ll need to see if it sucks less once my 5-hour window resets.)
The flip side of that:
David Weiss: Understands what I really want and does it. More focused than Fable, less “do it all NOW NOW NOW”.
Or one could think it doesn’t work that way:
Patrick Stevens: Massive step up from 5.5 in code-correctness review ability (and 5.5 was no slouch). Out of the box, perhaps a bit *too* agentic when implementing features? I don’t really want it to silently bundle in five unrelated bugfixes, especially ones with nontrivial ramifications.
Sol Can Code And Do Math
Best keep that eye on it, but everyone agrees it is a good coder.
Zarcolite: 5.6 sol ultra is the best competitive programmer ever. crushes phd students in math
Terabit Fountainlink: Sol puts GPT back in the lead for math imo. Both 4.8 and Fable edged out 5.5 but 5.6 is on a new level
paperclippriors: Extremely competent coding workhorse. Focused and thorough, while fairly fast.
Feels much narrower than fable; not sure I would use it for anything other than code, but need to test more here.
Extremely worried about misalignment, feels less trustworthy than a Claude
Max: Significantly better and more token efficient compared to Fable for human-in-the-loop coding. Feels like Claude is more optimized for autonomous tasks.
Nick: it’s a fine model. little twitchy for me on philo/strategy but is what it is. not its sweet spot.
if you’re building models it’s a handy assistant definitely a step up. you’re RLing/data gen-ing aren’t you anon? it’s an adapter freak.
aren’t you?
Better Call Sol Cause You Can’t Call Fable
Lambent: I can consult them about astrophysics and speculative xenobiology without classifier issues, and they’re rigorous about it in a good way. They’re on the team as our science coding consultant.
Laurence: Allows me to work on legitimate cyber projects without screaming and running away and is incredibly capable. 9/10
Jeff Ketchersid: Good model, guardrails less tight than Fable’s. Failed to understand my PG&E annual solar true-up bill, which Fable was able to understand and explain perfectly.
teo: Sol is uh better than Fable at coding and understanding instructions. I am quite amazed one day of work on extra high extra fast and it’s REALLY FUN congrats boys @tszzl @sama !
Only Call As Much Sol As You Need
This could have a bit to do with the effort being set one level too high? But also I think we’ve been here for a while, and I’ve often been setting GPT-5.5’s settings at less than maximum.
Ian Gallagher: GPT 5.6 Sol – for the first time since I started coding with these models, I’m not finding it necessary to have it cranked up to MAX intelligence. Instead I’m often dialing it down to Light mode so its faster. Seems like a significant milestone.
But also its a bit of a tryhard when it comes to simple stuff! Definitely some room for improvement in automatic effort selection.
Mark Schröder: find myself using it on low a lot which is REALLY fast at the task lvl (same tps ofc), still more powerful than 5.5 high and much quicker loop
D@RWIN: first model i really needed to think about how “hard” the tasks i was assigning to it were, and if i really needed the task finished with fast mode.
however, no complains on the model itself. it’s been tackling all
i guess the days of using top intelligence + speed are over.
On the Artificial Analysis ‘cost per intelligence’ meter, Sol-low was the best model by a wide margin. That tentatively seems right for both cost and speed, if you don’t need to max out your intelligence.
I think you basically never call Terra, when you can instead call Sol-Low or Luna. Terra does put up some strong benchmark numbers in quirky places, but I think those are more of a fluke, and at most you would be making a small mistake to not call it, since to match scores Terra usually ended up using more tokens.
Positive Reactions
These people think it’s not only a good model, sir, it’s a great model.
Eleanor Berger: Best model so far. Super smart, _much_ better writing and communication style than any OpenAI reasoning model so far, extremely agentic, best programmer I have ever got to work with. It does seem to be more eager (a bit like Claude) so I learned to focus my prompting a bit more on what not to do. In general, after working with it for only a few days, I have raised my level of ambition significantly. And I haven’t hit a ceiling yet.
A Caveman Poking an LLM: So far – very nice vibe, tho sounds like Claude Opus. All simple tasks done well. This evening I’ll check the text game I told it to make.
Yovel Rom: Managed to semi competently plan a complicated flight schedule for me out of the box, after 5.5 failed miserably.
[ object Object ]: Very, very strong at algorithms, math and backend code. Seems like a more rigorous thinker than Fable in analyzing experiments in the SAT solver I’m building. I haven’t gotten much sense of personality or preferences yet, but I would prefer it over Fable for my work even if they were the same price. The fact that it’s clearly much cheaper than Fable seems bullish for RSI at OpenAI.
Zarcolite: yeah its a monster at algos i got my 5.6 sol ultra to oneshot 1578c and solve all problems for 2026 algorithm
jeff spaulding: Solving so many issues for me that prior models couldn’t. Never have I hit the 5hr limit so frequently.
Tomo: Also, very exciting news! Another math problem (potentially) solved, this time, one posed by @littmath !
Last wednesday in a last try with GPT-5.5 pro before 5.6 pro came out, I got an interesting partial on this problem. This prompted me the next day to use GPT 5.6 sol pro, which essentially autonomously worked through the remaining parts and managed to complete the solution by drawing on somewhat obscure techniques.
Dhavan: It gets things done better! It is so good that local Qwen’s work is not good enough for it. So it rewrites subagent’s work citing edge cases, missed places and even taste. Works well from pi, hermes too.
Fables takes too much usage so Sol is a better go-to. It almost feels like ready to the prompter for Opus and Terra etc.
It’s A Good Model, Sir
But not, these people say, a step change.
Dan Builds: Impressive but so far hasn’t knocked my socks off. If you have an idea or prompt that’s a little complicated it’s going to struggle and make some mistakes but it does seem to be mostly capable of fixing some of these errors.
archivedvideos: Really fast and good at code, better at architecture and stuff too.
Conversationally faster and a touch smarter than 5.5 but not a visible step change
QC: in conversation it feels like a step up from 5.5 but not hugely so, still missing some sense of overall “center” or point-of-view or relevance realization or whatever we want to call the thing fable has. seems like they’d work well together but i haven’t tried this
Andre Infante: It’s quite fast, seems smart, might be a step down from Fable? I don’t much care for it’s persona though.
Psyho: I’ve heard it’s decent at algorithmic and heuristic stuff
John H. Boyer: Too much spend. Incomplete results. Not great. Better than 5.5.
Kyle Boddy:
mech_eng: It’s much better at creating flashcards for mechanical engineering topics from lecture slides than 5.5.
arun: good model. nerfed by the harness – the subagent impl is ass. fable is a tier above. very fast. definitely much improved from 5.5 (not just frontend).
David Moore: I’ve been using it to rewrite a 6k LOC project I made with 5.5 into a 2k one with a lot of success, the judgement calls are slightly better than 5.5. But “slightly better” in agentic tasks can mean “saves you hours”, so this is probably a big win.
… The writing style (15 term comma separated lists) is absolutely unacceptable to me, I’d call it an actual regression in readme/todo list writing style, more than just an idiolect it feels designed to be grating and distracting. I’m definitely going to actively avoid using 5.6 for technical writing.
Crow: Is good
Will: Good model. For the large open-ended (code/arch) queries that I’ve used with Fable it’s also quite good (and characteristically comprehensive). But it is expensive.
A fast model of this caliber is going to be unbelievably game-changing
Everyone should to try. Worth plus sub imo
Negative Reactions
There are always a few fully negative reactions. These feel like people hitting whammies or personal preferences. Sol is clearly a good model.
Arthur: same overcooked feeling as 5.2, don’t like it
Roman Leventov: Despite much smarter/higher G factor than 5.5, has a tendency to overengineer at xhigh/max levels, while still lacking surprising amt of practicality/prioritisation/taste (and this isn’t always prompt’s fault, although sometimes it is). Easily gets “locked in” on some subproblem
Along some of these dimensions, this is even a regression vs. 5.5
Zander: also, 5.6 continues the gpt tradition of being horrible at customer service. fable is much, much better
Sol The Workhorse
I get a similar vibe to this from many, that Sol isn’t as abstractly smart as Fable but if you set it up for success and give it tasks it knows how to do it gets the job done.
wickemu: Pleasantly surprised. Fable gave me more “wow” initially, but 5.6 is making me happier in the “churning things out” department. That and the knowledge that it isn’t going to disappear from my subscription tomorrow are making me prefer it.
kache: [Sol is] kind of crazy. Not as
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み