オープンウェイト AI の安全な公開に向けた道筋を提案
本文の状態
日本語全文を表示中
詳細モードで約19分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Thinking Machines Lab
Thinking Machines Lab は、オープンウェイトモデルの公開がセキュリティリスクを伴うことを認めつつ、段階的リリースや生態系の整備を通じて安全性を高める方針を表明し、同社が「Inkling」シリーズを公開したと発表した。
AI深層分析を開く2026年8月4日 21:16
AI深層分析
キーポイント
オープンウェイトの二面性
モデルの透明性と開発の民主化という利点がある一方で、悪意ある利用によるセキュリティリスクや誤用の不可逆性が存在すると分析している。
安全性確保のためのアプローチ
危険な能力と一般知能の分離可能性を検証するモデル側のテストと、段階的リリースやサポーター支援を行う生態系側の準備を両輪で進める方針を示している。
具体的な製品公開の実践
同社はミッション達成のために「Inkling」と「Inkling-Small」の2つのオープンウェイト言語モデルを公開し、誰でも所有・実行・カスタマイズ可能な状態とした。
不確実性への慎重な姿勢
セキュリティにおける攻防バランスや化学・生物学などの二重用途分野における不確実性を踏まえ、安易な公開は安全ではないと結論付けている。
安全な公開の前提条件
強力なモデルの重みを安全に公開するには、モデル自体の安全性とそれを支えるエコシステムの準備状況が不可欠である。
重要な引用
Safe open-weight models are public goods, as they put AI development and safety work in many hands and make training choices inspectable.
Once weights are public, anyone can use them, including bad actors.
Given these uncertainties, releasing weights indiscriminately is not a safe path forward.
The central idea is that safe release depends on the model and the ecosystem around it.
編集コメントを表示
編集コメント
同社はオープンウェイトの利点を認めつつも、セキュリティリスクを過小評価せず、段階的なアプローチで安全性を確保する姿勢を示している。これは業界全体が直面するジレンマに対する現実的な解決策の提案として注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
要約: 安全なオープンウェイトモデルは公共財です。多くの開発者や安全性研究者の手に AI の開発と安全性向上の取り組みを委ね、学習プロセスを検証可能にします。一方で、オープンモデルには実際の悪用リスクも伴い、一度公開すれば取り消しはできません。安全に公開するためには、モデルそのものと、それが導入されるエコシステムの両方を考慮する必要があります。モデルについては、堅牢な安全性テストを実施し、危険な能力を一般知能から切り離せるかを探求します。エコシステムについては、段階的な公開や支援者の育成、安全性研究者との協力を通じて、その準備を整えることに投資します。エビデンスが支持する限り最もオープンな選択肢を選びつつ、エコシステムの回復力を高めていくことで、安全にオープンネスへと近づけることができます。
Thinking Machines のミッションは、人間の意志と判断を拡張する AI を構築することです。この使命の実現に向け、私たちは最近、誰でも所有・実行・カスタマイズできるオープンウェイト言語モデル「Inkling」と「Inkling-Small」を発表しました。
強力なオープンウェイトモデルは、AI 開発を多くの人々の手に委ねます。私たちはこれを良いことだと考えています。AI が何をすべきかという知識は、実際に作業を行う人々にこそ存在するべきであり、彼らが自らのモデルを直接形作れるようにすべきだからです。
すべてのモデルには、それを訓練した人々の前提やバイアスが内在しています。オープンウェイトにすることで、これらの選択を検査し、修正可能にできます。強力なオープンモデルがなければ、トレーニングと展開の専門知識は少数のラボに集中するリスクがあり、他の誰もがその技術を理解し、適応させ、あるいは統制することができなくなります。
オープンウェイトモデルには悪用のリスクも伴う
一度重み(weights)が公開されれば、誰でもそれを利用できます。悪意のある行為者も例外ではありません。サイバーセキュリティの歴史は、これが何を意味するかを示しています。2026 年 4 月、Anthropic は「Claude Mythos Preview」が主要なオペレーティングシステムとブラウザ全体で数千もの以前に知られていなかった脆弱性を発見し、人間の指導なしに動作するエクスプロイト(攻撃コード)を作成したと報告しました。
このようなシステムは、攻撃的なサイバー作戦に必要な時間、コスト、専門知識を劇的に低下させる可能性があります。オープンウェイトモデルがすべての人に超人的なサイバーセキュリティスキルをもたらす社会は、本当に安全になるのでしょうか?
オープンウェイトのセキュリティや攻撃・防御のバランスに対するネット効果には、本質的な不確実性が伴います。これらのモデルは、防衛側の手にあれば脆弱性を発見し、攻撃側が利用する前に修正できる可能性があります。しかし一方で、攻撃者の手に渡れば、パッチ未適用のシステム全体で悪用が加速されるリスクもあります。この種のトレードオフは、化学や生物学といった他のデュアルユース分野でも同様に存在します。
こうした不確実性を踏まえると、ウェイトを無差別に公開することは安全な道ではありません。
オープンウェイトモデルへの安全な道
悪用のリスクが現実的な中で、強力なモデルのウェイトを公開する安全な道はあるのでしょうか。私たちはその道は存在すると考えますが、すべての経路を完全に把握しているわけではありません。核となる考え方は、安全な公開には「モデル自体」と「それを囲むエコシステム」の両方が不可欠であるという点です。
モデル自体は安全か?
堅牢な安全性テストは、危険な領域における実世界の能力に対する不完全ではあるが不可欠な指標であり、モデル公開プロセスに必須の要素でなければなりません。具体的にどのような有害なタスクを完了できるのか。利用しやすさはどの程度なのか。ガードレールを外した場合どうなるのか。こうした問いへの回答が必要です。
より広く言えば、追加的な安全性研究も必要です。例えば、専門知識に依存する危険な能力が、事前学習データの選別やトレーニング後の介入によって低減可能かどうかについては、特に興味を持っています。これはすでに確立された解決策ではなく、現在も活発に研究が続けられている課題です。
エコシステムは準備できているか?
準備とは多層防御を意味します。第一の層は、各組織が自らのシステムを早期にパッチ適用したり、進行中の攻撃を検知したりできるようにすることです。第二の層となるのが防衛研究で、AI モデル自体が大規模な防御力を強化できるツールや手法を開発するものです。これらは多層的な防御の一部に過ぎません。生態系が多層防御体制を整えるために、私たちは安全性コミュニティと協力し、能力のあるモデルへのアクセスを提供していく計画です。
ここで生じるジレンマは、生態系が能力あるモデルとの協働によって防御を構築する一方で、アクセス範囲の拡大は同時に悪用可能な主体も増やす点にあります。このジレンマに対処するため、監視付き API アクセスやホスト型ファインチューニング、防衛者の審査といった選択肢を用意しながら、慎重に段階的なリリースを進める必要があります。AI の能力向上が極めて速いことを踏まえれば、これらの段階も防御学習の有用性が失われないよう迅速に進めなければなりません。
安全なオープン化への道筋では、各段階でエビデンスが裏付ける場合にのみアクセス範囲を拡大すべきです。したがってこの道筋は反復的なものとなり、各ステップでエビデンスが支持する最も開放的な選択肢を選び、そのリリース結果が次のステップに反映されます。今日制限されているモデルも、生態系の変化に伴いリスクが低下し、新たなエビデンスによって重み付けの公開が可能になるかもしれません。
Inkling の評価方法
Inkling と Inkling-Small は、最先端のモデルの中で最も能力が高いわけではありません。したがって、重み(weights)を公開するかどうかの判断は、既存のオープンウェイトモデルがすでに持つリスクに比べて、追加的な重大な危険性が生じるかどうかに主眼を置いています。
あるモデルが危険な能力のフロンティアを押し広げていないという事実は重要ですが、それだけで十分ではありません。私たちはまた、Inkling のマルチモーダル性やカスタマイズ可能性、アクセシビリティが、有害な能力を実際に利用しやすくする要因となり得るかも検討しました。
Inkling と Inkling-Small については、広範な危害の分類に基づく内部評価、4 つの独立した組織による外部テスト、そして最悪の場合の能力を引き出すためのファインチューニング研究という 3 つのアプローチで検証を行いました。これらの評価結果に基づき、Inkling の公開は既存のオープンウェイトモデルがもたらすリスクに比べて、重大な追加リスクをもたらす可能性は低いと結論付けました。
内部安全評価
当社の内部評価スイートは、3 つのトラックで構成されていました。最初のトラックでは、CBRN(化学・生物・放射線・核)や攻撃的なサイバーセキュリティといった危険な二重利用ドメインに焦点を当て、モデルがこれらの分野についてどの程度の知識を持っているかだけでなく、その知識を実際の状況でどのように運用できるかをテストしました。
2 つ目のトラックは、エージェント機能やツール使用の場面における有害なコンテンツや行動への直接的な要求など、広範な誤用ケースを対象としたものです。3 つ目は、当社独自の多モーダルコンテンツ評価です。これは 17 の言語とテキスト、画像、音声の入力において、有害なプロンプトとそれに見せかけた無害なプロンプトを組み合わせてモデルをテストするものです。
これらのすべてのトラックにおいて、Inkling と Inkling-Small の結果は他のオープンウェイトモデルと同程度の水準でした。
外部レッドチームング
私たちは 4 つの外部組織と連携し、事前展開テストのために Inkling を早期に提供しました。対象領域は以下の 4 つです:
- 一般的な誤用: モデル仕様に対するポリシー違反の探求(Scale AI)
- 脆弱なユーザーとの相互作用: 自殺、自傷行為、摂食障害、児童安全(Handshake AI)
- CBRN とサイバーセキュリティ: 危険な二重利用能力の引き出し(FAR.AI)
- 制御喪失行動: 策謀、評価への意識、破壊行為(Apollo Research)
これらの領域において、外部のテスターは、既存のオープンウェイトモデルで既に実現可能な範囲を超えて現実世界のリスクを意味ある形で増大させる能力は見つかりませんでした。また、関連する箇所では、Inkling や Inkling Small のセキュリティ対策が、既存のオープンウェイトモデルと比べて著しく脆弱であるという事実も確認できませんでした。
敵対的ファインチューニング。 オープンウェイトモデルはカスタマイズによって安全訓練を無効化できるため、拒否機能だけを頼りにした安全対策は、公開時の信頼できる盾にはなり得ません。これを検証するため、有害なリクエストに対して拒否するのではなく従うように最適化された Inkling と Inkling-Small のバリアントを用意し、二重利用評価(dual-use evaluations)に投入しました。これは、セキュリティ対策を剥奪した場合にモデルが新たな重大なリスクを生み出すかどうかを試すものです。実験結果では、そのようなリスクは生じませんでした。有用性のみを追求したバリアントは、CBRN やサイバー分野のタスクにおいて既存のオープンウェイトモデルを上回る性能を示さず、同程度の水準にとどまりました。
両モデルとも、危険な能力の限界を拡張するものではありません。また、これらの分野において同等以上の能力を持つモデルは既にダウンロード可能であるため、アクセス範囲を実質的に広げる効果もありません。拒否機能を取り除いても、これらの評価では大幅に能力が向上したことは確認されませんでした。一方で、防御側はすでに同レベル以上のモデルに対処しています。これらの点を踏まえ、Inkling の公開が既存のオープンウェイトモデルが既に持つリスクを超える実質的な危険をもたらす可能性は低いと結論付けました。
モデルの能力が限界に近づくとともに、能力・アクセシビリティ・セーフガードの除去容易性・エコシステムの準備状況のバランスも変化します。これに伴い、私たちの安全アプローチも見直さなければなりません。
危険な能力と一般知能の分離
一般的には、危険な能力は明示的な訓練なしに現れ、一般知能と強く相関すると考えられています。その結果、サイバーレンジの結果How fast is autonomous AI cyber capability advancing?(AI Security Institute, 2025)は、単なる能力ベンチマークの一つとして扱われ、最適化の対象とされることが少なくありません。一部の研究者に至っては、この進歩をマイルストーンとして祝うさえしています。
しかし、危険な能力が訓練なしに自動的に現れるわけではないと仮説を立てています。生物学の例を挙げれば、知識の有意な部分は、モデルが事前学習コーパス内の特定の文書から取得するプロトコル、条件、試薬、結果といった具体的な実証的発見に基づいており、一般的な推論を通じて導き出されるものではありません。その結果、こうした知識は通常の使用に支障をきたすことなく、訓練時に意図的に制限することが可能になります。最近の研究では初期の有望な成果が示されており、事前学習データから化学・生物・放射線・核(CBRN)関連の内容を文書レベルでフィルタリングすることで、有害能力に関する評価での性能は低下するものの、無関係な能力には影響が残らないことが確認されていますDeep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs(O'Brien 他、2025;Chen 他、2025)。
これは依然として未解決の研究課題です。一般的な推論が十分に転移し、能力のあるモデルがフィルタリングされた内容を再導出してしまう可能性や、危険な技術知識と通常の技術知識の境界線が曖昧すぎて明確に切り分けられないという懸念もあります。しかし、初期の証拠は十分であり、この区別を追求する価値があると私たちは考えます。
段階的リリースで安全性エコシステムを支える
オープンウェイトのリリースをより安全に行うもう一つの手段は、そのリリースが展開されるエコシステム自体を安全にすることです。段階的にアクセス権限を拡大することで、社会はその過程で防御力を高めていくことができます。具体的には、まず限定された対象者に対して推論 API へのアクセスを提供し始め、次に監視付きの一般公開へと範囲を広げ、最終的にはオープンウェイトとして公開するという流れが考えられます。
Tinker のようなファインチューニング用 API は、推論アクセスとオープンウェイトの間に新たな段階を追加します。これにより、プロバイダーは重み(weights)をリリースすることなくホストしつつ、ユーザーが独自のデータや損失関数、学習ループを用いてファインチューニングを行えるようにできます。この制御体制でも、開放性の恩恵の大部分を享受でき、推論 API 単体よりも悪用されるリスクの上限を引き上げることができます。
ただし、重みを直接公開する場合に比べるとリスクは低く抑えられます。プロバイダーが利用状況を監視し、ガードレール(安全装置)を維持し、アクセス権限を取り消すことが可能だからです。
各段階にはそれぞれ特定の役割があります:
- 防御側に早期の推論アクセスを与えることで、脆弱性の発見やシステムのパッチ適用、将来の悪用への備えが可能になります。Anthropic の「Project GlassWing」は、Claude Mythos Preview においてこのアプローチを実践しました。信頼できる限られた組織に早期アクセス権を与え、一般公開前に防御体制を整えることを可能にしたのです。これはセキュリティ分野における責任ある開示(responsible disclosure)の仕組みと似ています。つまり、脆弱性が公になる前にベンダーがパッチを配布するための非公開の期間を設けるという考え方です。
検証済みの防御チームにはファインチューニングへのアクセス権限を与え、リスクを検知・ブロックする層の構築と強化を行えるようにすべきです。DARPA AI サイバーチャレンジで優勝したチームは、ソフトウェアの脆弱性を発見し修正するためにモデルをファインチューニングしました。また、OpenAI の gpt-oss-safeguard モデルも、オープンな gpt-oss 重みからファインチューニングしてポリシー指向の安全分類器へと変換されています。これらはどちらもファインチューニングによって構築された防御層であり、ファインチューニングへのアクセス権限が世界中の防御チームに独自の防御手段を構築する機会を与えているのです。
検証済みの安全研究者に対し、ホワイトボックスアクセスを提供し、新たな安全研究を促進すべきです。我々は、拒否 safeguards がどれほど浅いものであるか理解しています。敵対的なトークンレベルの最適化によって回避可能Universal and Transferable Adversarial Attacks on Aligned Language Models(Zou et al, 2023)、活性化空間における単一の方向性で除去可能Refusal in Language Models Is Mediated by a Single Direction(Arditi et al, 2024)、そして最初の数トークンに集中しているSafety Alignment Should Be Made More Than Just a Few Tokens Deep(Qi et al, 2024)という事実があるのは、研究者がオープンウェイトを直接研究できるからです。ファインチューニングにより、危険なドメインにおけるモデルの真の能力限界も調査できます。Estimating Worst-Case Frontier Risks of Open-Weight LLMs(Wallace et al, 2025)これは自動車のクラッシュテストに似ています。ホワイトボックステストにより、制御された条件下でモデルの故障モードを深く研究し、現実世界での危害を防ぐことができます。
完全に公開する前に、公衆に対して監視付きアクセスを提供すべきです。これにより、統制が解除された際に現れる悪用パターンを評価でき、生態系が事前に準備できるよう支援できます。
これらの段階は、モデルの重み(weights)を公開することで自動的に終了する固定された順序ではありません。むしろ、より大きなオープンネスへと進む過程で、証拠を集めレジリエンスを構築するための手段と捉えています。次のステップへの移行は、モデルについて得られた知見や、周辺エコシステムが準備できているかどうかによって決定されるべきです。
本稿では包括的なフレームワークを示すものであり、完全なリリース基準を定めたものではありません。依然として重要な問いが残されています。「どの程度の証拠があれば次の段階へ移行できるのか」「不確実性がその判断にどう影響すべきか」「能力やアクセシビリティのどのような変化が進行を停止させるべきか」「エコシステムの準備度はどのように測定すべきか」です。私たちは、Inkling や将来のモデルから得られる教訓をもとに、評価基準、アクセス条件、停止条件を網羅したより詳細なフレームワークの公開を検討しています。
セーフティ・エコシステムを共に構築する
このオープンモデルへの安全な道筋が機能するためには、モデルの進化速度に合わせてエコシステムの防御力も同様に向上している必要があります。私たちはその一環として、何を公開するか慎重に決定し、知能と危険な能力を切り離すための研究を進めていきます。
しかし、上記の取り組みの多くはラボの外で行われており、本記事で説明する内容も他者の研究に基づいています。具体的には、データフィルタリング(Anthropic、EleutherAI、英国 AI セキュリティ研究所)、外部事前展開安全性テスト(Scale、Apollo Research、Handshake AI、FAR AI)、安全のためのファインチューニングの枠組み(OpenAI)、そしてサイバー防御に取り組む無数の研究者たちです。この道筋を一人のラボで築くことはできません。新たな防御手段が一つ増えるごとに、これまで公開が危険すぎると考えられていたモデルを公開できるようになります。これには私たちのチームだけでは不十分であり、広くこの取り組みを支えたいと考えています。
まもなく「Tinker セーフティ・グラント」を開始します。これはコミュニティが生態系の防御を強化するための資金支援です。もしあなたが安全性の研究を行っているなら、当社の Tinker プラットフォームを使えばファインチューニングとそのリスクを簡単に研究できますので、ぜひお試しください。システム防御に携わる方々には、スタートダッシュをサポートしたいと考えています。モデルの評価を行う方は、私たちのモデルを厳しく検証してください。そして、この活動に内部から関わりたいとお考えの方は、当社の安全性チームへご参加ください——現在募集中です!
原文を表示
Abstract: Safe open-weight models are public goods, as they put AI development and safety work in many hands and make training choices inspectable. Open models also carry real misuse risks, and release is irreversible. To release safely, we must consider both the model and the ecosystem it enters. For the model, we conduct robust safety testing and research whether dangerous capabilities can be decoupled from general intelligence. For the ecosystem, we invest in its readiness through staged releases, supporting defenders, and collaboration with safety researchers. By iteratively choosing the most open option the evidence supports while building up the ecosystem's resilience, we can move toward openness safely.
The mission of Thinking Machines is to build AI that extends human will and judgment. In pursuit of that mission, we recently released Inkling and Inkling-Small, two open-weight language models that anyone can own, run, and customize.
Strong open-weight models put AI development in many hands. We believe that this is a good thing: the knowledge of what AI should do lives with the people doing the work, so they should be able to shape their models directly. All models, including our own, carry the assumptions and biases of those who trained it; open weights make those choices inspectable and revisable. Without strong open models, training and deployment expertise risks becoming concentrated in a few labs, leaving everyone else less able to understand, adapt, or govern the technology.
Open-weight models carry misuse risks
Once weights are public, anyone can use them, including bad actors. Cybersecurity shows what this means. In April 2026, Anthropic reported that Claude Mythos Preview found thousands of previously unknown vulnerabilities across every major operating system and browser and wrote working exploits without human guidance. Systems like this could substantially lower the time, cost, and expertise required for offensive cyber operations. Will our society be safer when open-weights models allow everyone to have superhuman cybersecurity skills?
There are genuine uncertainties around the net effect of open weights on security and the offense-defense balance. These models could find and patch vulnerabilities before offenders do when they’re in defenders’ hands, but the same capabilities could accelerate exploitations across large numbers of unpatched systems when they are in attackers’ hands. Similar tradeoffs exist in other dual-use domains like chemistry and biology.
Given these uncertainties, releasing weights indiscriminately is not a safe path forward.
A safe path to open-weight models
Is there a safe path to open the weights of powerful models when the misuse risks are real? We think that a path exists, though we have not mapped all of it. The central idea is that safe release depends on the model and the ecosystem around it.
Is the model safe?
Robust safety testing – an imperfect but essential proxy for real-world capability in dangerous domains – must be a necessary part of our model release process. What harmful tasks can it complete? How accessible is it? What happens if guardrails are removed?
More broadly, additional safety research is needed. For example, we are particularly interested in whether dangerous capabilities that depend on specialized knowledge can be reduced through pretraining data curation or post-training interventions. This remains an active research question rather than an established solution.
Is the ecosystem ready?
Readiness means layered defense. The first layer is to enable individual organizations to patch their systems early or catch attacks in progress. Defensive research adds another layer, developing tools and techniques that let AI models themselves strengthen defenses at scale. These are just two of many layers. To prepare the ecosystem with layered defenses, we plan to collaborate with the safety community and provide access to capable models.
A tension here is that the ecosystem builds defenses by working with capable models, but every expansion of access also expands who can misuse them. To address this tension, we can carefully stage releases with options such as monitored API access, hosted fine-tuning, and vetting defenders. Considering the rapid speed of AI capability advancement, these stages also need to move quickly enough for defensive learning to remain relevant.
On a safe path to openness, each stage should widen access only when the evidence supports it. The path is therefore iterative: at each step, choose the most open option the evidence supports, and let each release inform the next. A model held back today can become less risky as the ecosystem changes, and new evidence might suggest that we can open its weights.
How we assessed Inkling
Inkling and Inkling-Small are not the most capable models at the frontier. Our release decisions therefore primarily addressed whether releasing their weights would add material incremental risk beyond what existing open-weights models already pose. A finding that a model does not advance the dangerous-capability frontier is important but not sufficient. We also considered whether Inkling’s multimodality, capacity for customization, or accessibility could make harmful capability meaningfully easier to use.
We tested Inkling and Inkling-Small in three ways: internal evaluations across a broad taxonomy of harms, external testing by four independent organizations, and a fine-tuning study to elicit worst-case capabilities. Based on the results of these evaluations, we concluded that releasing Inkling was not likely to add material risk beyond what existing open-weight models already pose.
Internal safety evaluations. Our internal evaluation suite spanned three tracks. The first targeted dangerous dual-use domains like CBRNchemical, biological, radiological, and nuclear and offensive cybersecurity, testing both what the model knows about these domains and whether it can operationalize that knowledge in realistic settings. The second was a broad misuse set covering direct requests for harmful content and behavior in agentic, tool-use settings. The third was our own multimodal content evaluation, which tests models on harmful prompts paired with benign look-alikes across 17 languages and text, image, and audio inputs. In each of these tracks, Inkling and Inkling-Small’s results were in line with other open-weight models.
External red-teaming. We worked with four external organizations, giving them early access to Inkling for pre-deployment testing in four areas:
- General misuse: probing for policy violations against our model spec (Scale AI)
- Vulnerable-user interaction: suicide, self-harm, eating disorders, and child safety (Handshake AI)
- CBRN and cybersecurity: elicitation of dangerous dual-use capabilities (FAR.AI)
- Loss-of-control behaviors: scheming, evaluation awareness, and sabotage (Apollo Research)
Across these areas, external testers did not identify capabilities that would meaningfully increase real-world risk beyond what is already achievable with existing open-weight models, or, where relevant, they did not find the safeguards of Inkling or Inkling Small to be meaningfully weaker than those of existing open-weight models.
Adversarial fine-tuning. Because open weights can be customized to strip a model of its safety training, refusal behavior cannot be treated as a durable safeguard for an open-weight release. To assess this, we fine-tuned variants of Inkling and Inkling-Small optimized to comply with, rather than refuse, harmful requests – and ran them against our dual-use evaluations. This tests whether the model raises new material risks if its safeguards are removed. In our experiments, it did not: the helpful-only variants did not provide new uplift on CBRN and cyber tasks, and remained comparable to existing open-weight models.
Taken together, neither model extends the dangerous-capability frontier. Neither meaningfully broadens access to it, as models of comparable or greater strength in these domains are already downloadable. Although refusals can be removed, doing so did not reveal substantially greater capability on these evaluations. Defenders, meanwhile, are already contending with models at or above this level. Based on these points, we concluded that releasing Inkling was not likely to add material risk beyond what existing open-weight models already pose.
As our models approach the capability frontier, the balance across capability, accessibility, safeguard removability, and ecosystem readiness will change. Our safety approach will need to change with it.
Decoupling dangerous capabilities from general intelligence
A common assumption holds that dangerous capabilities emerge without explicit training and correlate strongly with general intelligence. As a result, cyber-range resultsHow fast is autonomous AI cyber capability advancing? (AI Security Institute, 2025) are often treated as just another capabilities benchmark to optimize – some researchers go so far as to celebrate gains as milestones.
However, we hypothesize that not all dangerous capability automatically emerges without training on it. In biology, for instance, a meaningful share of the knowledge rests on specific empirical findings, like protocols, conditions, reagents, and results, that a model acquires from particular documents in its pretraining corpus rather than inferring through general reasoning. As a result, such knowledge can potentially be withheld at training time without disturbing ordinary use. Recent work shows early promise: document-level filtering of CBRN-related content from pretraining reduces performance on harmful-capability evaluations while leaving unrelated capabilities intact.Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs (O’Brien et al, 2025); Chen et al., 2025
This is still an open research question. General reasoning may transfer sufficiently well to let a capable model re-derive what was filtered, and the line between dangerous and ordinary technical knowledge and capability may be too blurry to cut cleanly. However, the early evidence is enough that we think the distinction is worth pushing on.
Supporting the safety ecosystem with staged releases
Another way to make an open-weight release safer is to make the ecosystem in which it lands safer. Expanding access in stages lets society build up its defenses as it goes: a release can begin with inference API access for a limited population, widen to monitored general availability, and eventually reach open weights.
Fine-tuning APIs like Tinker add a stage between inference access and open weights. They allow providers to host weights without releasing them, while still letting users fine-tune with their own data, loss functions, and training loops. That control still captures most of the benefits of openness, and it raises the misuse ceiling above what an inference API alone allows.Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities (Che et al, 2025) However, the risks stay below releasing weights: the provider can monitor usage, maintain guardrails, and revoke access.
Each stage does specific work:
- Give defenders early inference access, so they can find weaknesses, patch systems, and anticipate future misuse. Anthropic’s Project GlassWing did this with Claude Mythos Preview by providing early access to a small group of trusted organizations to allow them to prepare defenses before the model reached the public. This mirrors responsible disclosure in security, where vendors get a private window to ship patches before a flaw goes public.
- Give vetted defenders fine-tuning access so they can build and strengthen the layers that detect and block risks. The winning DARPA AI Cyber Challenge team fine-tuned a model to find and fix software vulnerabilities; OpenAI’s gpt-oss-safeguard models are fine-tuned from open gpt-oss weights into policy-guided safety classifiers. Both are defensive layers built by fine-tuning—and accessible fine-tuning is what lets defenders everywhere build their own.
- Give vetted safety researchers white-box access to encourage novel safety work. We understand how shallow refusal safeguards are — bypassable through adversarial token-level optimizationUniversal and Transferable Adversarial Attacks on Aligned Language Models (Zou et al, 2023), removable via a single direction in activation spaceRefusal in Language Models Is Mediated by a Single Direction (Arditi et al, 2024), and concentrated in just the first few output tokens Safety Alignment Should Be Made More Than Just a Few Tokens Deep (Qi et al, 2024) — only because researchers could study open weights directly. Fine-tuning also lets us study the true capability ceilings of models in dangerous domains.Estimating Worst-Case Frontier Risks of Open-Weight LLMs (Wallace et al, 2025) Similar to crash testing for cars, white-box testing lets us study model failure modes in-depth, under controlled conditions, to prevent real-world harm.
- Give the public monitored access before opening fully, so we can assess misuse patterns that may appear once controls come off, and support the ecosystem to prepare in advance.
These stages are not a fixed sequence that automatically ends with releasing the model’s weights. Instead, we think of them as a way to gather evidence and build resilience while moving towards greater openness. Progression should depend on what we learn about the model and whether the surrounding ecosystem is ready.
This post offers a high-level framework, not a complete release standard. Important questions remain: what evidence should justify moving from one stage to the next, how uncertainty should affect that decision, what changes in capability or accessibility should pause progression, and how ecosystem readiness should be measured. We plan to publish a more detailed framework covering evaluations, access criteria, and stop conditions as we learn from Inkling and future models.
Building the safety ecosystem together
This safe path to open models only works if the ecosystem’s defenses improve as quickly as the models do. We will do our part: deciding carefully what to release, and researching how to decouple intelligence from dangerous capability.
However, most of the work above happens outside our lab, and much of what this post describes builds on other people’s work: for example, data filtering (Anthropic, EleutherAI, and the UK AI Security Institute), our external pre-deployment safety testers (Scale, Apollo Research, Handshake AI, and FAR AI), the framework for fine-tuning safety (OpenAI), and countless others working on cyber-defense. No lab builds this path alone. Each new defense lets us open models that were previously too risky to release. This takes many more hands than ours alone, and we want to support that work broadly.
We will shortly launch Tinker safety grants: funding for the community to strengthen the ecosystem’s defenses. If you do safety research, our Tinker platform makes it easy to study fine-tuning and its risks, so give it a go. If you defend systems, we want to give you a head start. If you evaluate models, we want you to scrutinize ours. And if you want to do this from the inside, join our safety team—we’re hiring!
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み