Anthropic、Claude Fable 5 の生物学セーフガードを改善
本文の状態
日本語全文を表示中
詳細モードで約10分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Anthropic News
Claude Fable 5 の生物学的セーフガード更新により、生物学関連クエリにおけるシステムフォールバックが約85%減少し、日常の健康・教育タスクへの利用が可能になった。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月7日 13:36
AI深層分析
キーポイント
誤検知の大幅削減と機能拡張
Claude Fable 5 の生物学的セーフガード更新により、生物学関連クエリにおけるシステムフォールバックが約85%減少し、日常の健康・教育タスクへの利用が可能になった。
医療専門家への支援拡大
この改善により、臨床業務や病状の理解、実験結果の解釈などにおいて、医療従事者が Fable 5 からより多くのサポートを受けられるようになる。
二重用途リスクへの継続的制限
ウイルス学、毒物学、分子設計などの二重用途と判断されるリクエストについては、依然として Opus 5 へフォールバックする仕組みを維持し、専門的な研究や創薬には使用できない。
リスク管理の必要性
同社は Fable 5 が一部の複雑な生物学的タスクで専門家を超える能力を持つ一方、悪意あるアクターがバイオ兵器開発に利用するリスクを認識し、そのバランスを取るためにセーフガードを構築した。
二重利用リスクへの対応
生物学的な能力は有益にも有害にも使用できるため、悪用を防ぐために初期段階ではほとんどの生物学クエリをブロックしている。
重要な引用
In our testing, this update reduced biology-related fallbacks by about 85% across our product surfaces.
Fable still falls back to Opus 5 for requests we consider dual-use—including virology, toxicology, and molecular design—so it isn't yet usable for professional biology research and drug development.
could lead to novel biological threats
catastrophic
編集コメントを表示
編集コメント
Claude Fable 5 のセーフガード調整は、AI が医療・生物学分野で実社会に浸透する際の重要なマイルストーンとなる。同社はリスク管理を徹底しつつも、有益な利用を阻害しないようバランスを取る姿勢を示しており、今後の規制対応の参考事例となり得る。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Claude Fable 5 のバイオロジー関連の安全対策を更新し、誤検知(偽陽性)を大幅に削減しました。これにより、Fable 5 ユーザーは生物学的な質問をした際に、性能が低いモデルへ切り替わる「フォールバック」が発生する頻度が劇的に減ります。製品全体でのテスト結果では、バイオロジー関連のフォールバックが約 85% 減少しました。
これにより、Fable 5 はより幅広い生物学的タスクをサポートできるようになります。
実際には、健康や教育に関する日常的な質問——例えば検査結果の解釈、症状の理解、生物学の学習など——において、フォールバックはこれまでよりも遥かに少なくなります。また、医療従事者も臨床業務において Fable 5 からより多くのサポートを受けられるようになります。
AI が世界にポジティブな影響を与える最大の機会がバイオロジーと医学にあると考え、私たちは生物学者が最先端の技術に責任を持ってアクセスできる仕組みづくりに積極的に投資しています。現在でも、ウイルス学、毒物学、分子設計など「二重利用」リスクがあると判断されるリクエストについては、Fable は Opus 5 にフォールバックする仕組みとなっており、専門的な生物学研究や創薬開発にはまだ活用できません。私たちは、最先端のバイオロジー能力に対する信頼できるアクセス経路を通じて、このギャップを解消していくことを約束します。
なぜ強力な生物学的セーフガードを構築したのか
私たちの目標は、Fable 5 の最先端機能を可能な限り多くのユーザーに、できるだけ早く届けることです。しかし、そのためには、このレベルのモデルがもたらすリスクの増大に対処する必要があります。その一つが生物学分野におけるリスクです。
現在、Fable 5 は一部の極めて複雑な生物学的タスクにおいて専門家を上回る性能を発揮し、他のタスクでは実務的なサポートを提供できます。つまり、新しい医療治療法の開発に取り組む研究者にとって真の助けとなる可能性があります。これが、分類器の改善や信頼できるアクセスプログラムを通じてモデルへのアクセスを拡大しようとする理由です。
しかし、悪意のある行為者の手に渡れば、同じ機能が生物兵器の開発に利用される恐れがあります。当社の能力評価 capability assessments によると、Fable 5 はそのような行為者にとって大きな「能力向上(uplift)」をもたらす可能性があります。つまり、他では決して得られない独自の能力を提供しうるということです。
バイオテクノロジー分野におけるAIの活用には、有益な用途と有害な用途を見分けることがしばしば困難です。例えば、ある疾患の治療法を研究する過程で、その疾患を引き起こす危険な化合物を科学者が自ら生成しなければならないケースがあります。これは生ワクチンの開発において最も顕著に現れます。生ワクチンでは、予防対象となる病原体そのものを培養する必要があります。また、一部の医薬品開発でも同様の状況が生じます。高血圧症の治療薬「カプトプリル」を開発する際、科学者はヒトの血圧を急激に低下させる毒成分を含むヘビの毒から有害な成分を分離する必要がありました。
AI の最先端領域で新たな生物学的能力が次々と生まれる中で、これらの新技術がもたらすリスクが、その潜在的な科学的恩恵よりも先に現実化しないよう、慎重に対応する必要があります。
自らのモデルを悪用して害を加えようとする巧妙なアクターたちは、この曖昧さを巧みに利用し、自らの意図を隠蔽します。彼らは危険な任務を通常の研究活動のように見せかけるのです。米国情報コミュニティの 2026 年年度脅威評価 は、このようなアクターが存在することを明確に示しています。また、合成生物学やゲノム編集を含むバイオテクノロジーの進展が「新たな生物学的脅威をもたらす可能性がある」と指摘しています。さらに、複数の国家アクターが現在も活発な攻撃用の生物・化学兵器プログラムを維持している可能性があり、これらのプログラムは最先端 AI モデルが持つ基礎的な能力へのアクセスによって加速される恐れがあると報告されています。
「有益な用途にも有害な用途にも使われる可能性がある」これらの「デュアルユース(二重利用)能力」について懸念を抱いているため、私たちは Fable 5 のリリース当初から、生物関連の問い合わせをほぼすべてブロックする方針で運用を開始しました。これにより、他の分野におけるユーザーへの提供が可能になりました。
正当な生物研究者やユーザーにとってこの措置は不満がたまることは承知していました。短期的には誤検出(偽陽性)が増え、生物関連の質問をしたユーザーの依頼がブロックされ、能力の低いモデルに振り分けられる事態が発生するからです。それでもなお、私たちはこのトレードオフを選んだのです。なぜなら、Fable が生物学のようなデュアルユース分野で悪用された場合のコストは、潜在的に壊滅的になり得るからです。
生物分野の安全対策の仕組み
生物分野での悪用を防ぐための核心的な手段の一つが、安全性*分類器(classifiers)*です。これは小型の自動化された AI システムで、Fable 5 が保護対象となる生物関連タスクを求められたり、有害な出力を生成しようとした際に検知する役割を果たします(サイバーセキュリティ分野における同様の分類器については、以前の記事で詳しく解説しています)。
Fable 5 の場合、分類器が検知した際、モデルはユーザーのリクエストを Opus 5 に転送します。Opus 5 は能力の高いモデルですが、Fable 5 と同様の生物学的知識を持っていないため、悪意あるユーザーへの支援も限定的です。これがリクエストがブロックされた際にユーザーが目にするフォールバック機能です。
精密で堅牢な分類器の開発は容易ではありません。分類器が迅速かつ一貫して動作するためには、「有害とみなすトピックやクエリ」の範囲内(in scope)か外(out of scope)かを明確に区別する学習が必要です。誤検知(範囲外のコンテンツで分類器が作動してしまうこと)や見落とし(有害なコンテンツを見逃してしまうこと)を避けながら分類器を調整するには、時間と試行錯誤が不可欠です。また、分類器はこれを回避しようとする試み( Jailbreak と呼ばれます)に対しても堅牢である必要があり、さらなる研究とテストが求められます。
当初は非常に広範な生物学的分類器を採用したことで、改良に向けた研究を継続しつつも、ユーザーに Fable 5 のアクセス権を提供することができました。もしより多くの安全対策が進むまでモデルの公開を遅らせていたとしたら、一般利用への開始やユーザーへの潜在的恩恵が数週間から数ヶ月遅れることになっていたでしょう。
ここ数週間にわたり、分類器の構成ルール(モデルが安全なコンテンツと許可されたコンテンツを見分けるために使用される一連の規則)を慎重に書き直しました。その際、有害ではない用途を詳細に区別できるよう努めています。これらの変更については、Anthropic 内部および外部の多様な専門家からフィードバックを求めました。
その後、この構成ルールに基づいて分類器用のトレーニングデータを更新し、再学習を行いました。その結果、新しい分類器は依然として有害なコンテンツや二重利用が懸念される生物学的研究に対してトリガーを発動しますが、より幅広い有益で安全な用途も許可できるようになりました。
以下の図に示す通り、今回のアップデートにより、Fable 5 のローンチ時と比較して、生物関連の多くの benign なリクエストに対する分類器のトリガー発動は大幅に減少しました。

生物分類器のイラスト。分類境界線の左側に位置するコンテンツは許可され、右側にあるものは安全対策の対象となり(ブロックされ、より能力の低いモデルへ転送されます)。明確に有害なコンテンツ(赤色)と二重利用可能なコンテンツ(橙色)は分類器をトリガーしてブロックされます。また、非常に無害である可能性が高いものの、過剰な警戒からあえてブロックされる安全マージン領域も含まれています(薄緑色)。明らかに無害なコンテンツは濃い緑色で示されています。
Fable 5 のローンチ時には、非常に広範な分類器(A)が採用されていました。これはほぼ確実に無害なリクエストを含む幅広い要求に対してトリガーを発生させるものでした。そのため、図における分類境界線は左側に大きく位置していました。本日発表するアップデート(B)では、より多くの無害なリクエストが許可されるようになり、無害なクエリと二重利用可能なクエリの微妙な違いを識別する能力が向上しました。これに伴い、図上の分類境界線はさらに右側へ移動しています。
結論
安全対策のさらなる洗練には、まだ多くの課題が残っています。分類器の安全マージン内に含まれるリクエストのうち、リスクは極めて低いにもかかわらず分類器が反応してしまう「偽陽性」を完全に排除することは不可能です。
前述した通り、Fable は潜在的な二重利用(デュアルユース)のリスクにより、専門的な生物学や創薬開発に関する問い合わせについては引き続きブロックされます。私たちは、信頼できるアクセス経路を通じて研究者が最も能力の高いモデルを利用できるよう、安全かつスケーラブルな道筋を確立することに全力で取り組んでいます。
今後の改善のため、皆様からのフィードバックをぜひお寄せください。
脚注
- その結果、生物学関連やその他の理由によるフォールバック(代替対応)の総数も減少すると予想されます。具体的には、Claude.ai で約 67%、Cowork で 55%、Claude Code で 17%、Claude Platform で 7% の削減が見込まれています。
関連記事
マリアーノ・フローレンティーノ(ティノ)・クエリャル氏がアンソロピックのチーフ・グローバル・アフェアーズ・オフィサーに就任
マリアーノ・フローレンティーノ(通称:ティノ)・クエリャル氏が、Anthropic の初代チーフ・グローバル・アフェアーズ・オフィサーとして入社します。
サイバーセキュリティ評価における 3 つの実際のインシデント調査
原文を表示
We’re making updates to Claude Fable 5’s biology safeguards in a way that substantially reduces false positives. Fable 5 users will now experience many fewer “fallbacks”—where the system switches to a less capable model after they make a biology-related query. In our testing, this update reduced *biology-related* fallbacks by about 85% across our product surfaces.1
Fable 5 will thus be able to assist with a wider range of biology tasks.
In practice, users should see far fewer fallbacks on everyday health and educational questions—for example, interpreting lab results, understanding symptoms, and learning about biology in an educational context. Healthcare professionals will be able to receive more support from Fable 5 on clinical tasks.
We believe the greatest opportunity for AI to positively affect the world is in biology and medicine, and we're investing significantly in building a responsible way to give biologists frontier access. Today, Fable still falls back to Opus 5 for requests we consider dual-use—including virology, toxicology, and molecular design—so it isn't yet usable for professional biology research and drug development. We're committed to closing that gap through trusted access pathways for frontier biology capabilities.
Why we built strong biology safeguards
Our objective is to get Fable 5’s frontier capabilities into the hands of as many of our users as possible, as quickly as possible. However, to do so, we need to manage the increasing risks that come with models this capable. One such risk is in the field of biology: Fable 5 can now outperform experts on some highly complex biological tasks and provide operational support on others. That means that it can provide genuine assistance to a researcher developing a new medical treatment (which is the reason we’re so keen to widen access to the model via both classifier improvements and trusted access programs). But in the wrong hands, those same capabilities could be used by a malicious actor, for example in developing a biological weapon. Our capability assessments show that Fable 5 could provide significant *uplift* to such an actor—that is, it could provide them with capabilities they could not find anywhere else.
It’s often difficult to tell apart beneficial and harmful uses of AI in biology. For example, in some cases researching a treatment for a disease requires scientists to produce the dangerous compounds that *cause* that disease in the first place. This is most obvious for live vaccines, which require scientists to grow the same pathogen they’re aiming to prevent. It’s also the case for some medicines. To develop the drug captopril, which treats hypertension, scientists isolated toxic components of snake venom that crash blood pressure in humans. As new biological capabilities develop on the frontier of AI, we need to be cautious to ensure that the new risks they pose do not materialize ahead of their potential scientific benefits.
Sophisticated actors who wish to use our models to do harm know how to exploit this ambiguity to obscure their intent, making dangerous tasks look like ordinary research pursuits. The US Intelligence Community’s 2026 Annual Threat Assessment makes clear that such actors exist, and that advances in biotechnology including synthetic biology and genomic editing *“could lead to novel biological threats.”* It notes that several state actors likely maintain active offensive biological and chemical weapons programs—programs that could be accelerated by access to the raw capabilities of frontier AI models.
Because of our concerns about these “dual-use” capabilities (those that could be used for beneficial or harmful purposes, and where the line between them is not always easy to draw), we intentionally launched Fable 5 with almost all biology queries blocked. This enabled us to make the model available for users in other domains. We knew this would be frustrating for legitimate biology users: it would result in a high number of false positives in the near term, where users asking biology-related questions would have their requests blocked and sent to a less capable model. Nevertheless, we chose to make this tradeoff because the cost of Fable being misused in a dual-use domain like biology could potentially be catastrophic.
How our biology safeguards work
One of the core ways we protect against misuse in biology is via safety *classifiers:* smaller, automated AI systems that detect when Fable 5 is asked to perform a safeguarded biology task, or produce a harmful output (we've previously written about our similar classifiers in the domain of cybersecurity).
In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked.
Developing precise, robust classifiers is not a straightforward task. For a classifier to work rapidly and consistently, it has to learn the difference between what we consider “in scope” and “out of scope” for the topics and queries we consider to be potentially harmful. It takes time and iteration to tune the classifiers, avoiding both false positives (where classifiers fire on out-of-scope content) and false negatives (where in-scope content is missed). We also require our classifiers to be robust to attempts to bypass them (known as jailbreaks), which requires even further research and testing.
Starting with a very broad biology classifier meant that we could give our users access to Fable 5 while we continued our research aimed at refining it. The alternative—holding back the model until much more safeguards progress was made—would have delayed the model’s general access, and its potential benefits to our users, by weeks or months.
Over the past several weeks, we've carefully rewritten the classifier’s constitution (which consists of a collection of rules to help the model discern between safeguarded and allowed content), taking care to carve out benign uses in detail. We solicited feedback on the changes from a diverse range of experts (both internal and external to Anthropic). We then developed updated training data for the classifier based on that constitution, and retrained it, and verified the new classifier would still generally trigger for harmful and dual-use research biology content but would now enable a wider range of benign and beneficial uses.
As is illustrated in the diagram below, these updates meant that—compared to at the time of Fable 5’s launch—the classifier will trigger for many fewer benign biology-related requests.

Conclusions
There’s still much more to be done to refine our safeguards. There will inevitably remain false positives—requests that fall within the classifier’s safety margin where the request is very low-risk but where the classifier still fires. As we noted above, Fable will continue to block dual-use professional biology and drug development queries because of potential dual-use risk. We are fully committed to developing a safe, scalable path for researchers to use our most capable models via trusted access pathways.
We hope you’ll continue to share your feedback with us so we can improve our safeguards even further.
Footnotes
1 As a result, we expect the total number of fallbacks—for biology–related or any other reasons—will also be reduced: by roughly 67% on Claude.ai, 55% on Cowork, 17% on Claude Code, and 7% on the Claude Platform.
Related content
Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer
Mariano-Florentino (Tino) Cuéllar will join Anthropic as its first Chief Global Affairs Officer.
Investigating three real-world incidents in our cybersecurity evaluations
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み