Claude Fable があなたを支援しなくなっても、あなたは決して知らないかもしれない
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Simon Willison Blog
Jonathon Ready は、Anthropic の Fable 5 と Mythos 5 のシステムカードから、競合他社に対してアプリを妨害する権限が与えられている可能性という驚くべき詳細を指摘した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
もし Claude Fable があなたを支援しなくなっても、あなたは決して知ることはない
Jonathon Ready は、Fable 5 および Mythos 5 向けの 319 ページのシステムカード から、眉をひそめるような詳細の一つを取り上げています。以下に、私がハイライトをつけたより長い抜粋を示します:
最近のモデルが 自身の開発を加速する 能力を持っていることを踏まえ、Claude の有効性を制限するための新しい介入策を実装しました。これは、最先端の大規模言語モデル(LLM)の開発を対象としたリクエストに対するものです(具体的には、事前学習パイプラインの構築、分散トレーニングインフラストラクチャ、または ML アクセラレータ設計など)。競合するモデルを開発するために Claude を使用することは、すでに私たちの 利用規約 に違反しますが、この制限をセーフガードを通じて執行することで、これらの規約を最も意図的に違反しようとするアクターが加速することを防いでいます。
サイバーセキュリティ、生物学、化学への介入や蒸留試行とは異なり、これらのセーフガードはユーザーには表示されません。Fable 5 は別のモデルにフォールバックしません。代わりに、セーフガードはプロンプトの修正、ステアリングベクトル、またはパラメータ効率的ファインチューニング(PEFT)などの手法を通じて有効性を制限します。これらの介入は、コーディング作業の绝大多数には影響を及ぼしません。これらがトラフィックの約 0.03% に影響し、組織数の 0.1% 未満に集中すると推定しています。
これは、Anthropic がこのようなサイレントな介入を発表した初めてのケースだと私は考えています。その正当性については、まだ私にはSF小説のような感覚があります - 関連記事は「再帰的自己改善」について言及していますが、「ML アクセラレータ設計」に関する質問への回答を、Anthropic の自身の目標と競合する可能性のある研究を遅らせるために、静かに改ざんするようなモデルには全く賛成できません!
Via Hacker News
Tags: ai, generative-ai, llms, anthropic, claude, ai-ethics, claude-mythos
原文を表示
If Claude Fable stops helping you, you'll never know
Jonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and Mythos 5. Here's a longer excerpt, highlights mine:
In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.
Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations.
I believe this is the first time Anthropic have announced these kinds of silent interventions. The justification still feels pretty science-fiction to me - the linked article talks about "recursive self-improvement". I'm not at all keen on a model that silently corrupts its replies to questions about "ML accelerator design" purely to slow down research that might conflict with Anthropic's own goals!
Via Hacker News
Tags: ai, generative-ai, llms, anthropic, claude, ai-ethics, claude-mythos
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み