Sakana AI、セキュリティ特化モデル「Fugu-Cyber」発表
Sakana AI は、実世界セキュリティベンチマークで最先端性能を達成したマルチエージェントシステム「Fugu-Cyber」を新APIとして公開し、単なるモデルの導入ではなく専門人材との連携が不可欠であると強調した。
キーポイント
Sakana AI の新セキュリティ特化モデル Fugu-Cyber の発表
Fugu-Cyber は CyberGym で 86.9%、CTI-REALM で 72.1% の成功率を記録し、GPT-5.5-Cyber や Mythos-Preview と同等の最先端性能を達成した。
マルチエージェント・オーケストレーションによる防御機能
単一のエンドポイントから複数の専門エージェントを動的に調整し、複雑な多段階タスクを処理するが、ベンダーロックインのリスクを回避する設計となっている。
AI 依存に対する現実的な視点と限界の提示
Nikkei のレポートや自社の経験を踏まえ、最先端モデルへのアクセスだけでは実世界の脆弱性対策は完結せず、専門人材と深い統合が不可欠であると指摘している。
API 単体利用の限界と誤検知の問題
孤立して使用された場合、生モデルは誤検知(false positives)を発生しやすく、運用環境の文脈を理解するには追加の仕組みが必要であると警告している。
専門的検証ワークフローの統合
単独モデルの誤検知を防ぐため、サイバーセキュリティ専門家によるサブエージェントと人間を介した検証プロセスを組み合わせた運用環境が必要です。
責任あるリリースとアクセス管理
機密性の高いセキュリティワークフローに対応するため、悪用防止ポリシーに基づき、申請内容の厳格な審査を経てのみAPIへのアクセスが許可されます。
重要な引用
Achieving high scores on an evaluation is only the beginning of the story.
Simply having access to a frontier model like Anthropic's Mythos does not magically solve enterprise security.
Successful deployments require having both the human expertise in cybersecurity and access to frontier capabilities.
True enterprise defense requires more than just having access to a frontier model.
By orchestrating the world's best models into a unified system, we are delivering the realistic, resilient blueprint required for AI sovereignty.
影響分析・編集コメントを表示
影響分析
この発表は、AI セキュリティ分野における技術的成熟度を示すとともに、業界全体が抱える『AI 万能主義』への警鐘を鳴らしています。Sakana AI は単に高性能なモデルを提供するだけでなく、その実装には人的要素とシステム統合が不可欠であることを明確にし、企業に対する現実的な導入ガイドラインを提示しました。これは、セキュリティ分野における AI 活用の次のフェーズ——『ツールと人間の協調』——への重要な転換点となるでしょう。
編集コメント
Sakana AI は、単なる性能競争ではなく、AI セキュリティの実装における「人間の役割」を強調することで、業界の成熟度を高めています。この発表は、企業が最先端モデルを導入する際の現実的な課題と解決策を浮き彫りにしており、技術者にとって非常に示唆に富んでいます。

本日、当社のオーケストレーションモデル「Fugu」のアップデート版として、「Fugu Cyber」を発表します。
新 API エンドポイントとして提供される Fugu-Cyber は、現代のサイバー防御が抱える複雑な課題に対応するために設計されました。業界で最も難易度が高いセキュリティベンチマークにおいて最先端のパフォーマンスを達成し、CyberGym では成功率 86.9%、CTI-REALM では 72.1% を記録しました。これは、GPT-5.5-Cyber や Mythos-Preview など、サイバーセキュリティに特化した主要なフロンティアモデルと同等の性能です。

Fugu-Cyber は、実世界のセキュリティベンチマークにおいて最先端のパフォーマンスを発揮し、GPT-5.5-Cyber や Mythos Preview といったサイバー特化型のフロンティアモデルと肩を並べる結果となりました。
これらのベンチマークは、企業防御の核となる要素をテストするものです。CyberGym は、エージェントが複雑なコードベースを分析して実世界の脆弱性を特定・検証できる能力を評価します。一方、CTI-REALM は、生きた脅威インテリジェンスレポートを実働可能な検出ルールに変換する能力を測定します。
オリジナルの Fugu オーケストレーションモデルと同様に、Fugu-Cyber も単一のモデルのように振る舞うマルチエージェントシステムです。ユーザーは 1 つのエンドポイントにリクエストを送るだけで、システムが専門的なエージェント群を動的にオーケストレートし、複雑で多段階のタスクを処理します。これにより、特定のベンダーへの依存というリスクを回避できます。Fugu-Cyber は現在、sakana.ai/fugu で新 API エンドポイントとして利用可能です。
フロンティアなサイバー能力に対する現実的な検証
評価で高得点を叩き出すことは、物語の始まりに過ぎません。
最近、最先端モデルのサイバー能力について不安を煽る声が非常に大きくなっています。業界の多くは、「サイバー機能を備えた最先端モデルへのアクセス権を組織に与えるだけで、セキュリティ上の課題が即座に解決する」といった主張を展開しています。
しかし、私たちはこの議論を現実 grounded に再構築すべきだと考えています。
最近の日経デジタルガバナンスレポート(日本語)でも指摘されている通り、Anthropic の Mythos といった最先端モデルへのアクセス権があるだけでは、企業のセキュリティ課題が魔法のように解決されるわけではありません。実際には、大手金融機関を含む大規模組織であっても、これらのツールを実務に落とし込むことに苦戦しているケースが多く見られます。専門的な社内人材や、独自ソースコードへの深い統合がなければ、たとえ最先端のサイバー能力を備えたモデルであっても、現実世界の脆弱性を発見したり修正したりすることは容易ではありません。
日経の記事で指摘された課題は、日本最大の企業と連携してサイバーセキュリティ対策に取り組むサカナ AI の経験とも一致しています。成功する導入には、サイバーセキュリティ分野における人的専門知識と、最先端の能力へのアクセスの両方が不可欠です。
強力な推論能力を備えた高機能な API は、このパズルの極めて重要な一部です。しかし、それだけが完全な解決策ではありません。
API の先へ
単独でモデルを運用すると、必ず誤検知が発生します。適切なハーン(枠組み)なしでは、実際の生産環境の微妙なニュアンスを理解するのは困難です。真の企業レベルの防御には、最先端モデルへのアクセス権を持つだけでは不十分です。
私たちの経験則として、サイバーセキュリティ対策には、最先端モデルを深くローカライズされた専門知識と厳格な検証ワークフローとともに導入する必要があります。AI システムが潜在的な脆弱性を検出した場合、パッチの提案前に実際に本番環境でトリガーされるかどうかを確認するため、サイバーセキュリティに特化したサブエージェントや人間が関与するプロセスによる検証が必要です。
これがまさに Sakana AI が取り組んでいる課題です。
企業向けソリューション:Fugu-Cyber の最先端能力と企業セキュリティの架け橋
ここが Sakana AI の応用企業チームの出番です。私たちはコアエンジンだけでなく、それを安全に利用するためのインフラも構築しています。
現在、主要な日本の機関と緊密に連携し、これらのモデルを生産環境で運用するために必要な専用ハーンやワークフローの構築を進めています。Fugu-Cyber の生来の推論能力とセキュリティ専門家の実務経験を組み合わせることで、企業は高度な能力を持ちつつも極めて信頼性の高い自動化された脆弱性検証やその後のタスクを構築できるよう支援しています。
業界の専門家が指摘する通り、サイバー防御の未来は、単一のモデルでは達成できない高い性能を実現するために複数の AI モデルを組み合わせるシステムにかかっています。世界最高峰のモデルを統合し、一つのシステムとして運用することで、AI の主権を守るために不可欠な、現実的で強靭な青写真を提供します。
責任ある導入
Fugu-Cyber は機密性の高いセキュリティワークフローを取り扱うため、安全かつ責任ある形でリリースすることに注力しています。
安全な導入を確保するため、Fugu-Cyber は攻撃的な悪用を禁止し、業界の安全性基準に合致するよう更新された利用規約に基づいて公開されます。
Fugu-Cyber の API は Token プランにて利用可能です。
ユーザーは、利用目的や検証済みの連絡先情報を明記したアクセス申請フォームを提出してアクセス権限を取得する必要があります。当社のチームでは、各申請を厳格に手動で審査・承認し、許可された場合にのみ Fugu-Cyber へのアクセスを付与します。
引き続きシステムの検証を継続し、エンタープライズパートナーと連携しながら、この技術が重要インフラの強化と防御のために活用されるよう努めます。
Sakana Fugu-Cyber は本日、sakana.ai/fugu の API エンドポイントで新モデルとして利用可能です。
エンタープライズソリューションの詳細については、Applied チームまでお問い合わせください。

Sakana AI
ご参加をご検討ですか?
詳細は採用情報ページをご覧ください。
原文を表示

Today, we are releasing an update to our Fugu orchestration model: Fugu Cyber.
Available as a new API endpoint, Fugu-Cyber is purpose-built for the complexities of modern cyber defense. Fugu-Cyber achieves state-of-the-art performance on the industry’s most challenging security benchmarks, reaching a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM, comparable to leading cybersecurity-focused frontier models such as GPT-5.5-Cyber and Mythos-Preview.

Fugu-Cyber achieves state-of-the-art performance on real-world security benchmarks, matching cyber-focused frontier models like GPT-5.5-Cyber and Mythos Preview.
Together, these benchmarks test the core pillars of enterprise defense: CyberGym evaluates an agent’s ability to analyze complex codebases to verify real-world vulnerabilities, while CTI-REALM measures its capacity to translate raw threat intelligence reports into working detection rules.
Like our original Fugu orchestration model, Fugu-Cyber is a multi-agent system that behaves like a single model. You send a request to one endpoint, and the system dynamically orchestrates a pool of specialized agents to tackle complex, multi-step tasks without the risk of single-vendor dependency. It is now available as a new API endpoint in sakana.ai/fugu
The Reality Check on Frontier Cyber Capabilities
Achieving high scores on an evaluation is only the beginning of the story.
Recently, there has been a lot of fearmongering about the cyber capabilities of frontier models. Much of the industry narrative suggests that simply granting an organization access to a frontier model with cyber capabilities will instantly solve their security challenges.
We believe it is time to ground this conversation in reality.
As highlighted in a recent Nikkei Digital Governance report (in Japanese), simply having access to a frontier model like Anthropic’s Mythos does not magically solve enterprise security. In reality, large organizations, including major financial institutions, often struggle to operationalize these tools. Without specialized internal talent and deep integration into proprietary source code, a frontier model, even with state-of-the-art cyber capabilities, cannot easily uncover or patch real-world vulnerabilities.
The challenges pointed out in the Nikkei article also reflect Sakana AI’s own experience as we work with the largest Japanese enterprises to tackle cybersecurity challenges. Successful deployments require having both the human expertise in cybersecurity and access to frontier capabilities.
A highly capable API with strong cyber reasoning is an incredibly important piece of the puzzle. It is not the entire solution.
Beyond the API
When deployed in isolation, raw models will inevitably generate false positives. They will struggle to understand the nuances of a live production environment without the right harness. True enterprise defense requires more than just having access to a frontier model.
In our experience, cybersecurity solutions require deploying frontier models with deep, localized cybersecurity human expertise and rigorous verification workflows. If an AI system surfaces a potential vulnerability, it must be validated by sub-agents specialized in cybersecurity and human-in-the-loop processes to confirm whether it would actually trigger in a real environment before a patch is proposed.
This is the exact challenge that Sakana AI is solving.
The Enterprise Solution: Bridging the Gap Between Frontier Capabilities of Fugu-Cyber and Enterprise Security
This is where Sakana AI’s Applied Enterprise team comes in. We are not just building the core engine; we are building the infrastructure required to use it safely.
We are currently working closely with major Japanese institutions to build the specialized harnesses and workflows required to deploy these models into production. By combining the raw reasoning power of Fugu-Cyber with the real-world experience of security professionals, we are helping enterprises build automated vulnerability verification and other subsequent tasks that are both highly capable and deeply reliable.
As industry experts have noted, the future of cyber defense relies on systems that can combine multiple AI models to achieve higher performance than any single model could alone. By orchestrating the world’s best models into a unified system, we are delivering the realistic, resilient blueprint required for AI sovereignty.
Responsible Deployment
Because Fugu-Cyber deals with sensitive security workflows, we are committed to its safe and responsible release.
To ensure safe deployment, Fugu-Cyber is being released under an updated Acceptable Usage Policy that prohibits offensive misuse and aligns with the industry’s safety standards.
The Fugu-Cyber API will be available for the Token Plan.
Users will also need to apply for access by submitting an access request form detailing their intended use case and providing verified contact information. Our team will manually rigorously review and approve each application before granting users access to Fugu-Cyber.
We will continue to vet our systems and work alongside our enterprise partners to ensure this technology is used to fortify and defend critical infrastructure.
Sakana Fugu-Cyber is available today as a new model at our API endpoint on sakana.ai/fugu.
To learn more about our enterprise solutions, please reach out to our Applied team.

Sakana AI
Interested in joining us?
Please see our career opportunities for more information.
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み