Anthropic、新ツール「ダイナミック・ワークフロー」を搭載した Opus 4.8 をリリース
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TechCrunch AI
AI企業 Anthropic が、新しいツール「ダイナミック・ワークフロー」機能を追加した大規模言語モデル「Opus 4.8」を正式に発表した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
木曜日、Anthropic は最も高度な公開モデルの最新バージョンである Opus 4.8 をリリースしました。このモデルはあらゆる場所で利用可能で、標準価格は前回の Opus リリースと同じ水準です。
この新モデルは、Opus 4.7 がリリースされてからわずか 41 日後に登場しました。これは Anthropic にとって通常よりもはるかに短いアップグレードサイクルです(最新の Sonnet および Haiku モデルはそれぞれ 3 ヶ月と 7 ヶ月前のものです)。この迅速な対応は、一部のユーザーが失望したという Opus 4.7 の冷ややかな反応に関係している可能性があります。
その間隔には、OpenAI の Codex と Google の Gemini Flash モデル(注:Gemini Flash model)の重要な新リリースもあり、Anthropic が競争に遅れぬよう努力する圧力が高まっています。
Opus 4.8 は期待通り最上位クラスのベンチマーク結果を示していますが、特に注目されているのは、このモデルが不正確または不確実なデータをどのように処理するかという点です。発表記事において、Anthropic の初期テスターたちは、新モデルが「自身の作業に関する不確実性を指摘する可能性が高まり、根拠のない主張を行う可能性は低くなった」と見解を述べています。
この点を踏まえて、ブリッジウォーターの担当者による証言では、今回のアップグレードにおける最大の違いは「Opus 4.8 が分析の入力と出力に関する問題を積極的に指摘する傾向がある点」であり、他のモデルでは通常見落とされ、ユーザーが自ら発見しなければならない箇所だった」と述べています。
新しいモデルと共に、Anthropic は Dynamic Workflows という新機能をリリースしました。これは研究プレビューとして利用可能になります。このシステムは、Opus などの大規模モデルが数百の並列サブエージェントにわたる複雑なタスクを管理するのを支援するために設計されています。
「Claude Code と Opus 4.8 を組み合わせれば、既存のテストスイートを基準として、数千万行規模のコードベースに対する移行を、開始からマージまで一貫して実行できるようになりました」と、同記事は説明しています。
Anthropic は先月行った暫定的なプレビュー a tentative preview last month でサイバーセキュリティ上の懸念が浮上したため、最も高度な Mythos モデルの公開はまだ保留しています。しかし、同社は今日の Opus リリースにおいて、必要な安全対策が整い次第、Mythos のプレビュー期間もまもなく終了する可能性を示唆しました。
「これらの安全対策の開発は急速に進んでおり、今後数週間で Mythos クラスのモデルをすべての顧客に提供できるようになる見込みです」と同社は述べています。
当記事内のリンクを通じてご購入いただいた場合、私たちは少額のコミッションを獲得する可能性があります。これは私たちの編集の独立性には影響しません。
Russell Brandom は 2012 年以来テクノロジー業界を取材しており、プラットフォームポリシーと新興技術に焦点を当てています。以前は The Verge や Rest of World で勤務し、Wired、The Awl、MIT の Technology Review に寄稿しています。
russell.brandom@techcrunch.com または Signal(412-401-5489)までご連絡いただけます。
原文を表示
On Thursday, Anthropic released Opus 4.8, the newest version of its most advanced publicly available model. The model is available everywhere, with standard pricing at the same level as the previous Opus release.
The new model comes just 41 days after Opus 4.7 was released, a much faster upgrade cycle than normal for Anthropic. (The most recent Sonnet and Haiku models are three and seven months old, respectively.) The fast turnaround may have something to do with the chilly reception to Opus 4.7, which some users found disappointing.
That interval has also seen significant new releases for OpenAI’s Codex and Google’s Gemini Flash model, increasing the pressure on Anthropic to keep pace.
Opus 4.8 comes with the expected best-in-class benchmark results, but there’s also particular attention to how the model manages bad or uncertain data. In the launch post, Anthropic’s early testers found that the new model is “more likely to flag uncertainties about its work and less likely to make unsupported claims.”
Echoing this point, a testimonial from Bridgewater associates said the biggest difference in the upgrade was “Opus 4.8’s tendency to proactively flag issues with the inputs and outputs of an analysis, something other models routinely missed and left to the users to catch.”
Together with the new model, Anthropic launched a feature called Dynamic Workflows, which will be available in research preview. The system is designed to help larger models like Opus manage complex tasks across hundreds of parallel subagents.
“Claude Code alongside Opus 4.8 can now carry out codebase-scale migrations across hundreds of thousands of lines of code from kickoff to merge, with the existing test suite as its bar,” the post explains.
Anthropic is still holding back its most advanced Mythos model after a tentative preview last month raised cybersecurity concerns. However, the company hinted in today’s Opus release that the Mythos preview period might soon end, once necessary safeguards are complete.
“We’re making swift progress on developing these safeguards and expect to be able to bring Mythos-class models to all our customers in the coming weeks,” the company wrote.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review.
He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み