Artificial Analysis、知能指数インデックス v4.1.1 を更新し評価モデルをアップグレード
本文の状態
日本語全文を表示中
詳細モードで約2分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
Artificial Analysis は評価指標 v4.1.1 を公開し、グラダーモデルの更新と 𝜏³-Banking のバージョンアップを実施したが、スコアへの影響は限定的である。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月7日 05:18
AI深層分析
キーポイント
評価指標のバージョンアップ
Artificial Analysis が Intelligence Index を v4.1.1 に更新し、グラダーモデルと 𝜏³-Banking の最新バージョンを適用した。
評価手法の刷新
HLE, AA-LCR, AA-Omniscience の各チェックにおいて、GPT-5.6 Luna (medium) を採用し、従来のモデルに代えた。
スコアへの影響
評価の堅牢性が向上した結果、全体的なスコアはわずかに増加したが、モデル間の順位変動は限定的である。
重要な引用
We have updated the Artificial Analysis Intelligence Index to v4.1.1 - this patch release upgrades our grader models, and brings the latest 𝜏³-Banking version to Artificial Analysis
HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively.
Claude Opus 5 remains in the #1 position with an Index of 63.
編集コメントを表示
編集コメント
評価指標の更新は、モデル性能の変化を正確に捉えるために不可欠なプロセスである。今回の変更はグラダーの能力向上によるものであり、開発者はスコアの絶対値よりもトレンドや相対的な順位変動に注目すべきだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Artificial Analysis Intelligence Index を v4.1.1 に更新しました。今回のパッチリリースでは、採点モデルのアップグレードと、最新の 𝜏³-Banking バージョンの導入を行いました。
開発者にとって最も有用な合成指標であり続けるため、Artificial Analysis Intelligence Index には定期的なアップデートを実施し、評価項目や独自の手法を改善しています。今回の更新は既存の評価セットの信頼性を高めるためのマイナーリリースです。
モデル全体のランキングは概ね安定しており、採点の堅牢性が向上したことでスコアがわずかに上昇しました。Claude Opus 5 は引き続き Index 63 でトップの座を守っています。
主な変更点:
➤ 𝜏³-Banking が v1.0.1 に更新されました(Sierra ベース)。これにより、最新のアップストリームタスクバージョンに対応し、不具合が発生した経路からの回復時に正答性を確保する採点パイプラインが改善されています。
➤ HLE、AA-LCR、AA-Omniscience の評価を GPT-4o、Qwen3 235B A22B 2507、Gemini 3 Flash Preview からそれぞれ GPT-5.6 Luna(ミディアム)へ変更しました。これらはいずれも、採点検証において人間の判断との合致度が高い現代の高性能モデルとして統一されました。
➤ スコアへの影響は限定的で、ほとんどのモデルは Intelligence Index 上で 1 ポイント未満の変動にとどまりました。最大のスコア上昇は Muse Spark 1.2(xhigh)で +2.7 ポイントとなり、上位の順位は依然として同じモデルが占めています。
公開スコアは v4.1.1 を採用しました。これにより、Artificial Analysis 上のすべてのモデル結果が今回の変更を反映し、最新の整合性のある独立した評価手法に基づいて算出されるようになりました。

原文を表示
We have updated the Artificial Analysis Intelligence Index to v4.1.1 - this patch release upgrades our grader models, and brings the latest 𝜏³-Banking version to Artificial Analysis
To keep the Artificial Analysis Intelligence Index the most useful synthesis metric for developers, we make regular updates to the included evaluations and our independent methodology. Today’s update is a minor one to keep our existing evaluation set as reliable as possible.
Overall model rankings remain largely consistent, with a slight increase in scores due to improved grading robustness across the updated evaluations. Claude Opus 5 remains in the #1 position with an Index of 63.
Key changes:
➤ 𝜏³-Banking now runs v1.0.1 from Sierra, updating to the latest upstream task versions and improved grader pipeline that resolves correctness errors in trajectories that recover from unhappy paths
➤ HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation
➤ The effect on scores is small: most models move by less than a point on the Intelligence Index. The largest increase occurred for Muse Spark 1.2 (xhigh, +2.7 points), and the same models hold the top of the leaderboard
Published scores are now using v4.1.1, so all model results on Artificial Analysis now reflect these changes and use our latest consistent, independent methodology.

関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み