MLPerf Endpoints v0.7、AI 推論性能評価基準の進化版として基盤リリース
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MLCommons Inference
MLCommons は AI インフラ調達を支援する新ベンチマーク「MLPerf Endpoints v0.7」を発表し、市場の最新動向に即した比較可能な評価基準の確立に向けた基盤整備を完了した。
AI深層分析を開く2026年8月1日 13:54
AI深層分析
キーポイント
新ベンチマークの目的と背景
MLPerf は従来のハードウェア購入判断から、あらゆる規模の企業がクラウドやニュークラウドを選定する際の意思決定支援へと進化し、エンドポイント評価を正式に開始した。
4 つの開発原則
市場の最新動向に追従する「Current」、競合他社を含む包括的な「Comprehensive」、コストや電力で正規化された比較可能な「Comparable」、そして分析を支援する「Commentary」の 4 原則が掲げられた。
初期結果と参加企業
Coreweave、Google、Intel、KRAI、Nvidia の 5 社が初期ベンチマークに参加し、3 つの基準で数桁にわたる性能差を示す結果を提出した。
今後のロードマップ
年内には v1.0 をリリースし、エージェントワークロードを含むベンチマークの拡大と、より広範なメンバーによるローリング型提出プロセスを開始する予定である。
30以上の支援企業との協力体制
AMDやOracleなど30社以上の支援企業がルールやプロセスのストレステスト、インフラ改善に貢献した。
重要な引用
MLPerf Endpoints is designed to meet this more diverse and complex demand for benchmarks to inform AI system procurement decisions.
Results must keep pace with the market. Buyers can't wait months for a benchmark round to include new hardware or models.
Later this year, we will deliver MLPerf Endpoints v1.0 with more buyer-centric rules, normalization, and an expanded set of benchmarks including agentic workloads.
With this release, we are just getting started.
編集コメントを表示
編集コメント
MLPerf の「Endpoints」化は、AI サービスが日常業務に浸透した現在、ベンダー選定プロセスをデータ駆動型で効率化する重要な転換点となる。特にローリング型の提出プロセス導入により、ベンチマーク結果の鮮度が市場速度と同期される点は実務において極めて価値が高い。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2018 年の設立以来、MLPerf は AI システムのパフォーマンスを測定する業界標準として君臨してきました。その間、大規模言語モデルにおけるワットあたりの推論性能は 100 倍以上に向上し、トレーニング速度も 50 倍の改善が記録されています [1] [2]。
過去 8 年間で AI 産業は成熟し、今や世界中の企業や消費者が日常的に AI サービスを利用する時代となりました。そこで本日、この成熟した業界により有用となるよう MLPerf を進化させることを発表します。それが「MLPerf Endpoints」の最初のバージョンです。
当初、MLPerf は限られた数のクラウドプロバイダーや下流システム構築者が、AI ハードウェアの購入判断を下す際に役立てるために設計されました。しかし現在では、推論用の計算リソースを調達することは、規模に関わらずすべての企業にとって重要な経営判断です。これは、ネオクラウド、従来のクラウドプロバイダー、マネージドサービスなど、複数の選択肢を同時に評価することを意味します。
こうした購入決定を下す側は、信頼性が高く、比較可能で、独立したパフォーマンスベンチマークを必要としています。また、業界が週次で新モデルをリリースするスピードに追いつけるよう、動的に変化し続けるベンチマークであることも不可欠です。
MLPerf Endpoints は、こうした多様かつ複雑なニーズに応えるために設計されました。企業購入者中心のベンチマーク開発を進める上で、以下の 4 つの基本原則が指針となっています:
結果は市場のスピードに追いつく必要があります。購入者は、新しいハードウェアやモデルを反映したベンチマークラウンドのために数ヶ月も待てません。
包括性:購入者は、購入可能な多くの競合する推論プロバイダー、システム、ワークロードを網羅するベンチマークを必要としています。
比較可能性:購入者は、調達判断を下すために、コストや電力で正規化されたベンダー間での「リンゴとリンゴ」の比較結果を確認する必要があります。
解説:これらの意思決定を購入者を支援するため、Endpoints はビジュアライゼーション、データフィルタリング、分析を通じてベンチマーク結果に追加の文脈を提供します。
MLPerf Endpoints v0.7 は基盤リリースであり、Coreweave、Google、Intel、KRAI、Nvidia からの初期結果が含まれています。3 つのベンチマークにまたがり、性能が数桁に及ぶ優れた成果を収めたこれらのメンバーに祝意を表します。これは、データセンター向けにより動的で包括的、かつ比較可能な推論ベンチマークスイートを構築するための基盤です。Endpoints では現在、mlcommons.endpoints.com で確認できる自動提出パイプライン、継続的なレビューツール、結果の動的ビジュアライゼーションをサポートしています。また、購入者中心のベンチマークへとルールを進化させています。
今年後半には、MLPerf Endpoints v1.0 をリリースします。より購入者視点に立ったルールや正規化を導入し、エージェントワークロードを含むベンチマークセットを拡大する予定です。その後、広範な会員を対象に継続的な提出プロセスを開始します。この継続的提出プロセスにより、Endpoints は市場のスピードに合わせて常に最新かつ正確な結果を提供し続けます。
MLPerf Endpoints の開発を導いてくださった 30 社以上の支援企業、AMD、アーゴンヌ国立研究所、Broadcom、Core 42、Dell、HPE、Lambda、Oracle、Red Hat に感謝申し上げます。彼らのパートナーシップは、ルールのストレステストやプロセスの改善、インフラの強化において極めて貴重なものでした。今回のリリースは、まさにスタートに過ぎません。
システムまたはサービスプロバイダーの皆様には、今こそご参加いただく時です。ルールの策定に参加し、ご都合の良い時期に提出を計画してください。興味と、ご提出予定のハードウェアについてフォームにご記入ください。コミュニティが共同で開発したルールは、今年後半の v1.0 リリースに向けて精査され続けています。この取り組みへの早期参加者が、Endpoints の方向性を形作ることになります。また、エンタープライズ購入者の方には、Endpoints が貴社と組織のために構築されていることをお伝えします。どのようにサポートすべきか、alejandro@mlcommons.org までメールでお知らせください。
推論における「基準となるベンチマーク」として、会員皆様から寄せられる信頼を何より重く受け止めています。MLPerf のこの未来を、皆様と共に築いていくことを心より楽しみにしています。
MLPerf エンドポイント v0.7 は、持続可能な AI を目指すための基盤リリースです。このバージョンでは、マイクロワットからメガワットまでの広範な電力消費を評価する「MLPerf Power」ベンチマークの重要性が強調されています。
関連する研究として、A. Tschand 氏らによる「2025 IEEE International Symposium on High-Performance Computer Architecture (HPCA)」での発表があります。彼らは MLPerf Power を通じて、機械学習システムのエネルギー効率を µWatts から MWatts の範囲で定量化し、持続可能な AI の実現に向けた指標を提供しました。
また、MLCommons は 2024 年 11 月 13 日、「MLPerf Training v4.1 Results — Press Briefing」を開催し、D. Kanter 氏らによってトレーニング結果の概要が発表されました。このプレスブリーフィングでは、最新のベンチマークデータと技術的知見が共有されています。
これらの情報は、MLCommons の公式ウェブサイト上で公開されており、持続可能な AI 開発における重要な指針となっています。
原文を表示
MLPerf has been the standard-bearer for measuring AI system performance since its launch in 2018. During that time, MLPerf has tracked over a 100X improvement in inference performance per watt for large language models and over a 50X improvement in training speed [1] [2]. Over the past eight years, the AI industry has matured, with AI services now used daily by enterprises and consumers worldwide. Today, we are announcing an evolution of MLPerf that will make it even more useful for this maturing industry: we are launching the first version of MLPerf Endpoints.
MLPerf was originally designed to help a relatively small number of cloud providers and downstream system builders inform their AI hardware purchasing decisions. Today, procuring inference compute is a salient business decision for companies of all sizes – and it means evaluating options across neoclouds, cloud providers, and managed services simultaneously. These buyers need reliable, comparable, and independent performance benchmarks to inform their purchasing decisions, and these benchmarks must be dynamic enough to keep pace with an industry that launches new models weekly.
MLPerf Endpoints is designed to meet this more diverse and complex demand for benchmarks to inform AI system procurement decisions. Four key principles are guiding our development of enterprise-buyer-centric benchmarks:
Current: Results must keep pace with the market. Buyers can’t wait months for a benchmark round to include new hardware or models.
Comprehensive: Buyers need benchmarks that cover the many competing inference providers, systems, and workloads available for purchase.
Comparable: Buyers need to see apples-to-apples results across vendors, normalized for cost or power, to inform procurement decisions.
Commentary: To support buyers in these decisions, Endpoints provides additional context for benchmark results through visualizations, data filtering, and analysis.
MLPerf Endpoints v0.7 is a foundation release with initial results from Coreweave, Google, Intel, KRAI, and Nvidia. We want to congratulate these members on their excellent results, spanning several orders of magnitude in performance across 3 benchmarks. This is the infrastructure on which we are building a more dynamic, comprehensive, and comparable inference benchmark suite for the data center. Endpoints currently supports automated submission pipelines, continuous review tooling, and dynamic visualization of results that you can see at mlcommons.endpoints.com. We are also evolving our benchmarking rules towards more buyer-centric benchmarks.
Later this year, we will deliver MLPerf Endpoints v1.0 with more buyer-centric rules, normalization, and an expanded set of benchmarks including agentic workloads and then open the rolling submission process to our broader membership. The rolling submission process ensures Endpoints will provide current and up-to-date results that move at the pace of the market.
We’re grateful to our 30+ supporters who have helped guide the development of MLPerf Endpoints, including AMD, Argonne National Laboratory, Broadcom, Core 42, Dell, HPE, Lambda, Oracle, and Red Hat. Their partnership has been invaluable in stress-testing our rules, processes, and improving our infrastructure. With this release, we are just getting started.
If you’re a system or service provider, now is the time to get involved, help shape the rules, and plan your submission at a time of your choosing. Complete this form to let us know you’re interested in submitting, and what hardware you want to submit on. The community-developed rules are actively being refined for the 1.0 release later this year, and early participants in this effort will shape Endpoints’ direction. If you’re an enterprise buyer, Endpoints is being built for you, and we want to know how we can best support you and your organization. Let us know by emailing alejandro@mlcommons.org.
We deeply value the trust our members place in us as the benchmark of record for inference. We are excited to build this future of MLPerf with all of you.
[1] A. Tschand, A. T. R. Rajan, S. Idgunji, et al., “MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from µWatts to MWatts for Sustainable AI,” in 2025 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2025. arXiv:2410.12032.
[2] D. Kanter, M. Ahmad, H. Kassa, and S. Rishab, “MLPerf Training v4.1 Results — Press Briefing,” MLCommons, Nov. 13, 2024. [Online]. Available: https://docs.google.com/presentation/d/1KSIJBvIV9OcswF1mVbhGRN0nUWbSgbYSoxu6dHCwajM/
The post MLPerf Endpoints v0.7: A Foundation Release appeared first on MLCommons.
AI算出
主要ニュースainew評価標準
AI ベンチマークの新たなバージョンが正式に発表され、具体的なベンダー名や設計原則が明記されているため新規性は高い。ただし、日本企業固有の情報や日本語一次情報がないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み