NVIDIA、AI ネイティブストレージの暗号化・圧縮・整合性チェックを高速化
本文の状態
日本語全文を表示中
詳細モードで約22分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
NVIDIA Developer Blog
NVIDIA は「Vera Storage Benchmarks」を発表し、AI ネイティブストレージにおける暗号化、圧縮、整合性チェック、復旧処理の速度向上を示した。
AI深層分析を開く2026年8月4日 08:14
AI深層分析
キーポイント
AI エージェントワークフローにおけるストレージの役割拡大
エージェンティアイが知識を検索し、永続的メモリにアクセスする際、ストレージはデータ供給と保存を継続的に担うため、従来の読み書き以上の機能が求められる。
CPU 処理ボトルネックの解消に向けた新アーキテクチャ
NVIDIA は Vera CPU の性能をストレージデータパスに直接導入し、暗号化や圧縮などの処理がアプリケーション応答性を阻害しないようにする。
従来の x86 CPU に対するベンチマークでの優位性
NVIDIA の発表によると、Vera は暗号化・復元・整合性チェック・圧縮などの処理において x86 CPU を上回る性能を示すことが確認されている。
インフラコストとパフォーマンスのバランス改善
従来の CPU によるスケーリングではコア数や電力、冷却コストが増大するが、Vera の導入により処理速度を維持しつつこれらの課題を緩和できる。
Vera CPU のベンチマーク結果
Vera は x86 CPU を上回る暗号化・復元・整合性チェック・圧縮性能を示し、CPU と電力のオーバーヘッドを削減しながらデータ処理能力を向上させる。
重要な引用
Storage is an active part of every agentic AI workflow.
Supplying and preserving this data requires more than basic reads and writes.
Faster SSDs and networks cannot deliver their full potential if the processor securing, protecting, and preparing the data cannot keep pace.
The benchmark results show Vera outperforming the x86 CPU across encryption and decryption, recovery, integrity checking, compression and decompression, and a multi-stage storage pipeline.
編集コメントを表示
編集コメント
エージェンティアイの普及に伴い、ストレージシステムの役割が単なる保存から積極的なデータ処理へ移行している背景を反映した発表である。CPU リソースの最適化は、大規模な AI ワークフローの実用化において不可欠な要素となっている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ストレージは、すべてのエージェント AI ワークフローにおいて能動的な役割を果たしています。エージェントがエンタープライズ知識の取得や永続メモリへのアクセス、キーバリュー(KV)キャッシュデータの再利用、ツールの実行、そして新たな結果の生成を行う際、ストレージシステムはエージェントの推論ループを動かすデータを継続的に供給し、維持し続ける必要があります。
各エージェントのステップでは複数のストレージ操作がトリガーされ、これらの操作は数千もの並行して動作するエージェントや、ますます拡大するコンテキストウィンドウにわたって繰り返されます。このデータを供給し維持するには、単純な読み書き以上の処理が必要です。AI 推論は GPU で実行されますが、エージェントプロセス、ツール呼び出し、データ管理タスク、そしてこれらを支えるストレージサービス自体は CPU で動作します。
書き込み時には、ストレージはデータの圧縮や暗号化、チェックサムの計算、冗長性の算出を行います。読み込み時には、データをアプリケーションに返す前に、その検証、復号、解凍、あるいは再構築が行われます。これらの機能は AI システムのセキュリティとレジリエンスにとって不可欠です。しかし、各機能にはデータがストレージパスを通過する際に追加の CPU 処理が必要となります。
エージェントの並行性(複数のユーザーや AI エージェント、並列実行されるタスク)やコンテキスト量が拡大するにつれ、ストレージはアプリケーションの応答速度やトークン生成を制限することなく、これらの作業をより多くこなさなければなりません。加速されたコンピューティングに必要な速度でデータを供給できることが求められます。
これらの機能の多くはデータパス上に直接配置されており、1 つでも処理が遅れると、広範なデータのフロー全体が影響を受けます。従来の CPU でこれらをスケーリングしようとすると、コア数や電力消費、冷却コストが増大し、インフラコストが高騰する一方で、パフォーマンスは依然として最も遅い工程に依存したままとなります。データを保護・準備するプロセッサの処理速度が追いつかない限り、高速な SSD やネットワークもその真価を発揮できません。
ストレージ処理のギャップを埋める
NVIDIA Vera BlueField-4 STX Storage Processor は、AI ネイティブなデータプラットフォームを支える NVIDIA STX の基盤となる重要なコンポーネントです。これにより、NVIDIA Vera CPU の性能をストレージデータパス内に直接持ち込むことが可能になりました。NVIDIA Rubin GPU にデータを供給し続けるために設計された同じ Vera CPU アーキテクチャが、CPU 側でのストレージ処理も加速します。
ベンチマーク結果では、Vera は暗号化・復号化、回復処理、整合性チェック、圧縮・解凍、そして多段階のストレージパイプラインにおいて、x86 CPU を上回るパフォーマンスを示しました。これらの性能向上により、ストレージプラットフォームはより少ない CPU 負荷と電力消費で大量のデータを処理し、必須のエンタープライズサービスを提供できるようになります。また、圧縮スループットの向上は、必要なストレージ容量や帯域幅の削減にも寄与します。
本記事では、BlueField-4 STX に搭載された Vera CPU が、エージェント AI が要求するストレージ処理をどのように加速するかを解説します。これにより、AI ネイティブなストレージプラットフォームは、より多くのデータを安全に保護・検証・圧縮できるようになり、ストレージ処理のスループットと効率も向上します。
Vera CPU アーキテクチャ:ストレージの二重の要求に対応して設計
Vera CPU には、Armv9.2 命令セットと完全に互換性のある NVIDIA 設計の Olympus CPU コアが 88 基搭載されています。また、176 スレッドをサポートする NVIDIA Spatial Multithreading 技術も採用しています。
これらのコアは、NVIDIA Scalable Coherency Fabric (SCF) と Small Outline Compression Attached Memory Module (SOCAMM2) LPDDR5X メモリと組み合わされることで、AI ファクトリー規模での高いスループットを持つ CPU 実行を実現しつつ、強力なシングルスレッド性能を維持します。
SCF は、コア間、共有キャッシュ、メモリコントローラー、I/O を横断する一貫性のあるオンダイデータパスを提供し、最大 3.4 TB/s のバイセクション帯域幅と 164 MB の統合 L3 キャッシュを実現しています。これにより、ワークロードがプロセッサ全体にスケールしても、アクティブなコアは共有データへの高帯域幅かつ予測可能なアクセスを確保できます。
SOCAMM2 LPDDR5X メモリサブシステムはこのファブリックを補完し、合計で最大 1.2 TB/s のメモリ帯域幅、あるいはコアあたり最大 14 GB/s を提供します。これにより、帯域幅集約型かつ高度に並列なワークロードにおいても、NVIDIA Olympus コアへの供給が維持されます。
モジュール式でフィールド交換可能なメモリは、データセンターインフラに必要な保守性と信頼性を保ちつつ、LPDDR5X の高い電力効率を両立しています。
ストレージの基礎機能は、CPU に2 つの異なる要求を課します。まず、各データストリーム内では、暗号化、整合性チェック、回復処理、圧縮・解凍が、その後のストレージ処理に進む前に迅速に完了する必要があります。そのため、コアあたりの持続的なパフォーマンスが重要になります。次に、システム全体としては、これらの操作は多数の並行ストリームで実行され、データがキャッシュやメモリを繰り返し通過するため、帯域幅と予測可能なレイテンシも同様に重要です。
Vera はこの両方の要件に対応します。Olympus コアは、広範な命令スループット、高度な分岐予測、深いアウト・オブ・オーダー実行、ベクトル演算および暗号化リソースを組み合わせることで、制御処理とデータ処理のコードにおいて、各コアが命令スループットを維持できるよう支援します。
NVIDIA の空間マルチスレッディング、モノリス型の計算ダイ(SCF)、統合された L3 キャッシュ、そして高帯域幅の SOCAMM2 メモリは、アクティブなコアにデータを供給し続ける一方で、スレッド間の干渉を減らし、負荷下でのデータアクセスの予測可能性を高めます。これらの機能全体が、暗号化、整合性チェック、パリティ計算、圧縮、そして多段階ストレージパイプラインにおける測定された性能向上の説明となります。
これにより、BlueField-4 STX ストレージプロセッサは、CPU リソース、電力、冷却の比例増大なしに、並行するデータストリーム上でより多くの CPU 側ストレージ処理を維持可能になります。
基盤となるストレージ性能の測定
ストレージ処理は、読み込み・書き込み・回復の各パスで繰り返し実行されます。そのスループットと効率性は、CPU 側の処理が SSD やネットワークに追いついているか、あるいはデータ経路におけるボトルネックとなっているかを判断する重要な指標となります。本記事で紹介するストレージ用マイクロベンチマークは、これらの機能を分離して測定し、プロセッサの貢献度を明確にします。その結果から、Vera でより高スループットで効率的なストレージサービスを実現するために、CPU がどの程度の性能と余裕を持っているかがわかります。
各テストは、メモリ上に既に保持されているデータを用いて単一プロセス内で実行されます。特に明記がない限り、ファイル I/O、ディスク性能、ネットワーク、コマンド起動、および外部デバイスのボトルネックは除外されています。ベンチマークセットでは、OpenSSL、Zstandard、LZ4 といった一般的なライブラリに加え、Arm および x86 プロセッサで利用可能なネイティブ命令を活用するように最適化された同等の実装も含まれています。
専用テストフレームワークが各ワークロードを一定の条件で実行し、バッファサイズ、スレッド数、CPU の配置、計測タイミング、正誤検証、結果収集などを統括します。使用されるアルゴリズムやソフトウェア実装の多くは広く採用されています。このフレームワークは Vera と x86 間で同一のワークロード定義と制御条件を適用するため、プロセッサ間の公平な比較が可能となります。ソースコード、スクリプト、固定されたソフトウェアバージョン、設定、および結果ファイルを用いれば結果の再現も可能ですが、本ベンチマークセットが市販の公的ベンチマークとしてそのまま利用できるわけではありません。
これらの測定値は、安全なデータ転送やレジリエンス、容量効率、サービス密度に影響を与えるストレージの基盤ブロックにおけるVeraのパフォーマンスを明らかにしています。実際の運用環境では、同じデータに対して複数の操作が適用されるため、CPU処理要件が累積します。
運用環境のストレージソフトウェアでこれらの個別操作を組み合わせた場合、そのパフォーマンスが集積スループットとCPU効率に寄与します。個々のプリミティブや多段階パイプライン全体でのスループット向上により、ストレージソフトウェアはSSDやネットワークに追いつくためのCPUリソースを確保でき、並行データフローのサポートや必要なデータサービスの効率的な適用が可能になります。ただし、ストレージシステム全体の性能やGPUパフォーマンスの結果を定量化するには、エンドツーエンドのテストが依然として必要です。
より高速な暗号化でAIデータをさらに保護
AIファクトリーでは、モデル資産、企業知識、エージェントのコンテキスト、プロンプト、出力、顧客データなど、機密性の高い情報を処理します。AES-128は、保存時および転送中のデータの暗号化に広く採用されています。暗号化は書き込みパスに直接位置するため、そのスループットがボトルネックになる前にプラットフォームが毎秒処理できる安全なデータの量を決定します。Veraは、比較対象として使用されたx86 CPUと比較して、AES-128の暗号化スループットを最大1.43倍向上させます。

復号は読み取りパスで対応する処理を行います。Vera は比較対象の x86 CPU と比べて、AES-128 復号のスループットが最大 1.29 倍向上します。

暗号化スループットが高いと、ストレージシステムは書き込みを制限することなくより多くのデータを保護できます。また、復号が高速化すれば、保護されたデータをエージェントやアプリケーション、アクセラレータへ返すまでの時間を短縮できます。これにより、ストレージのパフォーマンス予算の多くを消費することなく、増加する AI データの保護と転送が可能になります。
AI データの保護と復旧を高速化
ストレージプラットフォームでは、ドライブやノード、データ断片が利用不能になった際にデータを保護するために、消去符号化(erasure coding)を採用しています。リード・ソロモン符号化は冗長性を生成し、その冗長性を利用して欠落または破損したデータを再構築します。符号化処理は通常、書き込みパスで行われます。一方、再構築や修復、劣化した状態での読み取り時に復旧処理が行われます。これらの操作では計算資源とメモリアクセスパターンが異なるため、パフォーマンス結果も異なります。
Vera は、復旧ワークロードにおいて x86 CPU と比較してリード・ソロモンスループットを最大 3.26 倍向上させます。

高いリード・ソロモンスループットにより、ストレージシステムは保護されたデータをより高速に書き込み、欠落したデータを迅速に再構築できます。特定の効率性測定では、Vera は利用可能な CPU パワーの範囲内でより多くの保護処理を完了し、復旧時間の短縮や通常のデータサービスとの競合低減に貢献します。
高いスループットでのデータ整合性の検証
データは移動、保存、検索される過程で正しさを保つ必要があります。巡回冗長検査(CRC)はチェックサムを作成し、ストレージシステムが偶発的な破損を検出するために使用します。
CRC とリード・ソロモン符号は補完的な役割を果たします。CRC はデータが期待された結果と一致しなくなったことを検出し、リード・ソロモン符号は欠落または破損した情報を再構築するために必要な冗長性を提供します。
ストレージシステムでは、バッファ間でのデータコピー中も含め、書き込み経路と読み取り経路の両方で CRC を計算することがあります。データ量が膨大になるにつれ、これらのチェックが CPU リソースの有意な割合を消費するようになります。
Vera は x86 CPU と比較して、CRC32C のスループットを最大 3.67 倍向上させます。

CRC32C のスループット向上により、ストレージシステムは整合性チェックによる読み書きの制限を受けることなく、より多くのデータを検証できます。エージェント型ワークロードにおいては、これにより CPU 側の処理遅延を抑えつつ、信頼性の高いコンテキスト、永続メモリ、およびエンタープライズデータを提供することが可能になります。
高速な圧縮・展開でデータフットプリントを削減
エージェント型 AI は、コンテキスト、ログ、チェックポイント、検索データ、中間出力、永続的メモリなど、膨大な量のデータを生成します。これらのデータを保存するために必要なストレージ容量を削減し、転送に必要な帯域幅を抑えるのが圧縮です。一方、読み取り時にデータを復元するのが伸長処理です。
アーキテクチャによっては、圧縮はデータ書き込みの直前に行われる場合もあれば、書き込み後に行われる場合もあります。一方、伸長は通常、圧縮されたデータを読み取る際に必要となります。圧縮スループットはデータをどれだけ速く圧縮できるかを決定し、伸長スループットはストレージシステムがエージェント型 AI アプリケーションにデータを返す速度に影響します。
以下のベンチマークでは、これらの操作をそれぞれ独立して測定しています。
Vera は x86 CPU と比較して最大で3.29 倍高い圧縮スループットを実現し、測定したすべてのスレッド数においてその優位性を維持しています。

伸長ベンチマークは、圧縮データを参照する際の対応する操作を測定したものです。Vera は並行処理下で x86 CPU と比較して最大1.72 倍高い伸長スループットを提供し、並列で動作するワーカースレッド数が増えるほどその優位性は拡大します。

圧縮パフォーマンスはアルゴリズム、データ特性、圧縮レベル、バッファサイズ、スレッド数によって異なります。したがって、これらの結果は測定されたワークロードに特化したものです。高い圧縮スループットを実現することで、ストレージシステムは CPU 側の処理性能を維持しつつ、書き込み・保存・転送するデータ量を削減できます。
高い圧縮解除スループットにより、ストレージシステムは解凍データをエージェントやアプリケーションへより迅速に返却できるようになります。これらの機能を組み合わせることで、ストレージ容量と帯域幅への負荷を軽減しつつ、アジェンティック AI ワークロードの増加に伴い、各プロセッサが処理するデータ量を増やすことが可能になります。
多段階のストレージ書き込みパスの加速
ストレージシステムでは、データサービスが独立して実行されることは稀です。セキュアな書き込みパスでは、まずデータを圧縮してフットプリントを削減し、その後書き込み前に暗号化することが一般的です。今回のベンチマークセットには、メモリ上に存在するパイプラインが含まれており、各データバッファに対して順に圧縮と暗号化を適用します。個々の操作を測定する先行するベンチマークとは異なり、このテストでは 2 つの CPU 集約型ストレージ機能が連続して実行される際の、全体のパイプラインスループットを測定します。
ベンチマークで比較対象となった x86 CPU を使用した 2 段階の圧縮・暗号化パイプラインに対し、Vera は最大で 3.21 倍の高いスループットを実現します。

この結果は、Vera の性能優位性が個々の操作測定を超え、データを削減し保護するストレージ書き込みパスを代表する多段階シーケンスにも及ぶことを示しています。*多段階パイプラインの高性能化により、ストレージシステムはプロセッサあたりの処理データ量を増やし、エージェント AI のデータ量が膨らんでも書き込みスループットを維持しやすくなります。
Vera を用いたエージェント実行とストレージのスケーリング
エージェント AI では、CPU での実行とストレージ処理が同じ AI ファクトリのデータパスの一部となります。
CPU はモデル呼び出しの間のツール実行、コード処理、検索、分析、データ処理ステップを担います。一方、ストレージシステムはこれらのステップが必要とするデータを確保し、保護・検証・圧縮して返却します。AI ファクトリ全体で利用可能な CPU、電力、冷却資源という限られた枠組みの中で、両者は同時にスケールする必要があります。
Vera は、エージェント AI 時代における CPU 依存タスクの高速化を目的に設計されました。NVIDIA Vera Rubin では NVIDIA GPU のホスト CPU として機能し、エージェントの実行をサポートします。単体での Vera は、エージェント向けツールにおいてコアあたりのパフォーマンスを最大 1.8 倍向上させます。また、BlueField-4 STX では、AI ネイティブストレージプラットフォームが使用する CPU 側の処理を駆動しています。
ベンチマーク結果は、Vera が以下の機能を加速することを示しています。セキュリティのあるデータアクセスのための暗号化・復号化、保存先の再構築や修復時に欠落または破損したデータの高速な再構成を実現するリードソロモン復元、高スループットな整合性検証のための CRC32C、そして容量と帯域幅の要件を削減しデータ検索を加速するための圧縮・展開です。
これらの機能を個別に、あるいは多段階の書き込みパス内で持続的に実行することで、Vera は AI ネイティブストレージプラットフォームが CPU 側の遅延を抑えつつ、安全で信頼性の高いデータを処理・返却することを可能にします。また、CPU リソースや電力、冷却能力に比例して増大することなく、より多くの並列データフローと高いサービス密度をサポートします。特定のワークロードではワットあたりの測定パフォーマンスも向上し、利用可能なプロセッサの電力予算内で CPU 側のストレージ処理をさらに多く実行できるようになります。
計算機とストレージで共通の Vera CPU アーキテクチャとソフトウェアツールチェーンを採用することで、エージェントの実行とそれを支えるデータインフラストラクチャの両方のスケーリングに対する一貫した基盤を提供します。
Vera CPU と BlueField-4 STX についてさらに詳しく
これらのストレージ処理結果を支えるアーキテクチャの詳細については、NVIDIA Vera CPU の記事「Agentic AI 向けに最大のスレッドパフォーマンスを構築した Olympus コア」をご覧ください。
また、NVIDIA BlueField-4 STX ストレージプロセッサのデータシートをダウンロードし、AI ファクトリーのストレージ、ネットワーク、セキュリティを加速する NVIDIA BlueField インフラストラクチャプロセッサについて詳しく学んでください。
原文を表示
Storage is an active part of every agentic AI workflow. As agents retrieve enterprise knowledge, access persistent memory, reuse key-value (KV) cache data, execute tools, and generate new results, storage systems must continuously supply and preserve the data that moves the agent reasoning loop.
Each agent step can trigger multiple storage operations, and those operations can repeat across thousands of concurrent agents with increasingly larger context windows. Supplying and preserving this data requires more than basic reads and writes. AI inference runs on GPUs, but agentic processes, tool calls, data management tasks, and the storage services that support them run on CPUs.
During writes, storage may compress and encrypt data, calculate checksums, and calculate redundancy. During reads, it may validate, decrypt, decompress, or reconstruct data before returning it to the application. These functions are essential to the security and resilience of AI systems. Each function also requires additional CPU processing as data moves through the storage path.
As agent concurrency (multiple users, AI agents, or tasks running in parallel) and context volumes grow, storage must perform more of this work without constraining application responsiveness or token generation; it must supply data at the rate required for accelerated computing.
Many of these functions sit directly in the data path; one delayed operation can slow the broader data flow. Scaling them with conventional CPUs can require more cores, power, and cooling, increasing infrastructure cost while still leaving performance dependent on the slowest step. Faster SSDs and networks cannot deliver their full potential if the processor securing, protecting, and preparing the data cannot keep pace.
Closing the storage processing gap
The NVIDIA Vera BlueField-4 STX Storage Processor, a key component of the NVIDIA STX foundation for AI-native data platforms, brings the NVIDIA Vera CPU performance directly into the storage data path. The same Vera CPU architecture designed to keep NVIDIA Rubin GPUs fed also accelerates the CPU-side storage processing.
The benchmark results show Vera outperforming the x86 CPU across encryption and decryption, recovery, integrity checking, compression and decompression, and a multi-stage storage pipeline. These gains enable storage platforms to process more data and apply essential enterprise services with less CPU and power overhead, while higher compression throughput helps reduce storage capacity and bandwidth demands.
This post explains how the Vera CPU in BlueField-4 STX accelerates the storage processing required by agentic AI, helping AI-native storage platforms secure, protect, validate, and compress more data while increasing storage-processing throughput and efficiency.
Vera CPU architecture: Built for storage’s dual demands
Vera CPU includes 88 NVIDIA-designed Olympus CPU cores that are fully compatible with the Armv9.2 instruction set. The CPU supports 176 NVIDIA Spatial Multithreading threads. It pairs these cores with the NVIDIA Scalable Coherency Fabric (SCF), Small Outline Compression Attached Memory Module (SOCAMM2) LPDDR5X memory to sustain strong single-thread performance, and high-throughput CPU execution at AI-factory scale.
The SCF provides a coherent, on-die data path across the cores, shared cache, memory controllers, and I/O, with up to 3.4 TB/s of bisection bandwidth and a 164 MB unified L3 cache. This gives active cores high-bandwidth, predictable access to shared data as workloads scale across the processor. The SOCAMM2 LPDDR5X memory subsystem complements the fabric with up to 1.2 TB/s of aggregate memory bandwidth, or up to 14 GB/s per core, helping keep the NVIDIA Olympus cores supplied across bandwidth-intensive and highly concurrent workloads. Modular, field-replaceable memory combines LPDDR5X power efficiency with the serviceability and reliability required for datacenter infrastructure.
Storage primitives place two distinct demands on a CPU. First, within each data stream, encryption, integrity checking, recovery, compression, and decompression must complete quickly before subsequent storage processing can proceed, making sustained per-core performance important. Second, across the system, these operations run over many concurrent streams and repeatedly move data through caches and memory, making bandwidth and predictable latency equally important.
Vera addresses both requirements. The Olympus core combines wide instruction throughput, advanced branch prediction, deep out-of-order execution, and vector and cryptographic resources to help each core sustain instruction throughput across control-heavy and data-processing code.
NVIDIA Spatial Multithreading, the monolithic compute die, SCF, unified L3 cache, and high-bandwidth SOCAMM2 memory help keep active cores supplied with data while reducing thread-to-thread interference and supporting more predictable data access under load. Together, these capabilities help explain the measured gains across encryption, integrity checking, parity calculations, compression, and the multi-stage storage pipeline.
This enables the BlueField-4 STX Storage Processor to sustain more CPU-side storage processing across concurrent data streams without proportional increases in CPU resources, power, and cooling.
Measuring foundational storage performance
Storage tasks execute repeatedly across storage read, write, and recovery paths. Their throughput and efficiency help determine whether CPU-side processing keeps pace with SSDs and networks or becomes the limiting stage in the data path. The storage primitive microbenchmarks in this post isolate these functions to measure the processor’s contribution. They show the CPU performance and headroom available in Vera for building higher-throughput, more efficient storage services.
Each test runs within a single process using data already held in memory. The tests exclude file I/O, disk performance, networking, command startup, and external-device bottlenecks unless otherwise identified. The benchmark set uses common libraries, including OpenSSL, Zstandard, and LZ4, along with comparable implementations optimized to use the native instructions available on Arm and x86 processors.
A purpose-built test framework runs each workload consistently and controls buffer sizes, thread counts, CPU placement, timing, correctness validation, and result collection. The algorithms and many of the software implementations are widely adopted. The test framework applies the same workload definitions and controls across Vera and x86, enabling a consistent processor comparison. The results can be reproduced using the source code, scripts, fixed software versions, configurations, and result files, although the complete benchmark set is not an off-the-shelf public benchmark.
These measurements establish Vera’s performance on the storage building blocks that influence secure data movement, resilience, capacity efficiency, and service density. Production storage paths often apply several of these operations to the same data, causing their CPU-processing requirements to accumulate.
When combined in production storage software, the performance of these individual operations contributes to aggregate throughput and CPU efficiency. Higher throughput across individual primitives and the multi-stage pipeline gives storage software more CPU headroom to keep pace with SSDs and networks, support concurrent data flows, and apply essential data services efficiently. End-to-end testing is still required to quantify complete storage-system or GPU-performance outcomes.
Securing more AI data with faster encryption
AI factories process sensitive information, including model assets, enterprise knowledge, agent context, prompts, outputs, and customer data. AES-128 is widely used for data-at-rest and data-in-flight encryption. Encryption sits directly in the write path, where its throughput can determine how much secure data a platform can process each second before encryption becomes a bottleneck. Vera delivers up to 1.43x higher AES-128 encryption throughput than the x86 CPU used for comparison.

Decryption performs the corresponding operation on the read path andVera delivers up to 1.29x higher AES-128 decryption throughput than the x86 CPU used for comparison.

Higher encryption throughput enables storage systems to secure more data without constraining writes, while faster decryption reduces the time required to return protected data to agents, applications, or accelerators. These help storage systems protect and return growing volumes of AI data without consuming an increasing share of the storage-performance budget.
Protecting and recovering AI data faster
Storage platforms use erasure coding to protect data when drives, nodes, or data fragments become unavailable. Reed-Solomon encoding creates redundancy, while recovery uses that redundancy to reconstruct missing or corrupted data. Encoding typically occurs on the write path. Recovery occurs during rebuild, repair, or a degraded read. Their performance results can differ because these operations use different compute and memory-access patterns.
Vera delivers up to 3.26x higher Reed-Solomon throughput than the x86 CPU in a recovery workload.

Higher Reed-Solomon throughput enables storage systems to write protected data and reconstruct missing data faster. In selected efficiency measurements, Vera also completes more protection work within the available CPU power envelope, helping shorten rebuilds and reduce contention with normal data services.
Validating data integrity at higher throughput
Data must remain correct as it is moved, stored, and retrieved. Cyclic redundancy checks, or CRCs, create checksums that storage systems use to detect accidental corruption.
CRC and Reed-Solomon perform complementary roles. CRC detects that data no longer matches the expected result. Reed-Solomon provides the redundancy used to reconstruct missing or corrupted information.
A storage system may calculate CRCs on both the write and read paths, including while copying data between buffers. As data volumes grow, these checks can consume a meaningful share of CPU resources.
Vera delivers up to 3.67x higher CRC32C throughput than the x86 CPU.

Higher CRC32C throughput enables storage systems to validate more data without integrity checking limiting reads or writes. For agentic workloads, this helps return reliable context, persistent memory, and enterprise data with less CPU-side processing delay.
Reducing data footprint with faster compression and decompression
Agentic AI creates growing volumes of context, logs, checkpoints, retrieval data, intermediate outputs, and persistent memory. Compression reduces the storage capacity required for this data and the bandwidth needed to move it. Decompression restores the data when it is read. Compression may run inline or after data is written, depending on the storage architecture, while decompression is typically required when compressed data is read. Compression throughput affects how quickly data can be reduced, and decompression throughput can affect how quickly storage systems return data to agentic AI applications. The following benchmarks measure these operations independently.
Vera delivers up to 3.29x higher compression throughput than the x86 CPU, while sustaining its advantage across the measured thread counts.

The decompression benchmark measures the corresponding operation when compressed data is read. Vera delivers up to 1.72x higher decompression throughput than the x86 CPU under concurrency, with its advantage increasing as more worker threads run in parallel.

Compression performance varies by algorithm, data characteristics, compression level, buffer size, and thread count, so these results apply specifically to the measured workloads. Higher compression throughput enables storage systems to reduce the amount of data written, stored, and transferred while sustaining CPU-side processing performance.
Higher decompression throughput helps storage systems return decompressed data to agents and applications more quickly. Together, these capabilities lower pressure on storage capacity and bandwidth while enabling each processor to compress and decompress more data as agentic AI workloads grow.
Accelerating a multi-stage storage write path
Storage systems rarely execute data services independently. A secure write path may compress data to reduce its footprint and then encrypt it before it is written. The benchmark set includes a memory-resident pipeline that applies compression followed by encryption to each data buffer. Unlike the preceding benchmarks, which measure individual operations, this test measures total pipeline throughput when two CPU-intensive storage functions execute in sequence.
Vera delivers up to 3.21x higher pipeline throughput than the x86 CPU used in the benchmark for the two-stage compression and encryption pipeline.

This result shows that Vera’s performance advantage extends beyond the individual operations measured earlier to a multi-stage sequence representative of a storage write path that reduces and protects data.* *Higher multi-stage pipeline performance enables the storage system to process more data per processor, helping sustain write throughput as agentic AI data volumes grow.
Scaling agent execution and storage with Vera
Agentic AI makes CPU execution and storage processing part of the same AI factory data path.
The CPU runs tools, code, retrieval, analysis, and data-processing steps between model calls. The storage system secures, protects, validates, compresses, and returns the data those steps require. Both must scale within the finite CPU, power, and cooling resources available across the AI factory.
Vera was designed to accelerate CPU-dependent work for the agentic AI era. In NVIDIA Vera Rubin, it serves as the host CPU for NVIDIA GPUs and supports agent execution. Standalone Vera delivers up to 1.8x higher performance per core in agentic tools. In BlueField-4 STX, Vera powers the CPU-side processing used by AI-native storage platforms.
The benchmark results show Vera accelerating encryption and decryption for secure data access, Reed-Solomon recovery for faster reconstruction of missing or corrupted data during storage rebuilds and repairs, CRC32C for high-throughput integrity validation, and compression and decompression to reduce capacity and bandwidth demands and accelerate data retrieval.
By sustaining these functions individually and within a multi-stage write path, Vera helps AI-native storage platforms process and return secure, reliable data with less CPU-side delay. It also supports more concurrent data flows and higher service density without proportional growth in CPU resources, power, and cooling. Across selected workloads, Vera also delivers higher measured performance per watt, enabling more CPU-side storage processing within the available processor power budget.
A common Vera CPU architecture and software toolchain across compute and storage provides a consistent foundation for scaling both agent execution and the supporting data infrastructure.
Learn more about Vera CPUs and BlueField-4 STX
- Read NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI to learn more about the architecture powering these storage-processing results.
- Download the NVIDIA Vera BlueField-4 STX Storage Processor datasheet and learn how NVIDIA BlueField infrastructure processors accelerate storage, networking, and security for AI factories.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み