Google、今月の AI インフラとオーケストレーションの最新動向を公開
本文の状態
日本語全文を表示中
詳細モードで約10分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Google Cloud AI
Google Cloud は AI エージェント時代のインフラ要件に対応するため、DDN と連携した高性能ストレージ「Managed Lustre」の一般提供を開始し、最大 8PB の容量と多段階のスケーラビリティを実現した。
AI深層分析を開く2026年8月4日 06:41
AI深層分析
キーポイント
Managed Lustre の一般提供開始
Google Cloud Managed Lustre が GA となり、4 つの異なるパフォーマンスティア(125MB/s〜1000MB/s)を提供している。
DDN EXAScaler との連携
本ソリューションは DDN の EXAScaler を基盤としており、同社の長年のハイパフォーマンスストレージ技術と Google Cloud のクラウドインフラ専門知識を組み合わせている。
大規模スケーラビリティの実現
単一システムで最大 8PB のストレージ容量まで拡張可能であり、AI エージェントや大規模トレーニングにおけるデータ処理のボトルネック解消を目指す。
C4Nネットワークおよびストレージ最適化VMの一般提供
C4Nはデータ転送ボトルネックを解消するために設計された最初のVMシリーズであり、Titaniumオフロードハードウェアと5世代Intel Xeonプロセッサを搭載している。
GKE Dataplane V2の最大15,000ノードへのスケーリング対応
標準的なGKEクラスターがアクティブなネットワークポリシーを維持しながら最大15,000ノードまでスケール可能となり、大規模エンタープライズやAI/ML顧客のインフラ要件に対応する。
重要な引用
At Google, AI is a soup-to-nuts endeavor.
This is critical in today's agentic era, where AI is evolving from answering questions to reasoning and taking action.
Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity
C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks.
編集コメントを表示
編集コメント
AI エージェントの普及に伴い、従来のストレージ要件を超えた高速かつ大規模なデータ処理能力が求められている。Google と DDN の連携による Managed Lustre の登場は、この課題に対する具体的な解決策として注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Google において、AI は一貫した取り組みです。もちろん、Gemini や Nano Banana など最先端の AI モデルを開発しています。また、Gmail、BigQuery、AlloyDB、Google Cloud Code、Google Cloud Assist など、あなたが毎日使うツールに AI を組み込んでいます。
さらに、Gemini Enterprise Agent Platform、JAX、MaxTest といったソフトウェアフレームワークを提供し、AI を活用した開発を支援します。そして、そのすべてを支える強力なインフラストラクチャプラットフォームも自社で設計しています。これには、幅広い標準コンピューティングリソース、TPU や GPU などのアクセラレーター、最適化されたネットワークとストレージ、GKE や Cluster Director といったオーケストレーションソフトウェアが含まれます。
最後に、これらすべてを AI Hypercomputer などのスーパーコンピューティングプラットフォームにパッケージ化し、業界全体が AI とエージェント指向の未来へ移行するための原動力となっています。
これは、AI が単なる質問への回答から推論や行動実行へと進化している現在の「エージェント時代」において極めて重要です。この次の段階で主導権を握りたい企業は、これらの新たな要件に設計・最適化されたコンピューティングインフラストラクチャが必要です。そうすることで、より迅速なイノベーションが可能になり、魅力的なユーザー体験と顧客体験を提供でき、大規模な運用においてもコストとエネルギー効率の最適化を実現できます。
これらを支えるため、AIインフラとオーケストレーションに関するニュースを急速に発信しています。本ブログでは、最新の動向を把握するために知っておくべき新製品やマイルストーンを毎月まとめるとともに、アーキテクチャやパフォーマンスチューニングの深掘り記事、特定のユースケースに関する議論も掲載します。各記事には、さらに詳しく学ぶためのリンクも必ず記載しています。このブログは毎月更新されるので、ぜひチェックしてください。
2026年7月
プロダクト・技術・ツールのアップデート
- プロダクト更新: Google Cloud Managed Lustre が一般提供(GA)を開始しました。このサービスは、4 つの明確なパフォーマンスティアを提供し、容量 1 TiB あたりのスループットを 125 MB/s、250 MB/s、500 MB/s、1000 MB/s から選択できます。また、ストレージ容量は最大 8 PB までスケール可能です。Managed Lustre は DDN の EXAScaler を基盤としており、DDN が長年築き上げてきたハイパフォーマンスストレージのリーダーシップと、Google Cloud のクラウドインフラにおける専門知識を融合させたソリューションです。
今月の AI インフラとオーケストレーションの最新動向
製品アップデート: C4N ネットワークおよびストレージ最適化 VM が一般提供開始。C4N はデータ転送のボトルネックを解消するために設計された、ネットワークとブロックストレージに特化した最初の VM シリーズです。第 5 世代 Intel Xeon スケーラブルプロセッサを搭載し、Google の Titanium オフロードハードウェア上で動作することで、Hyperdisk Extreme と組み合わせることで 400 Gbps のネットワーク帯域幅、1 秒あたり 9500 万パケット(MPPS)の処理能力、そして最大 25 GiB/s のブロックストレージスループットを実現します。
新機能: GKE Dataplane V2 が一般提供開始。ネットワークポリシーを維持したまま 1.5 万ノードまでスケール可能。この機能により、標準的な GKE クラスターは 15,000 ノードまで拡張が可能になりつつ、アクティブなネットワークポリシーの完全な適用を維持できます。これにより、大規模企業や AI/ML カスタマーが抱える巨大なインフラニーズに対応可能です。
新機能: llm-d における協調型タイムスライシング。強化学習(RL)ワークロードを実行中の方へ、独立した RL ジョブを共有物理ハードウェア上でインターリーブして実行できるようになりました。これにより、モデルの収束や精度への影響を与えずに、アクセラレータの稼働率を約 40% のベースラインから最大 70% まで向上させることが可能です。
- 新しい AI セキュリティツール: GKE で AI サプライチェーンの保護、AI ワークロードの安全なデプロイ、シャドウ AI の削減をお考えですか?私たちは軽量で特権を必要としない Kubernetes コントローラー「k8s-aibom」をオープンソース化しました。このツールはコンテナクラスターを常時監視し、vLLM や Triton などの実行中の AI ランタイムを自動的に検出するとともに、標準的な CycloneDX Machine Learning Bill of Materials(ML-BOM)を生成します。k8s-aibom プロジェクト をチェックして、ぜひご参加ください。
プラクティショナー向けガイドとチュートリアル
- チュートリアル: 7 月 27 日、Google は重み付けが公開された同日に、2.8 兆パラメータを持つオープンウェイトモデル「Moonshot AI の Kimi K3」に対する Day 0 サポートを発表しました。Model Garden を利用するか、カスタムオーケストレーションを行うか、あるいは GKE で llm-d レシピを使用するか、ご自身のデプロイパスに関わらず、本ガイドでは Google Cloud 上で Kimi K3 を評価し、パイロット運用するための詳細な手順を解説しています。
今月の AI インフラとオーケストレーションの最新情報
ハウツーガイド: Google Kubernetes Engine (GKE) の管理型 DRANET が、GPU と TPU の両方をサポートするようになりました。この機能を利用するには、フルコントロールが可能な「標準クラスター」や、Google が設定を自動処理してくれる「Autopilot クラスター」など、いくつかの構成オプションがあります。詳しくは、GKE Autopilot クラスターと TPUs、GKE 管理型 DRANET、そして Gemma 4 のハンズオンラボをご覧ください。
ハウツーガイド: GPU ではなく TPU で Ray を実行する方法を学びましょう。この 2 部構成シリーズの 第 1 回 では、TPU スライスについて解説します(Ray はこれをスケジューリング対象の別のアクセラレータと捉えています)。その後、Ray の各種 AI ライブラリの使い方を 第 2 回 で詳しく紹介します。
ハウツーガイド: 新しいマイクロベンチマークスイートを使って、サンプルワークロードで TPU を評価しましょう。このツールを使えば、デバイスが理論上の性能仕様を達成できているかを正確に判断し、具体的なパフォーマンスの差やアーキテクチャ固有のボトルネックを特定できます。詳しくは こちら をご覧ください。
- 実践ガイド: 予算を圧迫せずにエージェントの規模拡大を実現する方法。GKE Agent Sandbox と Pod スナップショットを活用して、限られた計算リソースに安全により多くのエージェントを収容する方法 を詳しく解説します。パフォーマンス向上かコスト最適化かの目的に応じて、最適なエージェント効率を実現するための設定ポイントを紹介します。
- 技術的解説: Google の Ironwood(TPU v7x)上で Mistral 3 Large モデルの推論を最適化する内幕。本記事では、Google チームがハイブリッドシャーディングの採用や線形 VPU 加算を木構造による集約へ置き換えること、GMM/MLA カーネルの最適化、非同期スケジューリングの実装などを通じて、Mistral 3 Large の MoE モデル推論をどのように最適化したかを詳述しています。その結果、ベンチマーク精度を維持したままスループットを最大 48% 向上させることに成功しました。詳細はこちらをご覧ください。
リサーチ、レポート、深掘り記事
レポート:Google は、初回の「GartnerⓇ AI インフラストラクチャ・マジック・クアドラント」でリーダーに選出されました。この評価では、「実行能力」において最高位、「ビジョンの完全性」においても最も高い位置を占めています。ガートナーは、Google の独自開発によるスケーラブルなコンピューティング資源、統合された AI ハイパーコンピュータアーキテクチャ、そして AI コンピューティング容量の規模を主要な強みとして挙げています。詳細は こちら からダウンロードできます。
レポート:私たちは最近、1,400 名以上のシニア IT リーダーを対象に「アジェンティック AI エラにおける AI インフラストラクチャの現状」に関する調査を行いました。その結果、明確な傾向が浮かび上がりました。AI への野心とインフラの実態との間に、格差が広がっているのです。実際、組織の 83% が、本格的に運用されるアジェンティック AI を支えるためにインフラのアップグレードが必要だと回答しています。アジェンティックアプリケーションがシステムに課す要件に応えるようインフラを適応させることが、パイロット段階から本番環境への移行をどう支援するかについては、関連するブログ記事 をご覧ください。
2026 年 6 月
プロダクト・技術・ツールのアップデート
製品アップデート:AI で使用する機密データの保護は、高度で安全なクラウドインフラにおいて極めて重要な要素です。Confidential Computing は、検証可能なデータ整合性を備えたハードウェアベースの信頼実行環境(TEE)内で、使用中のデータを暗号化して保護します。この機能は now G4 マシンシリーズ で利用可能になりました。同シリーズには、NVIDIA RTX PRO 6000 Blackwell Server Edition GPU が搭載されています。Confidential G4 VM や Confidential G4 GKE ノード を活用して、すぐに使い始めることができます。
開発者向けリソース:モデル構築者、最適化担当者、そして開発者が Google Cloud TPU の性能を最大限に引き出すための学習拠点として、新しい TPU Developer Hub が開設されました。詳しくは、こちらのブログ記事をご覧ください。
新製品:新しい OpenTelemetry ベースの TPU AI テレメトリ収集エージェント で、AI ワークロードのスケーリングを効率化しましょう。これで初めて、高忠実度の TPU ハードウェアのテレメトリデータを Google Cloud Monitoring や Google Managed Prometheus、あるいは自前の Grafana スタックへ転送できるようになります。
実践者向けガイドと手順書
- 手順書: TPUs、Cloud Storage FUSE、Dynamic Resource Allocation (DRA) を活用して GKE Inference Gateway で動作する AI 推論ワークロードに高可用性を構築する方法をご紹介します。概要は こちら のブログ記事で、技術的な詳細は ハンズオン・コドラボ で確認できます。
手順書: AI エージェントを非構造化データに接続できることをご存知ですか?
原文を表示
At Google, AI is a soup-to-nuts endeavor. Obviously, we make leading AI models like Gemini and Nano Banana. We incorporate AI into the tools you use every day (think Gmail, BigQuery, AlloyDB, Google Cloud Code and Google Cloud Assist). We make software frameworks to help you build with AI, like Gemini Enterprise Agent Platform, JAX, or MaxTest. And we co-design the powerful infrastructure platform that runs underneath it all, including a broad range of standard compute, accelerators like TPUs and GPUs, optimized networks and storage, as well as orchestration software like GKE and Cluster Director. Then we package them all up into supercomputing platforms like AI Hypercomputer to power the industry-wide transformation to the AI and agentic future.
This is critical in today’s agentic era, where AI is evolving from answering questions to reasoning and taking action. Companies that want to lead in this next phase of AI need computing infrastructure that’s designed and optimized for these new requirements, so they can innovate faster, deliver compelling user and customer experiences, and optimize for cost and energy efficiency — all at massive scale.
To support this, we are making AI infrastructure and orchestration news at a furious pace. In this blog, we provide a monthly snapshot of the recent launches and milestones that you need to know about to keep up-to-date, deep dives on architecture and performance tuning, and discussions of specialized use cases, always with pointers to where you can learn more. Keep an eye out for updates to this blog every month.
July 2026
Product, technology, and tools updates
- Product update: Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure.
- Product update: C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme.
- New feature: GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers.
- New feature: Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy.
- New AI security tool: Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the k8s-aibom project and get involved.
Practitioner guides and how-tos
- How-to guide: On July 27, Google announced Day 0 support for Moonshot AI’s Kimi K3 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud.
- How-to guide: Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4.
- How-to guide: Learn to run Ray on TPUs, not GPUs. In Part 1 of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (Part 2).
- How-to guide: Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in here.
- How-to guide: Scale your agents without killing your budget. Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency.
- Technical blueprint: Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog here.
Research, reports and deep-dives
- Report: Google was named a Leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy here.
- Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. Read the accompanying blog to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production.
June 2026
Product, technology and tool updates
- Product update: Protecting sensitive data used with AI is a critical part of advanced and secure cloud infrastructure. Confidential Computing cryptographically protects data in use in hardware-based Trusted Execution Environments (TEEs) with verifiable data integrity, and is now available on the accelerator-optimized G4 machine series, featuring NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. Get started with Confidential G4 VMs and Confidential G4 GKE Nodes.
- Developer resource: The new TPU Developer Hub is the place to go for model builders, optimizers, and developers to learn to unlock the full performance of Google Cloud TPUs. Read more in this blog.
- New product: Scale your AI workloads with the new OpenTelemetry-Based TPU AI Telemetry Collector Agent. For the first time, you can route high-fidelity TPU hardware telemetry to Google Cloud Monitoring, Google Managed Prometheus, or your own self-hosted Grafana stack.
Practitioner guides and how-tos
- How-to guide: Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This blog provides an overview, or you can get all the technical details in the hands-on codelab.
How-to guide: Did you know you can connect your AI agents to unstructured data in
AI算出
主要ニュースainew評価標準
AI エージェント時代を支える基盤インフラ(ストレージ、ネットワーク、オーケストレーション)に関する具体的な新製品の一般提供と技術詳細が報じられており、新規性と実用性が高い。ただし、日本企業固有の導入事例や価格・規制情報の記載がないため、日本の関連性は限定的となる。
6つの評価軸を見る
- AI関連度
- 75
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み