Dropbox AI エンジニアリング、AI 需要増大に伴うインフラ効率化の取り組みを報告
本文の状態
日本語全文を表示中
詳細モードで約17分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Dropbox AI Engineering
Dropbox は AI 需要の増大に対応するため、エネルギーや冷却などの制約下で既存インフラから最大限の効率を引き出すシステム全体最適化アプローチを推進している。
AI深層分析を開く2026年8月19日 04:35
AI深層分析
キーポイント
既存インフラの最適化重視
新設への依存だけでなく、エネルギー、冷却、ハードウェア利用可能性などの制約下で既存設備から最大限の価値を引き出すことが重要であるとしている。
システム全体でのトレードオフ管理
ストレージ密度の向上や高性能サーバー導入など、一層の変更が他層に与える影響を考慮し、インフラを単体ではなく統合システムとして最適化する。
長期的な需要予測と計画
顧客のアップロードや AI 機能の利用といった需要変化を数ヶ月から数年単位で予測し、リソースの追加時期と配置場所を事前に決定する。
ハイブリッドモデルによる可視化
自社ストレージシステム「Magic Pocket」と専門プロバイダが運営するコロケーションデータセンターを組み合わせることで、ソフトウェアから物理環境までの全体像を把握する。
AI 時代のインフラ計画とトレードオフの理解
AI 製品の成長に伴い、計算能力やストレージの増加がエネルギーや冷却要件を変化させるため、施設制約を考慮した早期のトレードオフ理解が必要である。
重要な引用
Engineering teams are also working within constraints on energy, cooling, hardware availability, and physical space
Rather than treating these as independent problems, we optimize them as parts of a single system
Understanding those tradeoffs allows us to make infrastructure decisions with the entire system in mind
The goal is to understand those tradeoffs early, add capacity deliberately, and preserve enough headroom for growth, maintenance, failures, and changing workloads.
編集コメントを表示
編集コメント
AI 需要の急増に伴い、ハードウェアの新設だけでなく既存リソースの効率化が競争力の源泉となっている。Dropbox の事例は、物理的な制約下でシステム全体を最適化する実践的なアプローチを示している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI への需要が拡大するにつれ、それを支えるインフラの必要性も高まっています。業界の注目は往々にして「より多くのもの」を構築することに集中しがちです。つまり、データセンター、サーバー、電力供給を増やすことに焦点が当たります。しかし、新たなインフラを整備することだけが課題のすべてではありません。
エンジニアリングチームは、エネルギー、冷却、ハードウェアの入手可能性、物理的なスペースといった制約の中で活動しています。そのため、すでに存在するインフラからいかに多くの価値を引き出すかが、ますます重要になっています。
過去 10 年以上にわたり、Dropbox のインフラおよびデータセンターエンジニアリングチームは、インフラの計画・運用・最適化の方法を継続的に改善してきました。この取り組みはストレージシステムに限ったものではありません。Dropbox のエンジニアたちは、キャパシティプランニング、フリート最適化、ハードウェアライフサイクル管理、電力供給、冷却、ラック設計、施設計画など、多岐にわたる分野で連携しています。
これらを独立した問題として扱うのではなく、一つのシステムの一部として最適化しています。ある層での意思決定が、他の層で何が可能になるかに影響を与えるからです。
その結果生み出されたのは、特定のシステムや最適化手法を超えたエンジニアリングの知見です。インフラのある一部を変更すれば、別の場所で新たな機会が生まれる一方で、制約が生じることもあります。例えば、ストレージ密度を高めることで必要なハードウェア量を減らせる一方、より高性能なサーバーを導入すると、エネルギーと冷却に関する新たな要件が発生します。
これらのトレードオフを理解することで、システム全体を見据えたインフラの意思決定が可能になります。
インフラ需要の増大に伴い、システムレベルのアプローチが効率性の優先と、データセンターの物理的拡張前に成長のための余地を創出する上で役立っています。
効率的なインフラの計画
インフラの効率性を形作る多くの決定は、新しいキャパシティが稼働する数ヶ月前、あるいは数年も前から行われます。Dropbox にファイルをアップロードしたり、Dash に対して質問をしたりする際、ユーザーは即座にレスポンスを期待します。そのような体験を提供するには、エンジニアが顧客の需要やワークロードの変動を予測し、追加リソースが必要となる時期を見極め、既存環境がそれを支えられるかを理解する必要があります。
Dropbox は過去 10 年以上にわたり大規模なインフラを運用してきました。当社のハイブリッドモデルは、コアストレージシステムである Magic Pocket と、専門プロバイダーが運営する施設内に設置されたコロケーションデータセンターを組み合わせたものです。これにより、エンジニアはソフトウェア、ハードウェア、そして物理的なデータセンター環境全体にわたる可視性を確保できます。
顧客の需要が高まり、特に AI 搭載製品や機能の普及に伴ってワークロードが変化する中、エンジニアは単に必要な追加容量だけでなく、それをどこに、どのように展開するかを計画する必要があります。新しいハードウェアは、施設のインフラ制約内に収まるとともに、ハードウェアの入手可能性と将来の製品ニーズも考慮して導入されなければなりません。例えば、より多くの計算能力やストレージを提供するサーバーは、エネルギー消費や冷却要件を増大させ、ラックや施設が支えられるハードウェアの総量に影響を及ぼします。
目指すべきは、こうしたトレードオフを早期に把握し、計画的に容量を追加するとともに、成長・保守・障害対応・ワークロードの変化に対応できる十分な余裕(ヘッドルーム)を確保することです。追加の計算能力とストレージが展開された後は、焦点はすでに稼働中のシステムを最大限活用することに移ります。
稼働中のファームを継続的に最適化する
計画により、インフラが予測される需要に備えられるようにすることはできますが、実際のワークロードが当初の想定通りに振る舞うことは稀です。顧客の行動は変化し、製品は進化し、新機能によって基盤システムへの新たな要求が生じます。AI はその典型例であり、AI 搭載機能はインフラ需要の規模と形状の両方を変化させ、計算・ストレージ・メモリ・ネットワークに対して新たな負荷を課します。
つまり、効率化は継続的なエンジニアリング課題です。展開されたインフラを固定のものとして扱うのではなく、Dropbox は需要の変化に応じてアクティブなファームウェアの運用方法を絶えず適応させています。時には需要が低い時期に使用されるハードウェア量を減らすこともあれば、リソースが利用可能なファームの一部へ作業をシフトさせることもあります。また、ハードウェアの進歩により、同じ物理インフラで大幅に多くのストレージをサポートできるようになるケースもあります。
これらのアプローチはシステムの異なる層で機能しますが、共通する目的は、新たなインフラを追加する前に、すでに存在するインフラからより有用なキャパシティを引き出すことです。
必要がない時にキャパシティを休ませる
インフラは需要が高まる期間に対応するために用意され、信頼性とメンテナンスのための余剰分も確保されています。つまり、すべての利用可能なキャパシティが常に必要なわけではありません。ハードウェアがアイドル状態にある場合や過剰なキャパシティが存在する場合、すべてのコンポーネントを完全に稼働させたままにすると、エネルギーを消費するだけで追加の価値は生まれません。Dropbox はこのオーバーヘッドを削減するために「Deep Sleep」という手法を採用しています。
Deep Sleep は、ハードウェアが実際に必要とされていない際に電力消費を削減する Dropbox のインフラ施策です。ハードウェアの種類によっては、ハードディスクをスタンバイモードにスピンダウンさせたり、アイドル状態のサーバー自体を完全に電源オフしたりします。
サーバーが Deep Sleep 対象となるかは、自動化されたフリート管理アルゴリズムによって判断されます。Dropbox は、サーバーが数分以内にサービス復帰できるため、エネルギー効率とパフォーマンス、信頼性のバランスを保つことが可能です。低レイテンシが求められるワークロードの場合には、サーバー全体を電源オフするのではなく、アイドル状態のハードディスクのみを選択的にスピンダウンさせることで対応します。
課題は、これらの省電力措置をどこで安全に適用できるかを判断することです。エンジニアは、運用と信頼性の要件を満たすために十分な利用可能容量を確保しつつ、完全な給電を維持する必要がないハードウェアを特定する必要があります。これにより、Dropbox は製品が依存する容量を損なうことなく、未活用インフラのエネルギー消費を削減できます。
アイドル状態のハードウェアによる電力消費を減らすことは、既存フリートの効率化の一つの方法です。もう一つの手法は、すでに稼働中のインフラをより効果的に活用することです。
フリート全体でのワークロードの最適配分
インフラの効率運用において、需要全体を満たす十分なキャパシティを確保することは、重要な要素の一つに過ぎません。重要なのは、そのキャパシティが実際に作業が行われている場所で利用可能であることです。
fleet(機体群)の一部はすでに限界に近づいている一方で、別の部分にはまだ余裕がある場合があります。この不均衡に対処する手段がなければ、Dropbox は既存のリソースを有効活用する代わりに、新たなインフラを追加してしまうリスクがあります。
これを防ぐため、Dropbox ではインフラ全体におけるリソースの使用状況を常時監視しています。チームは、fleet 全体の予備キャパシティとワークロードの分布をモニタリングすることで不均衡を検出し、利用可能な余力が減少している領域や、作業が偏って集中している箇所を特定します。
需要に偏りがある場合、チームはシステム間でワークロードを再バランスしたり、必要な場所に追加キャパシティをオンライン化したりできます。一部の変更は自動的に行われますが、大規模な変更についてはエンジニアによるレビューと検証を経て実施されます。これにより、局所的な制約を全体のインフラ増設の必要性と捉えるのではなく、他の場所で利用可能なリソースを活用することが可能になります。
fleet の構成は常に変化するため、このバランス調整も継続して行う必要があります。顧客の行動パターンが変化し、新製品によって新たなワークロードパターンが生まれる一方で、インフラ自体も進化を遂げます。どこで作業を実行するかを絶えず適応させることで、Dropbox はすでに稼働中のキャパシティからより多くの価値を引き出すことができます。
ワークロードのバランス調整は、既存の容量をより有効活用する手段ですが、効率化にはインフラ自体が支えられる能力を高めるアプローチも重要です。
ストレージ密度の向上
既存のインフラからより多くの価値を引き出すもう一つの手段は、ハードウェア1 台あたりのサポート能力を高めることです。ストレージ技術の進歩により、Dropbox は各ドライブに格納できる顧客データを大幅に増やすことに成功しました。その一例がシャングルド・マグネティック・レコーディング(Shingled Magnetic Recording)です。この技術は物理サイズを増やさずに、ハードドライブ上にデータをより高密度に記録します。
こうした成果はインフラ全体で相乗効果を生みます。各ドライブの容量が増えれば、同じ量のストレージを提供するために必要なドライブ数は減ります。その結果、サーバーやラックの数を削減でき、配線も少なくて済むようになります。さらに、電力消費量と冷却要件も低下します。つまり、単一コンポーネントの容量を高めることが、展開全体の資源必要量を削減することにつながるのです。
このシステムレベルでの影響こそが、総電力消費量だけでは効率性の全貌を語れない理由です。Dropbox が成長し顧客データをより多く格納するにつれ、インフラの効率が向上しても、エネルギー使用量は増加する可能性があります。より有用な指標は「ペタバイトあたりのワット数」、つまり 1 ペタバイトのストレージを支えるために必要な電力量です。2020 年以降、当社のストレージインフラにおけるペタバイトあたりの消費電力は 50% 以上改善されました。現在では、同じ量のストレージをサポートするために必要な電力は、2020 年の半分以下となっています。
上記のすべてのアプローチを組み合わせることで、Dropbox は不要な電力消費を抑え、利用可能なキャパシティを最大限に活用し、ハードウェア 1 台あたりの処理能力を引き上げることで、既存インフラからより多くの価値を生み出しています。ただし、そのハードウェアがどれだけの期間、安定して稼働し続けられるかも重要な要素です。
インフラのライフサイクル延長について
機器を早すぎると有用なキャパシティが余ってしまいますし、逆に長すぎると信頼性やパフォーマンスにリスクが生じます。ハードウェアは単に一定の年齢に達したからといって不具合を起こすわけではなく、長く使い続けることが常に最適とは限りません。こうした判断を下すため、Dropbox は本番環境におけるハードウェアのパフォーマンスを継続的に監視しています。年間故障率などの指標を用いることで、エンジニアは異なる部品や世代ごとの機器が時間経過とともにどう振る舞うかを把握できます。
インフラへの需要が増え続ける中、こうしたライフサイクルに関する判断の重要性はさらに高まっています。既存インフラからより多くの価値を引き出すことは、稼働中のハードウェアの効率性だけでなく、追加投資が必要になるまでどの程度安定して動作し続けられるかという理解にも依存しています。そのデータに基づき、機器を継続使用するか、修理すべきか、それとも交換が必要かを判断します。
ハードウェアのパフォーマンスを包括的に把握することで、機器が安定して動作し続ける限り、固定された交換スケジュールに頼るのではなく、その有用な寿命を延ばすことが可能になります。パフォーマンスの低下や、新しいハードウェアが容量・信頼性・効率において実質的な改善をもたらす場合、移行計画を立てることができます。ただし、何よりも重要なのは信頼性です。機器が運用上の基準を満たし続ける場合にのみ、ハードウェアのライフサイクル延長は価値を持ちます。
Dropbox では、可能な限りハードウェアを修理してその有用な寿命を延ばしています。もはやサービス継続が困難になった機器については、信頼できるパートナーと連携して再販売または責任あるリサイクルを行います。これらの取り組みを通じて、導入や交換といった特定の時点だけでなく、機器の全ライフサイクルを通じて価値を最大化することを目指しています。しかし最終的には、新しいハードウェアを導入せざるを得ません。サーバーがより強力になりストレージ密度が高まるにつれて、それらを支える物理インフラには新たな制約が生じる可能性があります。
ハードウェアを超えたエンジニアリング
新しいハードウェアの導入は、単にサーバーを交換するだけではありません。本番環境で稼働させる前に、その機器が動作できる物理的環境を整備する必要があります。チームは、各導入に必要な電力・空気流・ラックレイアウト・配線などの物理要件を計画し、コロケーションプロバイダーと密接に連携しながら進めています。
これらの判断は密接に関連しています。より多くのデータを保存したり、より高い計算能力を提供するサーバーほど、消費電力や発熱が増大します。その結果、ラックの設計方法や機器の冷却方式、さらにはデータセンター内の特定エリアでサポート可能なハードウェアの量まで影響を受けます。
インフラがより高性能かつ高密度化する中で、ハードウェアから最大限の性能を引き出すためには、周囲の環境も同時に進化させることが不可欠です。
最近の事例は、こうしたトレードオフが実際の運用でどのように現れるかを如実に示しています。Dropbox が第 7 世代サーバーの導入を進めた際、必要な電力が増大し、既存のラック電源設計の容量を超えてしまいました。施設全体のインフラを再構築するのではなく、ハードウェアエンジニアリングチームとデータセンターエンジニアリングチームはラック内の電源アーキテクチャを見直し、既存のバスウェイを使い続けながら、1 ラックあたりの電源分配ユニット(PDU)の数を倍増させました。
その結果、データセンター自体に大規模な改修を施すことなく、新しいハードウェアに対応することができました。
これは、インフラの改善はハードウェアそのものだけで完結するものではないという教訓です。需要が高まり、それを満たすためにハードウェアが進化するにつれて、次世代の機器は常に周囲の環境が持つ制約の中で収まるように設計されなければなりません。
長年の運用経験から得られたインフラに関する知見
この取り組みは、Dropbox のインフラに対する 10 年以上にわたるエンジニアリング投資の集約です。キャパシティプランニング、フリート最適化、ストレージシステム、ハードウェアライフサイクル管理、データセンター工学など、各領域が異なる課題に取り組んでいますが、これらを統合することで、製品を支えるインフラをより効果的に活用できるようになっています。
これらの取り組みは一度きりの作業ではありません。顧客のニーズの変化やハードウェアの進化、新技術による新たな機会と制約の登場に応じて、私たちはインフラを継続的に改善しています。各アップデートは、それまでの成果の上に積み重ねられていきます。
デジタルインフラへの需要がさらに高まる中、これらの教訓はこれまで以上に重要になっています。AI の台頭により業界全体でストレージとコンピューティングの必要性が加速していますが、根本的なエンジニアリングの課題は変わっていません。インフラは依然として、信頼性、効率性、回復力を保ちながらスケールする必要があります。技術が進化する中で、私たちは顧客が現在依存している製品を支えるシステムを構築し続ける一方で、将来のニーズにも柔軟に対応できる基盤も同時に整えていきます。
~ ~ ~
革新的な製品や体験、インフラの構築に情熱を持っている方、一緒に未来を築きませんか?dropbox.jobs で募集ポジションをご覧ください。
原文を表示
As demand for AI continues to grow, so does the infrastructure needed to support it. Much of the industry's attention has focused on building more: more data centers, more servers, and more power. But building new infrastructure is only part of the challenge. Engineering teams are also working within constraints on energy, cooling, hardware availability, and physical space, making it increasingly important to get more from the infrastructure that's already in place.
For over a decade, Dropbox’s Infrastructure and Datacenter Engineering teams have continually improved how we plan, operate, and optimize infrastructure. That work spans far more than storage systems. Engineers across Dropbox work together on capacity planning, fleet optimization, hardware lifecycle management, power delivery, cooling, rack design, and facility planning. Rather than treating these as independent problems, we optimize them as parts of a single system, where decisions in one layer influence what's possible in another.
The result is an engineering discipline that extends beyond any single system or optimization. A change in one part of the infrastructure can create opportunities or constraints elsewhere. Increasing storage density, for example, can reduce the amount of hardware needed, while more powerful servers can introduce new energy and cooling requirements. Understanding those tradeoffs allows us to make infrastructure decisions with the entire system in mind.
As infrastructure demand grows, that system-level approach helps us prioritize efficiency and create room for growth before expanding our data center footprint.
Planning for efficient infrastructure
Many of the decisions that shape infrastructure efficiency happen months or sometimes years before new capacity goes into production. Whether someone is uploading a file to Dropbox or asking Dash a question, they expect the product to respond without delay. Delivering that experience requires engineers to forecast how customer demand and workloads will change, determine when additional resources will be needed, and understand whether our existing environments can support them.
Dropbox has operated large-scale infrastructure for over a decade. Our hybrid model combines Magic Pocket, the core Dropbox storage system, with colocated data centers, where we manage our own servers and networking equipment in facilities operated by specialized providers. This gives our engineers visibility across software, hardware, and the physical data center environment.
As customer demand grows and workloads change, particularly with the growth of AI-powered products and features, engineers have to plan not only for how much additional capacity is needed, but where and how it can be deployed. New hardware has to fit within a facility's infrastructure constraints while accounting for hardware availability and future product needs. A server that provides more compute or storage, for example, may also require more energy or cooling, changing how much hardware a rack or facility can support.
The goal is to understand those tradeoffs early, add capacity deliberately, and preserve enough headroom for growth, maintenance, failures, and changing workloads. Once additional compute and storage capacity is deployed, the focus shifts to making the most of the systems already in production.
Continuously optimizing the active fleet
Planning helps ensure infrastructure is ready for anticipated demand, but workloads rarely behave exactly as they did when that infrastructure was first deployed. Customer behavior changes, products evolve, and new capabilities introduce different demands on the underlying systems. AI is a prime example because AI-powered features can change both the scale and shape of infrastructure demand, placing new demands on compute, storage, memory, and networking.
That makes efficiency an ongoing engineering problem. Rather than treating deployed infrastructure as fixed, Dropbox continually adapts how the active fleet operates as those demands change. Sometimes that means reducing how much hardware is in use when demand is lower. Sometimes it means shifting work to parts of the fleet with resources available. And in other instances, advances in hardware allow the same physical infrastructure to support substantially more storage.
These approaches work at different layers of the system, but they share the same objective: getting more useful capacity from the infrastructure already in place before adding more of it.
Letting capacity rest when it isn't needed
Infrastructure has to be provisioned for periods of higher demand, with additional headroom built in for reliability and maintenance. That means not all available capacity is needed at all times. When hardware is sitting idle or excess capacity is available, keeping every component fully powered consumes energy without providing additional value. Deep Sleep is one way Dropbox reduces that overhead.
Deep Sleep is a Dropbox infrastructure initiative that reduces power consumption when hardware isn't actively needed. Depending on the hardware, that can mean spinning down hard drives into standby mode or powering down idle servers altogether. A server becomes eligible for Deep Sleep through automated fleet management algorithms. Dropbox is able to balance energy efficiency with performance and reliability because servers can return to service within minutes. For workloads that require lower latency, we can also selectively spin down idle hard drives rather than powering down the entire server.
The challenge is determining where those power-saving measures can be applied safely. Engineers have to preserve enough available capacity to meet operational and reliability requirements while identifying hardware that doesn't need to remain fully powered. That allows Dropbox to reduce the energy consumed by underused infrastructure without compromising the capacity our products depend on.
Reducing the power consumed by idle hardware is one way to make the existing fleet more efficient. Another is making better use of the infrastructure that's already online.
Balancing work across the fleet
Having enough capacity to meet overall demand is only part of operating infrastructure efficiently. That capacity also needs to be available where the work is happening. One part of the fleet may be approaching its limits while another has room to take on more work. Without a way to address that imbalance, Dropbox could end up adding more infrastructure instead of making better use of what’s already available across the fleet.
To avoid that, Dropbox continually monitors how workloads are using resources across our infrastructure. The team identifies imbalances by monitoring spare capacity and workload distribution across the fleet, looking for areas where available headroom is falling or work is concentrating unevenly. When demand is uneven, teams can rebalance workloads across systems or bring additional capacity online where it's needed. Some adjustments happen automatically, while larger changes are reviewed and validated by engineers. This allows us to take advantage of available resources elsewhere rather than treating a localized constraint as a need for more infrastructure overall.
That balancing has to continue as the fleet changes. Customer behavior shifts, products introduce new workload patterns, and the infrastructure itself evolves. Continually adapting where work runs helps Dropbox get more from the capacity that's already online.
Balancing workloads helps make better use of the capacity we already have, but efficiency can also come from increasing how much that infrastructure can support in the first place.
Increasing storage density
Another way to get more from existing infrastructure is to increase how much each piece of hardware can support. Advances in storage technology have allowed Dropbox to store significantly more customer data on each drive. One example is shingled magnetic recording, which packs data more densely onto a hard drive without increasing its physical size.
Those gains compound across the infrastructure. When each drive holds more data, fewer drives are needed to provide the same amount of storage. That can mean fewer servers and racks, less cabling, and lower power and cooling requirements. In turn, increasing the capacity of a single component can reduce the resources required across an entire deployment.
That system-level impact is also why total power consumption doesn't tell the full story of efficiency. As Dropbox grows and stores more customer data, overall energy use may increase even as the infrastructure becomes more efficient. A more useful measure is watts per petabyte, or the amount of power required to support a petabyte of storage. Since 2020, watts per petabyte across our storage infrastructure have improved by more than 50%. Today, it takes less than half as much power to support the same amount of storage as it did in 2020.
Together, all of the approaches described above help Dropbox get more from existing infrastructure by reducing unnecessary power consumption, making better use of available capacity and increasing how much each piece of hardware can support. How long that hardware can reliably remain in service matters, too.
Extending infrastructure over time
Replacing equipment too early can leave useful capacity on the table, while keeping it too long can introduce reliability and performance risks. Hardware doesn’t become unreliable simply because it reaches a particular age, nor does keeping equipment longer always make sense. To make those decisions, Dropbox monitors how hardware performs in production. Metrics such as annual failure rate help engineers understand how different components and generations of equipment behave over time.
As demand for infrastructure continues to grow, those lifecycle decisions become increasingly important. Getting more from existing infrastructure isn't only about how efficiently hardware operates while it's in service. It's also about understanding how long that hardware can continue operating reliably before additional investment is needed. That data informs whether hardware can remain in service, should be repaired, or needs to be replaced.
Understanding hardware performance holistically allows us to extend the useful life of equipment when it continues to perform reliably instead of relying solely on a fixed replacement schedule. When performance declines or newer hardware provides meaningful improvements in capacity, reliability, or efficiency, we can plan a transition. Reliability comes first. Extending a hardware lifecycle is valuable only when equipment continues to meet our operational standards.
We repair hardware whenever practical to extend its useful life. When equipment can no longer remain in service at Dropbox, we work with trusted partners to resell or responsibly recycle it. Together, these decisions help maximize the value of equipment throughout its lifecycle rather than treating deployment and replacement as the only meaningful milestones. But eventually, new hardware does need to come online. And as servers become more powerful and storage becomes denser, deploying them can introduce a new set of constraints in the physical infrastructure that supports them.
Engineering beyond the hardware
Deploying new hardware isn't as simple as swapping one server for another. Before new hardware can go into production, the physical environment has to be able to support it. Our team plans for the power, airflow, rack layout, cabling, and other physical requirements each deployment needs, working closely with our colocation providers along the way.
Those decisions are closely connected. A server that stores more data or delivers more compute may also draw more power or generate more heat. That can change how racks are designed, how equipment is cooled, and even how much hardware a particular area of a data center can support. As infrastructure becomes more powerful and denser, getting more from the hardware depends on making sure the environment around it can evolve, too.
One recent example illustrates how those tradeoffs play out in practice. As Dropbox deployed its seventh-generation servers, the increased power requirements exceeded the capacity of the existing rack power design. Rather than rebuilding the underlying facility infrastructure, the Hardware Engineering and Datacenter Engineering teams redesigned the rack power architecture, doubling the number of power distribution units per rack while continuing to use the existing busways. The result supported the new hardware without requiring major changes to the data center itself.
It's a reminder that infrastructure improvements don't stop with the hardware itself. As demand grows and hardware evolves to meet it, each new generation has to fit within the constraints of the environment around it.
What years of operating infrastructure have taught us
The work reflects more than a decade of engineering investment across Dropbox's infrastructure. Capacity planning, fleet optimization, storage systems, hardware lifecycle management, and data center engineering each address different challenges, but together they help us make better use of the infrastructure that supports our products.
None of this work is one and done. As customer demands change, hardware evolves, and new technologies introduce opportunities and new constraints, we improve our infrastructure, each update building on the ones that came before it.
Those lessons have become even more relevant as demand for digital infrastructure continues to grow. AI is accelerating the need for storage and compute across the industry, but the underlying engineering challenge hasn't changed. Infrastructure still has to scale while remaining reliable, efficient, and resilient. As these technologies evolve, we'll continue building systems that support the products our customers rely on today while giving us the flexibility to support what's next.
~ ~ ~
If building innovative products, experiences, and infrastructure excites you, come build the future with us! Visit dropbox.jobs to see our open roles.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み