Hugging Face、地球規模の地理空間推論プラットフォーム「OlmoEarth」を発表
AI2 は地球規模の地理空間推論を実現する OlmoEarth プラットフォームを発表し、衛星画像処理におけるハードウェア最適化と大規模分散処理の技術的課題解決を提示した。
AI深層分析を開く2026年7月29日 02:51
AI深層分析
キーポイント
衛星推論の技術的課題
地球規模での衛星画像推論は、膨大なデータ量と複雑なインフラ要件により従来の手法では困難であることが指摘されている。
タスクに最適化されたハードウェア選定
特定の処理タスクに対して最適なハードウェアを選定・構成することで、推論効率を最大化するアプローチが採用されている。
大規模分散処理アーキテクチャ
1 つのリクエストに対して数百人のワーカーと数千のプロセスを動員するスケーラブルなシステム設計により、膨大なデータを並列処理する。
環境分野の組織におけるインフラ不足への対応
多くの環境関連組織はデータラベリングや大規模推論の実行に必要なエンジニアリングチームやインフラを欠いている。OlmoEarth Platform はこれらの課題に対応し、モデルの微調整から大規模推論までのライフサイクルを支援する。
大陸スケールでの効率的な推論とコスト削減
同プラットフォームは約1日で大陸規模の領域における推論を実行し、平方キロメートルあたり数セント未満のコストで数十テラバイトの画像を処理できる。
重要な引用
The right hardware for the right task
One request, hundreds of workers, and thousands of processes
Handling failure at scale
That's why we built the OlmoEarth Platform: infrastructure for taking geospatial models from fine-tuning and evaluation to large-scale inference.
編集コメントを表示
編集コメント
衛星画像処理における大規模推論の課題に対し、ハードウェア選定と分散処理アーキテクチャという具体的な技術的アプローチで解決策を提示した点は評価できる。同社の OlmoEarth プラットフォームが実環境でどの程度の性能を発揮するかは、今後の適用事例に注目する必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
記事一覧に戻る
- 衛星推論がなぜ難しいのか
- タスクに最適なハードウェア
- 1 つのリクエスト、数百人のワーカー、数千のプロセス
- 必要なピクセルの特定と取得
- スケール上の障害への対応
- 私たちの今後の展望
🌍 OlmoEarth プラットフォームの詳細はこちら:https://allenai.org/olmoearth
OlmoEarth モデル は、約 10TB のマルチモーダル衛星データで事前学習された地球観測用ファウンデーションモデルのファミリーです。政府機関や NGO、ミッション志向の組織はすでに、森林伐採の監視、食料安全保障、山火事のリスク評価などの用途に OlmoEarth を活用し始めています。
AI2 では、強力なオープンモデルの訓練と公開に長けています。エンジニアリングチームが堅牢な組織であれば、オープンモデルさえあれば十分です。しかし、環境分野で最もこれらのモデルを適用できる立場にある多くの組織には、データラベリングやモデルの微調整、大規模推論の実行など、フルライフサイクルを管理するためのインフラやエンジニアリングチームがありません。
私たちは過去 10 年以上にわたり、Skylight や EarthRanger といったプラットフォームを運営してきました。これらは世界中のユーザーが毎日頼りにしているソフトウェアであり、毎日確実に動作する必要があります。この経験から、真のインパクトを生み出すために必要なことがわかりました。それは、適切なタイミングと場所でコスト効率よくモデルを実行し、パフォーマンスを監視し、生データを実行可能な洞察に変換し、その結果がパートナーが求める成果につながっていることを検証することです。
そこで私たちは、地理空間モデルを微調整・評価から大規模推論までつなぐインフラとして「OlmoEarth Platform」を構築しました。
このスケールでの推論には独自の課題があります。衛星画像は複数のプロバイダー間で検索・アクセス可能にし、投影法や解像度を整合させる必要があります。また、効率的に処理された結果は、分散コンピューティングで頻発する障害からインフラが回復しながらも、地理的に一貫した地図として統合されなければなりません。
現在、このプラットフォームは約1日で大陸規模の地域にわたる推論を実行でき、数十テラバイトの画像データを処理しながら、1平方キロメートルあたりのコストは数セント未満で抑えています。このシステムを開発する過程では、大規模な地理空間システムに取り組む他者も直面する可能性のある一連のエンジニアリング課題に直面しました。本稿では、それらの課題と私たちが導き出した解決策について解説します。
*OlmoEarthプラットフォーム上で生成された最新の山火事リスクマップ(統計情報付き)*
衛星画像推論が難しい理由
一般的な機械学習モデルは数メガバイトのデータを処理し、1秒未満で結果を出力します。例えば、LLMがテキストの段落を読み込んだり、コンピュータビジョンモデルがスマートフォンの写真を解析したりするケースです。一方、地球観測データの推論は全く異なるスケールで行われます。最大限のパフォーマンスを発揮させるために基盤モデルを微調整する単一のジョブでも、テラバイト規模のデータを移動させ、数時間にわたって実行されることがあります。
入力データは複数のスペクトルバンドやセンサータイプ、そして広大な地理領域にわたる時間ステップを含む可能性があります。データは複数のプロバイダーから提供されることも多く、それぞれが異なる投影法と解像度を使用しています。また、雲に隠れて観測できないデータが含まれていることもあります。
出力結果自体が地図となるため、すべての予測値は周囲の地域と同じ投影法および座標グリッドと厳密に整合させる必要があります。
データ取得自体が大きな課題になることもあります。予測ジョブでは、モデルを実行する時間よりも画像のダウンロードや準備に多くの時間を費やすことが多く、効率的なデータパイプラインが不可欠です。これらのパイプラインは大量の I/O を処理しつつ、画像の座標変換やリサンプリングに必要な計算資源も提供する必要があります。
適切なハードウェアを適切なタスクに割り当てる
データ取得と準備にジョブの実行時間の大部分が割かれるため、この作業を GPU に任せてしまうと、最も高価なハードウェアが本来 CPU が得意とするタスクに費やされてしまいます。そこで、各ジョブは 3 つのステージに分割され、それぞれに最適なハードウェアプロファイルが割り当てられます。
- データ取得と前処理(CPU、高 I/O): 画像の取得、座標変換、整列、正規化を行い、推論時に高速読み込みが可能となる形式で保存します。
- 推論(GPU): モデルの順方向計算を実行し、最小限の加工を施した出力を直接ストレージに書き出します。
- 後処理(CPU): ウィンドウごとの出力をつなぎ合わせ、マスク適用やリサイズを行い、Zarr、GeoTIFF、GeoJSON などユーザーが扱いやすい形式でエクスポートします。
OlmoEarth Platform はこれらのステージを多数のマシンに分散させつつ、GPU の稼働率を最大化します。マルチプロセスデータローダーが各 GPU に継続的にデータを供給し、完了した出力は直接 Blob ストレージへストリーミングされます。
1 つのリクエストで数百人のワーカーと数千のプロセス
「OlmoEarth Run」は、大規模な推論ジョブを実行するためのプラットフォームのレイヤーです。この仕組みでは、各ジョブがカバーする地理領域を個別の計算インスタンス(ワーカー)に適したサイズのパーティションに分割し、さらにそれを OlmoEarth モデルが処理する小さなウィンドウへと細分化します。各ウィンドウは独立して順次パスで処理できるため、地図の一部に対する作業が他の部分の完了を待たずに並行して進められます。
実際には、州程度の広さが約 100 のパーティションに分割される一方、大陸規模の実行では数千ものパーティションが発生します。隣接するパーティションはわずかに重なり合っており、出力を組み立てる際にこの重複部分を調整することで、最終的なラスタデータに継ぎ目が見えないようにしています。
パーティションが独立しているため、同じ処理段階を数千の計算インスタンスで同時に実行可能です。最近ではこのアプローチを用いて、北米全域をカバーする山火事リスクマップを生成しました。ピーク時には約 19,600 の CPU と 994 の GPU を並列で使用し、ネットワークスループットは 168 GB/s を超えました。この並列処理のレベルにより、推定で 4,737 時間かかるシリアル計算を、実質的な経過時間で約 30.5 時間に短縮。155 倍の高速化を実現しました。
ただし、並列処理(ファンアウト)は無制限ではありません。ワーカー数を増やすとクラウドのクォータに抵触するため、並列度は実行ごとの設定項目であり、個々のジョブで調整可能なパラメータの一つです。
出力解像度を上げれば詳細度が向上する反面、データ量と計算リソースが増大します。モデルサイズを大きくすれば精度は上がりますが、GPU 使用時間が長くなります。また、生画像をキャッシュすれば、同じエリアを繰り返し処理する際の速度が向上しますが、その分ストレージ容量を消費します。
最適な設定は、現在のタスク内容と予算によって異なります。
適切なピクセルの検索と取得
地理的な領域と時間範囲が指定されると、プラットフォームはまず、モデルに供給すべき衛星画像シーンを特定する必要があります。これは、異なるカタログ、フォーマット、公開遅延を持つ複数のプロバイダー間で、「何が」「どこで」「いつ」撮影されたかを把握することを意味します。
選定基準はデータソースによっても異なります。Sentinel-2 などの光学画像では、一般的に雲の少ないシーンを優先しますが、合成開口レーダー(SAR)の場合には、利用可能な偏波チャネルの方が重要になることがあります。
可能であれば、私たちはパブリックな STAC カタログやオープン標準を利用しています。ただし、大規模な推論ジョブを実行すると、一度に数千件のメタデータクエリが発生する可能性があります。これは、ESA や Microsoft Planetary Computer の STAC API などが同時に処理できるように設計されている数値を遥かに超えるものです。
外部サービスの負荷を避けるため、OlmoEarth プラットフォームは独自のメタデータインデックスを維持し、新しい画像が公開されるたびに更新しています。AWS Open Data を通じてホストされているデータセットの場合、新しいシーンごとに SNS 通知を受信します。変更ストリームを提供しないプロバイダーについては、数分ごとに上位のインデックスをポーリングして対応しています。その結果、外部サービスへのリクエストは、大規模な推論ジョブによって発生する突発的なバーストではなく、新規公開のペースに合わせた安定した流れで処理されます。
各インデックスエントリには、シーンのメタデータと、基盤となるピクセルが利用可能なすべての場所へのポインタが格納されています。実行時にはプラットフォームが最適なソースを選択し、COG や Zarr といったクラウド最適化フォーマットに対してウィンドウ読み込みを実行します。これにより、必要なバイトのみを取得でき、シーン全体をダウンロードする必要はありません。
このインデックスは注釈ツールのサポートも担っています。Sentinel-1、Sentinel-2、Landsat、NISAR の画像がクラウド最適化フォーマットでポインタとして保持されているため、個別の取り込みパイプラインを構築することなく、同じウィンドウ読み込みシステムを通じてインデックスされたシーンのタイルを配信できます。
このワークフローを最も容易に実現したプロバイダーには、地球観測データの公開におけるベストプラクティスとして推奨できる 3 つの特徴がありました。それは、新しい画像が利用可能になった際のキューベース通知、専用レート制限や可用性のボトルネックのない主要クラウドプラットフォーム上でのストレージ、そして範囲読み込みをサポートするクラウド最適化フォーマットです。
OlmoEarth の衛星画像インデックスに対するサンプルクエリ。サンフランシスコ上空の 6 月第 1 週で雲量が最も少ない Sentinel-2 画像を検索するものです。このサービスは、利用可能な最良の画像を返すとともに、Amazon S3 バケット内の GeoTIFF ファイルへのポインタも提供します。
スケールにおける障害対応
OlmoEarth プラットフォームは、障害から自動的に回復するように設計されています。各ステージと地理的パーティション内のタスクごとに、ランナー Docker コンテナを実行する仮想マシンを動的にプロビジョニングします。ランナーはタスクパラメータを取得して実行し、結果を返した後にシャットダウンします。すべてのタスクが再入力でかつ冪等性を持つため、一時的な障害も安全に再実行することで処理できます。
この規模では障害は想定内の出来事です。プロバイダーの応答が遅れる、一時的に利用不可になる、メタデータには画像が存在すると示されていても必要なバンドやウィンドウが欠落している、雲の厚さによって利用可能な観測数が不足する、あるいはタスク自体がクラッシュするといったケースが発生します。プラットフォームは、タスク追跡機能、自動リトライ機能、代替プロバイダーへのフォールバック(利用可能な場合)、そして再試行可能なエラーと致命的なエラーを明確に区別する仕組みで対応します。また、独立した監視プロセスがスタックまたは停止したランナーを検知し、そのタスクを再起動します。
今後の展望
当社のロードマップは、パートナーから指摘されたギャップや最も必要とされている機能によって形作られています。現在取り組んでいる主な領域は以下の通りです:
自動モデル実行。関心のあるエリアに新しい映像が検出されたタイミングで、または事前にスケジュールを設定して推論ジョブをトリガーできます。
変化検知とアラート。監視対象の景観に変化が生じた際にユーザーへ通知し、森林伐採や洪水などの事象を手動でラスタデータを探して確認するのではなく、自動的にアラートとして表面化させます。
エージェントツールとインターフェース。エージェントは地理空間モデルの利用障壁を下げます。データのキュレーションから特徴量のエンジニアリング、微調整済みモデルの改善方法の特定まで担います。技術レベルに関わらず、従来は熟練した ML 研究者にしかできなかった作業を誰でも実行できるようにします。
高速化されたモデル。GPU 使用時間を削減するより効率的なアーキテクチャについて、当社の研究チームと連携して開発します。
多様なモダリティの追加。OlmoEarth モデルおよびそれを支える映像インデックスに、新たな衛星センサーやデータソースを追加します。現在は気象データ(ERA-5)や環境要因をより詳細に捉える衛星の統合に注力しています。
埋め込みベクトル。専用の埋め込みモデルを開発し、グローバルスケールで事前計算を行います。多くのタスクにおいて、生映像全体に対して順方向パスを実行する代わりに、これらの埋め込みベクトルに対する推論を行うことで、ワークロードを大幅に高速化・低コスト化できます。難易度の高いタスクでは微調整や直接推論が最大性能のために重要ですが、埋め込みベクトルはより広範な効率的なアプリケーションの可能性を開きます。
どこでも実行可能。OlmoEarth Run では、Docker イメージを実行できる仮想マシンとバロブストレージへのアクセス権限があれば十分です。現在は Google Cloud で運用していますが、アーキテクチャは複数のクラウドや、パートナーのアカウント内・計算環境での展開にも対応するように設計されています。
これは始まりに過ぎません。地理空間基盤モデル、特にその実用化はまだ発展途上の技術であり、これらを最も活用できる立場にある組織(保全、食料安全保障、災害対応、気候変動分野で活動する団体など)は、これまでこうしたインフラに触れる機会がありませんでした。これらのチームが地球について理解するために必要なことと、予算や技術リソースによって実際に実行可能なことの間に、依然として大きな隔たりがあります。
私たちはこの格差を埋めるために OlmoEarth を構築しています。
原文を表示
- Why satellite inference is challenging
- The right hardware for the right task
- One request, hundreds of workers, and thousands of processes
- Finding and fetching the right pixels
- Handling failure at scale
- Where we're headed
🌍 Learn more about OlmoEarth Platform: https://allenai.org/olmoearth
The OlmoEarth models are our family of Earth observation foundation models, pretrained on roughly 10 terabytes of multimodal satellite data. Governments, NGOs, and other mission-driven organizations are already adapting OlmoEarth for applications including deforestation monitoring, food security, and wildfire risk.
At Ai2, we know how to train and release powerful open models, and for organizations with strong engineering teams, an open model is all they need to run with. But most organizations in the environmental space – the ones best placed to apply these models – don't have the infrastructure or engineering teams that can manage the full lifecycle: labeling data, fine-tuning models, and running large-scale inference. We’ve spent more than a decade operating platforms like Skylight and EarthRanger, software that users around the world rely on every day, so it has to work every day. That experience taught us what delivering impact takes: running models cost-effectively at the right time and place, monitoring performance, turning raw outputs into actionable insights, and verifying those outputs drive the outcomes partners want.
That’s why we built the OlmoEarth Platform: infrastructure for taking geospatial models from fine-tuning and evaluation to large-scale inference.
Inference at this scale presents its own set of challenges. Satellite imagery must be found and accessed across multiple providers, aligned across projections and resolutions, and processed efficiently. Results then have to be stitched into geographically consistent maps while the infrastructure recovers from the routine failures of distributed computing.
Today, the platform can run inference across continent-scale areas in roughly a day, processing dozens of terabytes of imagery at a cost of fractions of a penny per square kilometer. Developing it meant confronting a series of engineering challenges that others working on large-scale geospatial systems are likely to encounter as well. This post walks through those challenges and the solutions we arrived at.
*A recent wildfire risk map generated on the OlmoEarth Platform, with statistics.*
Why satellite inference is challenging
Most ML models take in a few megabytes of data and produce a result in under a second—think LLMs processing a paragraph of text or computer vision models analyzing a photo from a smartphone. Earth observation inference operates at a very different scale—a single job fine-tuning a foundation model for maximum performance can move terabytes of data and run for hours. The inputs may span multiple spectral bands, sensor types, and time steps across a large geographic area. They can come from several providers, each using different projections and resolutions, and may include observations that are missing or obscured by clouds. The output is itself a map, so every prediction must remain precisely aligned with the same projection and coordinate grid as the areas around it.
Even acquiring the data can be a major challenge. Prediction jobs often spend more time downloading and preparing imagery than running the model itself, making efficient data pipelines critical. Those pipelines must handle high-volume I/O while providing the compute needed to reproject and resample imagery.
The right hardware for the right task
Because data acquisition and preparation often dominate an inference job’s runtime, assigning that work to GPUs would leave the system’s most expensive hardware doing tasks better suited to CPUs. We therefore divide each job into three stages, each matched to a distinct hardware profile:
- Data acquisition and preprocessing (CPU, high I/O): Fetch, reproject, align, and normalize imagery, then write it in a format optimized for fast loading during inference.
- Inference (GPU): Run the model’s forward pass and write minimally processed outputs directly to storage.
- Postprocessing (CPU): Stitch the per-window outputs together, apply masks or rescaling, and export them in user-friendly formats such as Zarr, GeoTIFF, or GeoJSON.
The OlmoEarth Platform distributes these stages across many machines while keeping GPUs fully utilized. Multiprocess data loaders continuously feed each GPU, while completed outputs stream directly to blob storage.
One request, hundreds of workers, and thousands of processes
OlmoEarth Run, the platform’s execution layer for large-scale inference jobs, divides the geographic region covered by each job into partitions sized for individual compute instances (workers), then subdivides those partitions into smaller windows that the OlmoEarth models process. Because each window can be handled independently in a separate forward pass, work on one part of the map does not need to wait for another.
In practice, a state-sized area might become a hundred or so partitions, while a continent-scale run can become thousands. Adjacent partitions overlap slightly, and we reconcile that overlap when the outputs are assembled so no seam appears in the final raster.
Because the partitions are independent, the same stage can run across thousands of compute instances at once. We recently used this approach to generate a wildfire-risk map covering all of North America. At peak, the run used roughly 19,600 CPUs and 994 GPUs in parallel, with network throughput exceeding 168 GB/s. That level of parallelism reduced an estimated 4,737 hours of serial compute to about 30.5 hours of wall-clock time—a 155× speedup.
Fan-out is not unbounded, though. More workers push against cloud quotas, so parallelism is a per-run knob, one of several that we can adjust on individual jobs. Output resolution trades data volume and compute for detail; model size trades GPU time for accuracy; caching raw imagery trades storage for speed across repeated runs over the same area. The right setting depends on the task – and budget – at hand.
Finding and fetching the right pixels
Given a geographic region and time range, the platform first has to determine which satellite scenes should feed the model. That means identifying what was captured, where, and when across providers with different catalogs, formats, and publication delays. The selection criteria also depend on the source: for optical imagery such as Sentinel-2, we generally want the least-cloudy scenes available, while for synthetic aperture radar, the available polarization channels may matter more.
We rely on public STAC catalogs and open standards wherever possible. But a large inference job can generate thousands of metadata queries at once—far more than external services such as ESA’s or Microsoft Planetary Computer’s STAC APIs are designed to handle concurrently.
To avoid overwhelming those services, the OlmoEarth Platform maintains its own metadata index, updated as new imagery is published. For datasets hosted through AWS Open Data, we receive an SNS notification for each new scene. When a provider does not offer a change stream, we poll its upstream index every few minutes. As a result, our requests to external services follow the steady pace of new publications rather than the sudden burst generated by a large inference job.
Each index entry stores scene metadata along with pointers to every location where the underlying pixels are available. At runtime, the platform selects the best source and performs windowed reads against cloud-optimized formats such as COG or Zarr, retrieving only the bytes needed for a given partition rather than downloading entire scenes.
The index also supports our annotation tools. Because it maintains pointers to Sentinel-1, Sentinel-2, Landsat, and NISAR imagery in cloud-optimized formats, we can serve tiles from any indexed scene through the same windowed-read system, without building a separate ingestion pipeline.
The providers that made this workflow easiest shared three characteristics that we would recommend as best practices for publishing Earth observation data: queue-based notifications when new imagery becomes available, storage on major cloud platforms without bespoke rate limits or availability bottlenecks, and cloud-optimized formats that support ranged reads.
*A sample query against the OlmoEarth satellite imagery index, searching for the least-cloudy Sentinel-2 image over San Francisco during the first week of June. The service returns the best available image along with a pointer to the GeoTIFF file in an Amazon S3 bucket.*
Handling failure at scale
The OlmoEarth Platform is designed to recover from failures automatically. For each task within a stage and geographic partition, it dynamically provisions a virtual machine running our runner Docker container. The runner retrieves the task parameters, executes the work, returns the result, and shuts down. Because every task is reentrant and idempotent, intermittent failures can be handled safely by rerunning it.
At this scale, failures are expected: a provider may be slow or briefly unavailable; metadata may indicate that imagery exists even when a required band or window is missing; cloud cover may leave too few usable observations; or a task may crash outright. The platform responds with task tracking, automatic retries, fallback to alternate providers when available, and clear distinctions between retryable and fatal errors. A separate monitoring process detects runners that have stalled or stopped and restarts their tasks.
Where we're headed
Our roadmap is shaped by the gaps our partners have identified and the capabilities they need most. Among the areas we are working on:
- Automated model runs. Schedule inference jobs in advance or trigger them whenever the imagery index registers a new scene over an area of interest.
- Change detection and alerts. Notify users when the landscapes they monitor change, so events such as deforestation or flooding surface as alerts rather than rasters someone has to find and inspect manually.
- Agentic tools and interfaces. Agents can lower the barrier to using geospatial models, from data curation and feature engineering to identifying ways to improve a fine-tuned model. We want users at any technical level to do work that previously required an experienced ML researcher.
- Faster models. Work with our research team on more efficient architectures that reduce GPU time per window.
- More modalities. Add new satellite sensors and data sources to both the OlmoEarth models and the imagery index that supports them. We’re currently focused on incorporating weather data (ERA-5) and satellites that provide more detail on environmental factors.
- Embeddings. Develop a dedicated embedding model and precompute embeddings at global scale. For many tasks, running inference against those embeddings could replace a full forward pass over raw imagery, making workloads substantially faster and less expensive. Fine-tuning and direct inference will remain important for maximum performance on challenging tasks, but embeddings could open up a much broader range of efficient applications.
- Run anywhere. OlmoEarth Run requires only virtual machines capable of running a Docker image and access to blob storage. We currently operate it on Google Cloud, but the architecture is designed to support multiple clouds and deployments within a partner’s own account and compute environment.
This is only the beginning. Geospatial foundation models, and especially operationalizing them, are still an emerging technology, and many of the organizations best positioned to use them – those working in conservation, food security, disaster response, and climate – have never had access to infrastructure like this. The gap between what these teams need to understand about the planet and what their budgets and technical resources allow them to do remains wide.
We are building OlmoEarth to help close it.
AI算出
主要ニュースainew評価高い
AI モデルと大規模インフラの具体的な実装・性能データ(10TB データ、数セント/km²など)を含む新規発表であり、既存記事との比較でも独自の実装課題や解決策という情報増分があるため高評価とした。日本固有の導入事例や規制情報は含まれていないため関連性は低めとした。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 78
- 日本での有用性
- 25
他社はどう報じたか
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み