F1、AWS上のエージェント型AIでデータ運用を加速し数週間から数分に短縮
本文の状態
日本語全文を表示中
詳細モードで約22分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
AWS Machine Learning Blog
F1 は AWS のアジェンティック AI を活用し、データソースの統合期間を最大 8 週間から約 40 分へ短縮する「Data Accelerator」を導入して運用効率を劇的に改善した。
AI深層分析を開く2026年8月4日 03:26
AI深層分析
キーポイント
データ統合プロセスの自動化
F1 は従来の手動工程(6〜8 週間)に代わり、AWS Bedrock AgentCore を活用したアジェンティック AI でデータソースのオンボーディングを約 40 分のコード生成と数時間のデプロイで完了させる仕組みを実現した。
運用課題の根本解決
エンジニアが手動でスキーママッピングやガバナンスポリシーを設定していた従来の方式では対応しきれなかった、頻繁な仕様変更への追従とデータ品質の問題を、再現性のある堅牢なシステムで解消した。
可視化とコラボレーションの強化
単なるアラートダッシュボードを超え、データラインジと根本原因分析を含むエンドツーエンドの可視性を提供し、アナリスト、エンジニア、科学者が同一環境で協働できる基盤を構築した。
エージェントによるデータソースの自動オンボーディング
限られた情報に基づくビジネス要件文書から、インフラコードやGDPR分類を含む生産準備完了パイプラインを生成する。エンジニアが手動でスキーママッピングやガバナンスポリシーを設定する必要がなくなる。
データソースの統合と可視性の課題
従来のプロセスでは新規ソースへの対応に6〜8週間を要し、仕様変更やログの断片化により問題解決に時間がかかっていた。
重要な引用
Our MarTech platform is the nervous system of F1's fan engagement. But every new data source required 6 to 8 weeks of manual engineering.
AWS worked backwards from our needs to implement an agentic solution that worked end to end, applying business logic at each step.
For the first time, we have end-to-end visibility across the entire MarTech platform with data lineage and root cause analysis, not just dashboards full of alerts.
First, onboarding each new data source was a heavily manual effort: engineers wrote schema mappings, built ingestion pipelines, configured data quality checks, defined General Data Protection Regulation (GDPR) classifications, and set governance policies by hand. This process took 6 to 8 weeks per source.
編集コメントを表示
編集コメント
F1 の事例は、アジェンティック AI が単なる実験段階を超え、大規模なデータ基盤の運用課題を解決する実用ツールとして成熟したことを示す重要なケーススタディである。特に「ビジネスロジックを各ステップで適用する」という設計思想は、複雑な業務プロセスにおける自律化の鍵となる要素と言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
F1(フォーミュラ・ワン)は、デジタルプラットフォーム、F1 TV、ソーシャルメディア、チケット販売、グッズ販売などを通じて、年間を通じて世界中で8億人以上のファンにリーチしています。レースは2週間に一度開催され、ファンのエンゲージメントを測る時間は数分単位です。商業的な意思決定も、ピットレーンと同じスピードで行わなければなりません。
その裏側では、F1のマーケティングテクノロジー(MarTech)プラットフォーム「Customer 360」が、すべての接点でのインタラクションを収集し、パーソナライゼーションやセグメンテーション、商業戦略を支えています。しかし、このプラットフォームには大きな運用上の課題がありました。F1のITディレクターであるChris Roberts氏はこう語ります。「当社のMarTechプラットフォームは、F1のファンエンゲージメントにおける神経系のようなものです。しかし、新しいデータソースを追加するたびに、手動でのエンジニアリングに6〜8週間を要していました。12の新しいソースを統合するために、すでに18ヶ月分のバックログを抱えていたのです。」
ビジネス側で生成されるデータの速度が、エンジニアチームが接続できる速度を上回っていました。そこでF1データ運用責任者のMatt Kemp氏は、効率化とデータ品質の向上に取り組むことにしました。「手動でのデータソース取り込みは時間がかかり、ソリューションにばらつきを生み、最終的にはデータの整合性に問題を引き起こします。私は、反復可能で堅牢かつ信頼性の高い解決策を求めていました。AWSは私たちのニーズから逆算してエンドツーエンドで機能するアジェンティック・ソリューションを実装し、各ステップでビジネスロジックを適用しました。」
2026 年初、F1 と AWS は「Data Accelerator」というソリューションを共同開発しました。これは Amazon Bedrock AgentCore で動作するエージェント AI を活用し、F1 の MarTech データプラットフォームを手動管理型のシステムから、自己管理・可視化・統合されたデータ基盤へと変革するものです。
本稿では、Data Accelerator がデータソースの導入に要していた最大 8 週間を、コード生成に約 40 分、デプロイに数時間を短縮した事例を紹介します。また、本番環境でのデータソースの不具合を検知・修正し、データプラットフォームの運用状況やエージェントの系譜(ラインージ)を単一の画面で追跡可能にしました。これにより、アナリスト、エンジニア、科学者が協働するための新たな窓口も開かれました。
「初めて、アラートだらけのダッシュボードではなく、データ系譜と根本原因分析を含む MarTech プラットフォーム全体のエンドツーエンドな可視性が実現できました」とロバーツは語ります。
F1 の「Customer 360」プラットフォームは、チケット販売パートナーやストリーミング連携、スポンサーのアクティベーションフィード、ソーシャルメディア、商品販売システムなどからデータを収集しています。この広範囲かつ高速なデータ基盤を運用する中で、チームは解決すべき 3 つの課題に直面しました。
まず、新しいデータソースを追加する際の手間は非常に大きかったことです。エンジニアが手作業でスキーマのマッピングを行い、取り込みパイプラインを構築し、データの品質チェックを設定し、GDPR(一般データ保護規則)の分類を定義し、ガバナンスポリシーを手動で設定する必要がありました。このプロセスには、ソースごとに 6〜8 週間かかりました。
次に、プラットフォームは絶えず変化する上流側のフィードに対応し続けなければなりませんでした。プロバイダーが通知なしに列名を変更したり、フィールドを追加したり、ペイロードの構造やスケジュールを再構成することが頻繁にありました。こうした変更は、レースウィークエンド中や重要なキャンペーンの立ち上げ時など、最もタイミングが悪い時に発生することが多かったです。
最後に、可視性が分断されていました。ログは各サービスに散在しており、統一されたデータ系譜(データラインジ)がありませんでした。ステークホルダーが特定の指標について疑問を呈した際、エンジニアは Amazon S3 のパスや Amazon Redshift の管理テーブル、Airflow のログ、DBT の出力などを手動で追跡するのに数時間を要しました。
解決策の概要
「Data Accelerator」は、以下の 5 つのワークストリームを同時に実行することでこれらの課題に対処しました:
- Amazon Bedrock AgentCore を活用したエージェントによるデータソースのオンボーディング。エージェントはランタイムコンテナ内でホストされます。
- スキーマの変化を検出・修復する自動化機能。
Amazon SageMaker Unified Studio を通じた統一されたデータアクセス。
根本原因分析ツール(RCA)とコンテキストグラフを活用したエンドツーエンドの観測性。
観測性ダッシュボードで障害を自動的に特定し、コード変更で解決可能な場合はエージェントによる操作を実行する仕組み。
6 つ目の作業ストリームでは、顧客アイデンティティ解決アルゴリズムを最適化し、複数のチャネルにまたがるファンの接点を統合しました。以下、各作業ストリームについて詳しく説明します。
エージェント型データソースのオンボーディング
Data Accelerator の中核となるのは、プラットフォームエージェント群です。これらのエージェントは、データソースに関する情報が限られたビジネス要件定義書(BRD)を受け取り、インフラストラクチャコード、データ変換処理、ガバナンスポリシー、GDPR 分類を含む、そのまま本番環境で稼働可能なオンボーディングパイプラインを生成します。人間がボイラープレートコードを一行も記述する必要はありません。エージェントは以下の2段階で動作します。
フェーズ1:設定の自動生成
新しいデータソースの導入が必要になった際、チームメンバーは BRD(ビジネス要件定義書)を Amazon S3 バケットにアップロードします。このアップロードがトリガーとなり、AWS Lambda 関数が起動し、Amazon Bedrock AgentCore の機能である「Amazon Bedrock AgentCore Runtime」を呼び出します。
エージェントは BRD を読み込み、一連の設定ファイルを生成します。その後、GitHub App を通じて GitHub にアクセスし、これらのファイルを標準化された Git リポジトリへのプルリクエストとしてプッシュします。また、REST API を介して Jira にアクセスし、該当する PR を参照したチケットを作成します。
エージェントのすべての会話とアクションは、Amazon Bedrock AgentCore の組み込み観測機能を通じて Amazon CloudWatch で追跡されます。割り当てられたエンジニアがこれらを確認し、必要に応じて調整・承認を行います。

フェーズ 1 のワークフロー:BRD のアップロードがトリガーとなり、エージェントが設定ファイルを生成してプルリクエストを開く
フェーズ 2:パイプライン全体の生成
設定ファイルの承認後、人間が次のステージを開始します。エージェントは承認された設定に基づき、3 つの独立したプルリクエストを生成します。
- AWS Glue アプリケーションおよびインフラストラクチャコード。
- DBT 変換フレームワーク。
- GDPR タグ付けを含むガバナンスポリシー。
3 つの PR はすべて、追跡可能性のために単一の Jira チケットにリンクされています。エンジニアはインフラストラクチャ、DBT、ガバナンスのリポジトリでそれぞれをレビューし、承認します。

フェーズ 2 のワークフロー:エージェントがインフラストラクチャ、変換、ガバナンスのプルリクエストを生成し、これらは単一の Jira チケットにリンクされます。
自動化された GDPR クラス分類
このシステムが単純なコードジェネレーターと異なる点は、統合された GDPR クラス分類機能にあります。エージェントはすべてのデータ列を能動的に分析し、個人データ、機微な個人情報、または擬似匿名化データを含まれているかを判定して、適切な GDPR カテゴリでタグ付けします。これらのタグは SageMaker Unified Studio のガバナンスレジストリに直接公開されるため、コンプライアンスチームは手動のレビューサイクルを経ずに即時に可視化できます。
モジュラー型スキルアーキテクチャ
本システムは、密結合したエージェントグラフではありません。単一のエージェントがモジュール化されたスキル定義を用いて動作し、それぞれが固有の機能をカプセル化しています。具体的には、スキーママッピングとデータ型の推論、データの品質検証、ガバナンスの適用、機微データの分類といった機能です。
実行時には、エージェントが入力された要件を評価し、関連するスキルを活性化します。これらは多段階の推論プロセスを通じて組み合わされます。0 段目はトークン管理(スクラビング)を担当し、1 段目でツールの出力を要約し、2 段目では全体の評価を統合します。このようにして、一度きりの回答に依存するのではなく、精度と完全性を段階的に高めていきます。
新しい機能は、コアとなるエージェントループを変更することなく、新たなスキルモジュールとして提供されます。これにより、プラットフォームが成長しても、アーキテクチャの保守性と組み合わせやすさが保たれます。
その結果、オンボーディングにかかる時間が従来の 6〜8 週間から、コード生成に約 40 分、デプロイとレビューに数時間を要する程度まで短縮されました。現在、作業の 95% を AI エージェントが自律的に処理しています。
スキーマ進化の自動化
新しいデータソースの導入は課題の一つですが、既存の連携を健全に保つこともまた重要な課題です。アップストリームのプロバイダーは頻繁にデータ構造を変更します。カラムの名前変更から新フィールドの追加まで多岐にわたります。
従来、F1 チームがこうした変更を発見できたのは、パイプラインが失敗した時だけでした。しかも、その多くが生レースウィーク中に発生していました。
導入プロセスを担うのと同じエージェントアーキテクチャが、現在はアップストリームのスキーマ変更を継続的に監視しています。プロバイダーがデータ構造を変更すると、エージェントは AWS Lambda と Amazon EventBridge を活用したイベント駆動型のトリガーを通じてこれを検知します。そして、ダウンストリームへの影響を評価し、どのパイプラインが影響を受け、どのコンシューマーが変更されたフィールドに依存しているかを特定します。
その後、影響を受けるすべてのリポジトリに必要なコード更新を生成し、文脈と関連するプルリクエスト(PR)を含めた Jira チケットを作成します。エンジニアには、何が変わったか、どのような影響があるか、そしてレビュー用の提案された修正内容が記載された通知が届きます。
エンドツーエンドでの解決にかかる時間は、従来は数日かかったものが、現在は数時間で済むようになりました。

スキーマ進化におけるエージェントワークフロー
Amazon SageMaker Unified Studio による統一データアクセス
データアクセラレーター導入前は、Customer 360 のデータを扱うために複数の分断された環境を行き来する必要がありました。データエンジニアは特定のアカウントでパイプラインを構築し、ファンの行動をモデル化したいデータサイエンティストは別のアカウントにアクセスする必要があり、アナリストはさらに全く異なる環境で作業していました。ツールや文脈の共有はなく、分析を開始する前に質問から回答を得るまでに数日間の調整が必要でした。
この解決策では、Amazon SageMaker Unified Studio を基盤としたデータメッシュ・フレームワークを採用しています。中央のガバナンスアカウントが複数のプロデューサーチーム間でのデータ発見とアクセス仲介を担当します。鍵となるのは、ガバナンスを手動コンソール操作ではなく宣言型設定としてコード化している点です。単一のデータソース定義により、カタログへのデータ公開と同時に、消費者が購読するために必要なアクセス制御も自動的に用意されます。これにより、エージェントはストレージからカタログ、そして管理されたアクセスまでを一貫して安全に新規データ製品に登録できます。これはフレームワーク自体がセキュリティ制約を構造的に強制しているためです。IAM ポリシーや AWS Lake Formation の付与権限を手動でレビューする必要はありません。プラットフォームが構造的な正しさを保証します。これが「単一のフロントドア」を実現する理由です。
データエンジニアは、1 つの場所でデータのキュレーションとガバナンスを行い、データサイエンティストも同じ環境でそのデータを発見します。そこでは、データが適切に管理され、文書化されており、モデリングの準備ができています。ファンセグメンテーションモデルを構築したり、顧客IDアルゴリズムを最適化したりするデータサイエンティストは、データがどこに保存されているか、パイプラインの所有者が誰か、どの S3 プレフィックスを使用すべきかなどを知る必要はありません。彼らは Unified Studio を開き、キュレーションされた Customer 360 データセットを見つけるとすぐにモデリングを開始できます。宣言的なガバナンスにより、コントロールを犠牲にすることなく安全なセルフサービスが可能になったため、共有ノートブックや一貫したツール、管理されたアクセス権が提供されます。これで、データのキュレーションと利用がついに並行して実行されるようになりました。
RCA とコンテキストグラフによるエンドツーエンドの可観測性
データプラットフォームの信頼性は、チームが「現在のデータは正しいのか」という 1 つの問いに答える能力にかかっています。Data Accelerator の導入前は、この質問に答えるには Apache Airflow にログインし、Amazon S3 のパスを確認し、Amazon Redshift の管理テーブルを照会し、DBT ログを読み込む必要がありました。「誰も全体像を持っていませんでした。ステークホルダーが『なぜこの数値がおかしいのか』と尋ねても、『数時間ください』という返事しか得られませんでした」とロバーツは付け加えます。「可観測性ダッシュボードはこの状況を根本から変えました。」
観測レイヤーでは、S3 からの生データ取り込みから処理済み層を経て Amazon Redshift の DBT ステージに至るまでの完全なデータ系譜が、健康状態を色分けした単一の対話型グラフとして可視化されています。ユーザーは任意のノードをクリックして詳細を確認でき、各ソースやテーブルにはパス/フェイルの状態、最終実行時刻、所要時間が表示されます。
パイプラインに障害が発生した場合、この系譜ビジュアライゼーションによって、どこで断絶が生じたか、そしてどの下流データが影響を受けるかを正確に特定できます。
F1 のプラットフォームには、システムログを読み込み、データ基盤全体における障害点を特定する「原因究明(RCA)」というエージェント型ツールが組み込まれています。単独の RCA ツールでは「何が失敗したか」を告げるだけですが、これを JSON 形式で定義されたビジネスコンテキストやシステムトポロジーと組み合わせることで、その能力を強化しています。
例えば、S3 にファイルが存在しないことがエラーの原因である場合でも、コンテキストグラフを活用すれば、RCA は「上流の提供者が配送ウィンドウを再設定したため、パイプライン実行時にファイルが存在しなかった」という背景まで説明します。これは単に失敗を知ることと、その理由を理解することの違いです。
これで F1 では、データ系譜、因果関係に基づく原因究明、そしてビジネスコンテキスト定義が一つの場所に統合され、ダッシュボードは 15 分ごとに自動更新されるようになりました。

データソースと処理ステージ全体にわたるパイプラインの健康状態を示すデータ系譜ビジュアライゼーション

障害の詳細を含む観測性ダッシュボード
カスタマーアイデンティティの解決
最終的な作業チームは、F1 のファンが関わるすべてのタッチポイントで顧客のアイデンティティを統合するアルゴリズムの最適化に取り組みました。同じファンでも、アプリでの利用、ウェブサイトのチケット購入、F1 TV での視聴、SNS でのエンゲージメントなど、複数の経路を通じて接触します。これらの多様なインタラクションを、誤った結合や見落としなく単一のアイデンティティとして統合することが、Fan Personalization Platform(FPP)における効果的なパーソナライゼーションを実現する鍵となります。
F1 にはすでに機能するアイデンティティ解決のプロセスが存在していましたが、チャンネルを超えたファンインタラクションの増加に伴い、処理速度が遅く、スケーリングに課題を抱えていました。パイプラインの再構築やコンポーネントの置き換えを行うのではなく、チームは既存の解決アルゴリズムの計算パフォーマンスの最適化に注力しました。実行上のボトルネックを特定し、マッチングロジックを調整した結果、処理時間を 50% 短縮することに成功しています。その際、解決パイプライン全体やその下流との連携には一切変更を加えず、完全な状態を維持しました。プロセスの変更も精度の妥協もなく、F1 の本番環境規模において、同じアルゴリズムが半分の時間で動作するようになっています。
解決が早くなれば、F1 は新しいデータソースをオンボーディングし、顧客データを収集する時間を半分に短縮できます。解決速度の向上は、より鮮度が高く統一されたプロファイルにつながり、それがすべてのマーケティングチャネルでタイムリーかつ関連性の高いパーソナライゼーションを実現します。
「要は、メール、F1 アプリ、チケット販売、ソーシャルメディアなど、どのチャネルであっても、適切なファンに、適切なタイミングで、正しいメッセージを届けることです。ソースのオンボーディングが数時間で可能になり、アイデンティティ解決も迅速になったことで、ファンの期待に応えるパーソナライズされた体験を、すべてのマーケティングチャネルで実際に提供できるようになりました」とケンプは語ります。
デザインによるセキュリティとガバナンス
Data Accelerator は「AI が提案し、人間が検証する」という原則に基づいて動作します。エージェントは Amazon Bedrock AgentCore で実行され、長期メモリを備えており、呼び出し間でも文脈を保持します。開発には構造化された仕様駆動型の開発ツールである Kiro と、基盤モデルとして Amazon Bedrock(Claude Sonnet 4.6)が使用されました。イベント駆動のバックボーンでは、計算に AWS Lambda を、ルーティングに Amazon EventBridge を、ワークフローオーケストレーションに Amazon Managed Workflows for Apache Airflow (MWAA) を、生データ層に Amazon S3 をそれぞれ採用しています。すべての AI モデルへのアクセスは、F1 の AI Gateway によって一元化されたアクセス制御、コスト管理、監査ログの観点から統制されています。しかし、アーキテクチャがすべてではありません。本番環境で運用可能にするのはセキュリティ体制です。
セキュリティ体制には以下が含まれます:
最小権限の原則:細粒度のアクセス制御、有効期限が 1 時間の短期トークン、特定のリポジトリとリソースへのアクセス制限。
完全な監査証跡:すべてのアクションはログに記録され、コンプライアンスのために誰が行ったかが追跡可能。
人間のレビュー:生成された Pull Request はすべてエンジニアの承認を経てからマージされる。
自動テスト:エージェントは自分自身の変更に対する包括的なテストを生成する。
ロールバック機能:マージ後に発生した問題も即座に元に戻せる。
ネットワーク分離:システム全体は、インターネットへの直接アクセスがない Amazon Virtual Private Cloud (Amazon VPC) 内のプライベートサブネット内で実行される。
暗号化された認証情報:すべてのシークレットは AWS Systems Manager Parameter Store に保存され、保存時に暗号化されている。
「アジェンティック AI を本番のデータパイプラインに導入する自信があったのは、『人間が舵を取る』という仕組みのおかげです。エージェントが重労働を担いますが、最終的な判断は人間が行います。すべての変更はエンジニアたちが普段使っているレビュープロセスを経るため、導入は即座に進みました」とロバーツ氏は語る。
影響
Data Accelerator は F1 の MarTech オペレーションにおいて、計測可能な大きな成果をもたらしました:
- データソースのオンボーディング:コード生成に要する時間が従来の 6〜8 週間から約 40 分に短縮され、その後は展開とレビューに数時間を要するのみ。
- 自律的な作業:AI エージェントがオンボーディングタスクの 95% を人間の介入なしで処理。
- 価値提供までの期間:約 99% の短縮を実現。
- スキーマ進化:解決にかかる時間が数日から数時間に短縮。
- インテグレーションのバックログ:18 ヶ月分の蓄積を数週間で解消。
以前はボイラープレート型のデータ取り込みコードの記述やスキーマの破綻対応に時間を費やしていたデータエンジニアたちは、今ではビジネスを推進する戦略的な業務に注力しています。
実装速度も劇的に向上し、単一の開発者がアジェンティック・ソリューションを概念証明(PoC)から本番リリースまでわずか 4 ヶ月で完了させました。
MarTech プラットフォームの信頼性、一貫性、データ整合性が改善され、運用オーバーヘッドは大幅に削減されました。ケンプ氏はこう語っています。「Data Accelerator は単に処理を速くしただけではありません。私たちの運用方法そのものを変えたのです。データエンジニアたちはもはやボイラープレート型のコード記述から解放され、戦略的な業務に集中できるようになりました。エンドユーザーが気づかないうちに問題を特定し、修正することが可能になっています。」
結論
Data Accelerator
原文を表示
Formula 1® (F1) engages an audience of over 800 million fans globally across digital platforms, F1 TV, social media, ticketing, and merchandise year-round. Races happen every two weeks. Fan engagement windows are measured in minutes and commercial decisions need to move at the speed of the grid. Behind the scenes, F1’s marketing technology (MarTech) platform, Customer 360, captures interactions across all of these touchpoints to power personalization, segmentation, and commercial strategy.
However, the platform faced a significant operational challenge. According to Chris Roberts, Director of IT at Formula 1, “Our MarTech platform is the nervous system of F1’s fan engagement. But every new data source required 6 to 8 weeks of manual engineering. We had an 18-month backlog just to integrate 12 new sources.” The business was generating data faster than the engineering team could wire it up. As a result, Matt Kemp, F1 Head of Data Operations, set to improve efficiencies and data quality. “Manually ingesting data sources is time consuming, creates solution variances, and ultimately results in data integrity issues. I wanted a solution that was repeatable, robust and reliable. AWS worked backwards from our needs to implement an agentic solution that worked end to end, applying business logic at each step.”
In early 2026, F1 and AWS worked together to build the Data Accelerator, a solution that uses agentic AI on Amazon Bedrock AgentCore to transform F1’s MarTech data platform from a manually maintained system into a self-managed, observable, and unified data estate. In this post, we show how the Data Accelerator reduced data source onboarding from up to 8 weeks to approximately 40 minutes of code generation plus hours of deployment. It also identified and fixed data source anomalies in production, tracked data platform operations and agent lineage in a single window, and opened a gateway for analysts, engineers, and scientists to collaborate. “For the first time, we have end-to-end visibility across the entire MarTech platform with data lineage and root cause analysis, not just dashboards full of alerts,” says Roberts.
The challenge
F1’s Customer 360 platform ingests data from ticketing partners, streaming integrations, sponsor activation feeds, social media, and merchandise systems. Operating a data estate of this breadth and velocity surfaced three areas of friction the team set out to solve. First, onboarding each new data source was a heavily manual effort: engineers wrote schema mappings, built ingestion pipelines, configured data quality checks, defined General Data Protection Regulation (GDPR) classifications, and set governance policies by hand. This process took 6 to 8 weeks per source. Second, the platform had to keep pace with constantly evolving upstream feeds. Providers frequently changed column names, added fields, or restructured and rescheduled payloads without notice. Those changes often surfaced at the worst possible moment, such as mid race-weekend or during a mission-critical campaign launch. Third, visibility was fragmented. Logs were scattered across services with no unified data lineage. When a stakeholder questioned a metric, engineers spent hours manually tracing the issue across Amazon Simple Storage Service (Amazon S3) paths, Amazon Redshift control tables, Airflow logs, and DBT outputs.
Solution overview
The Data Accelerator addressed these challenges through five workstreams delivered simultaneously:
- Agentic data source onboarding using Amazon Bedrock AgentCore, hosting agents in its runtime containers.
- Automated schema evolution detection and remediation.
- Unified data access through Amazon SageMaker Unified Studio.
- End-to-end observability with root cause analysis tool (RCA) and context graph.
- Automated identification of a failure in observability dashboard and agentic operation if they could be fixed with code changes.
A sixth workstream optimized the customer identity resolution algorithms that unify fan touchpoints across channels. The following sections describe each workstream in detail.
Agentic data source onboarding
The centerpiece of the Data Accelerator is a set of platform agents that take a Business Requirements Document (BRD) with limited information about the data source and produce a fully production-ready onboarding pipeline. This includes infrastructure code, data transformations, governance policies, and GDPR classification without a human writing a single line of boilerplate. The agents work in two phases:
Phase 1: Configuration generation
When a new data source needs onboarding, a team member uploads a BRD to an Amazon S3 bucket. The upload triggers an AWS Lambda function, which invokes Amazon Bedrock AgentCore Runtime, a capability of Amazon Bedrock AgentCore. The agent reads the BRD and generates a set of configuration files. It then accesses GitHub through a GitHub App to push these files as a pull request to the standardized Git repository, and accesses Jira through its REST API to create a ticket referencing the PR. All agent conversations and actions are traced in Amazon CloudWatch through built-in AgentCore observability. The assigned engineer reviews, adjusts if necessary, and approves.

Phase 1 workflow: a BRD upload triggers the agent to generate config files and open a pull request
Phase 2: Full pipeline generation
Once the configuration files are approved, a human triggers the next stage. The agent takes the approved configuration and generates three separate Pull Requests:
- AWS Glue application and infrastructure code.
- DBT transformation framework.
- Governance policies including GDPR tagging.
All three PRs link to a single Jira ticket for traceability. Engineers review each one across the Infrastructure, DBT, and Governance repositories and approve.

Phase 2 workflow: the agent generates infrastructure, transformation, and governance pull requests
Automated GDPR classification
What distinguishes this from a basic code generator is the integrated GDPR classification. The agent proactively analyzes every data column, determines whether it contains personal data, sensitive personal data, or pseudonymized data, and tags it with the appropriate GDPR category. These tags publish directly to the governance registry in SageMaker Unified Studio, giving the compliance team immediate visibility without manual review cycles.
Modular skill architecture
The system is not a tightly coupled agent graph. A single agent operates with modular skill definitions, each encapsulating a distinct capability: schema mapping and data type inference, data quality validation, governance enforcement, and sensitive data classification. At runtime, the agent evaluates incoming requirements and activates the relevant skills, composing them through a multi-pass reasoning process. Pass-0 handles token management through scrubbing, Pass-1 summarizes tool outputs, and Pass-2 rolls up an overall assessment, refining accuracy and completeness progressively rather than relying on a one-shot response. New capabilities ship as new skill modules without changing the core agent loop, keeping the architecture maintainable and composable as the platform grows.
The result is onboarding time dropped from 6 to 8 weeks to approximately 40 minutes of code generation plus hours of deployment and review. AI agents now handle 95% of the work autonomously.
Automated schema evolution
Onboarding new data sources is one challenge, but keeping existing integrations healthy is another. Upstream providers frequently modify their data structures, from renaming a column to creating a new field. Previously, the F1 team discovered these changes when a pipeline failed, often during a live race weekend. The same agent architecture that handles onboarding now continuously monitors for upstream schema changes. When a provider modifies their data structure, the agent detects it through event-driven triggers using AWS Lambda and Amazon EventBridge. It assesses the downstream impact, identifying which pipelines are affected, and which consumers depend on the changed fields. It then generates the necessary code updates across all affected repositories and creates a Jira ticket with full context and linked PRs. Engineers receive a notification that explains what changed, describes the impact, and presents a proposed fix for review. End-to-end resolution now takes hours instead of days.

Schema evolution agentic workflow
Unified data access with Amazon SageMaker Unified Studio
Before the Data Accelerator, working with Customer 360 data required navigating multiple disconnected environments. Data engineers curated pipelines in one account. Data scientists who wanted to model fan behavior needed access to a separate account, and analysts operated in a third world entirely. Nobody shared tooling or context, and getting from a question to an answer took days of coordination before any analysis could begin.
The solution uses Amazon SageMaker Unified Studio as the foundation for a data mesh framework where a central governance account brokers data discovery and access across multiple producer teams. The key enabler: governance is codified as declarative configuration, not manual console operations. A single data source definition simultaneously publishes data to the catalog and provisions the access control needed for consumers to subscribe. This means agents can safely onboard new data products end-to-end, from storage to catalog to governed access, because the framework enforces security constraints by construction. No human needs to review IAM policies or AWS Lake Formation grants. The platform guarantees correctness structurally. This is what makes the “one front door” possible.
Data engineers curate and govern datasets in one place, and data scientists find those same datasets in the same environment: governed, documented, and ready to model. A data scientist building a fan segmentation model or optimizing the customer identity algorithm doesn’t need to know where the data lives, who owns the pipeline, or which S3 prefix to use. They open Unified Studio, find the curated Customer 360 datasets, and start modeling. They get shared notebooks, consistent tooling, and governed access, because declarative governance made safe self-service possible without sacrificing control. The curation and the consumption finally live side by side.
End-to-end observability with RCA and context graph
A data platform is only as trustworthy as the team’s ability to answer one question: is the data correct right now? Before the Data Accelerator, answering that question meant logging into Apache Airflow, checking Amazon S3 paths, querying Amazon Redshift control tables, and reading DBT logs. “Nobody had the full view. When a stakeholder asked, ‘why does this number look wrong?’ the answer was always, ‘give us a few hours.’ The observability dashboard changes that entirely,” adds Roberts.
The observability layer presents full data lineage from S3 Raw ingestion through Processed layers into Amazon Redshift DBT stages as a single interactive graph, color-coded for health. Users click on any node to drill down to individual sources and tables, each showing pass/fail status, last run time, and duration. If a pipeline fails, the lineage visualization shows exactly where the break occurred, and which downstream data is affected.
Root cause analysis (RCA) is an agentic tool within F1’s platform that reads system logs and identifies failure points across the data estate. On its own, RCA can tell you what failed. We augment the RCA tool by passing through business context and system topology, codified as JSON. A missing file in S3 might be the error, but with the context graph, RCA tells you that the upstream provider rescheduled their delivery window, which is why the file wasn’t there when the pipeline ran. That’s the difference between knowing what failed and understanding why.
For the first time, F1 has full lineage, causal root cause analysis, and business context definitions in one place, and the dashboard auto-refreshes every 15 minutes.

Data lineage visualization showing pipeline health across sources and stages

Observability dashboard with failure details
Customer identity resolution
The final workstream optimized the algorithms that resolve customer identity across F1’s fan touchpoints. A single fan might interact through the app, buy tickets on the website, watch on F1 TV, and engage on social media. Unifying those interactions into a single identity without false merges or missed matches is what makes effective personalization possible within the Fan Personalization Platform (FPP).
F1 already had a working identity resolution process, but it was slow and struggled to scale with the growing volume of fan interactions across channels. Rather than re-architecting the pipeline or replacing components, the team focused on optimizing the existing resolution algorithm’s computational performance. By profiling execution bottlenecks and tuning the matching logic, the engagement reduced processing time by 50%, while keeping the entire resolution pipeline and its downstream integrations fully intact. No processes were changed, no accuracy trade-offs were made: the same algorithm now runs in half the time at F1’s production scale.
With faster resolution, F1 can onboard any new data source and gather new customer data in half the existing time. Faster resolution means fresher unified profiles, which in turn means more timely and relevant personalization across every marketing channel.
“The whole point is to deliver the right message to the right fan at the right time, whether that’s through email, the F1 app, ticketing, or social. Now that we can onboard sources in hours and resolve identities faster, we can actually deliver the personalized experiences our fans expect across every marketing channel,” says Kemp.
Security and governance by design
The Data Accelerator operates on the principle that AI proposes and humans review. The agents run on Amazon Bedrock AgentCore with long-term memory, retaining context across invocations. Development used Kiro for structured spec-driven development and Amazon Bedrock (Claude Sonnet 4.6) as the foundation model. The event-driven backbone uses AWS Lambda for compute, Amazon EventBridge for routing, Amazon Managed Workflows for Apache Airflow (MWAA) for workflow orchestration, and Amazon S3 as the raw data layer. All AI model access is governed through F1’s AI Gateway for unified access control, cost management, and audit logging. But the architecture is only half the story. The security posture is what makes this production-ready.
The security posture includes:
- Least privilege: fine-grained permissions, short-lived tokens with one-hour expiry, access limited to specific repositories and resources.
- Full audit trail: every action is logged and attributed for compliance.
- Human review: every generated Pull Request goes through engineer approval.
- Automated testing: agents generate comprehensive tests for their own changes.
- Rollback capabilities: issues surfaced post-merge can be reverted immediately.
- Network isolation: the entire system runs within private subnets in Amazon Virtual Private Cloud (Amazon VPC) with no direct internet access.
- Encrypted credentials: all secrets stored at rest in AWS Systems Manager Parameter Store.
“What gave us confidence to put agentic AI in our production data pipelines was what we call ‘Human at the helm.’ The agents do the heavy lifting, but humans make the decisions. Every change goes through the same review process our engineers already use, so adoption was immediate.” says Roberts.
The impact
The Data Accelerator delivered measurable impact across F1’s MarTech operations:
- Data source onboarding: reduced from 6 to 8 weeks to approximately 40 minutes of code generation plus hours of deployment and review.
- Autonomous work: AI agents handle 95% of onboarding tasks without human intervention.
- Time-to-value: approximately 99% reduction.
- Schema evolution: end-to-end resolution in hours instead of days.
- Integration backlog: 18-month backlog cleared in weeks.
- Data engineers who previously spent their time writing boilerplate ingestion code and chasing schema breaks now focus on strategic initiatives that advance the business.
- Implementation velocity: a single developer took the agentic solution from proof of concept to production release in 4 months.
The reliability, consistency, and data integrity of the MarTech platform were improved, while the operational overhead was reduced: “The Data Accelerator didn’t just speed things up. It changed how we operate. Our data engineers went from writing boilerplate ingestion code to focusing on strategic initiatives. Issues can be identified and fixed before our end users even notice.” says Kemp.
まとめ
The Data Accelerat
AI算出
導入事例ainew評価標準
記事は AWS の Agentic AI を活用した F1 のデータ運用改善事例に焦点を当てており、AI が主題であるため ai_relevance は 0.75。同クラスターに先行記事がないが、ベンダー導入事例としての性質が強いため novelty は 0.5 とし、検索機会も明確な製品名(Amazon Bedrock AgentCore)があるため 0.75。日本企業との直接的な関連性や独自データは乏しいため japan_relevance は 0.25。
6つの評価軸を見る
- AI関連度
- 75
- 情報源の信頼性
- 100
- 新規性
- 50
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み