Deepgram、Amazon SageMaker AI のサポートに AWS IAM 一時的委譲
Deepgram は Amazon SageMaker AI と連携し、パートナーが長期間のアクセス権限を保持せずとも迅速な診断と監査可能なサポートを提供する機能を導入した。
AI深層分析を開く2026年7月28日 02:07
AI深層分析
キーポイント
IAM Temporary Delegation の統合
Deepgram は新しい IAM 機能である「Temporary Delegation」を導入し、パートナーが顧客のアカウントに長期間の有効期限を持つロールを設けることなく、時間制限付きで特定のリソースへのアクセス権限を取得できるようにした。
サポートプロセスの劇的短縮
従来の顧客とパートナー間でスケジュール調整や画面共有を行う必要がなくなり、顧客が IAM コンソールで承認するだけで初期調査を開始できるため、対応時間が数日から数分に短縮された。
SageMaker AI による管理機能
Deepgram は Nova、Flux、Aura-2 モデルを AWS Marketplace で提供し、VPC 分離や自動スケーリングなどエンタープライズ向けの検証済みアーキテクチャと Terraform モジュールを SageMaker AI を介して利用可能にした。
セキュリティとコンプライアンスの維持
データ所在、ネットワーク分離、規制遵守を重視する企業が、マネージドクラウドサービスの運用効率を損なうことなく自前ホストを実現できる環境を整備した。
重要な引用
Deepgram has integrated with IAM temporary delegation, a new IAM capability that grants partners scoped, time-limited, customer-approved access to specific resources, with no long-lived credentials, no cross-account roles to provision, and no shared secrets.
With this integration, Deepgram has reduced the time for initial investigation on a SageMaker AI support ticket from days to minutes.
SageMaker AI is how we close that gap: a single AWS native control plane for deploying, scaling, and observing Deepgram speech models inside the customer's own account.
編集コメントを表示
編集コメント
Deepgram が AWS の最新 IAM 機能を即座に活用し、自前ホスト型 AI サービスのサポート課題を解決した点は実用的である。特に、顧客が権限付与を自己完結的に実行できる仕組みは、セキュリティ意識の高い組織にとって大きなメリットとなる。
自社で音声 AI を運用する企業にとって、迅速かつ監査可能なサポートを受けながら、パートナーに長期のアカウントアクセス権を付与しないことは重要です。モデルが予期せぬ結果を返したり、エンドポイントが異常な動作を示したりした場合、それを診断するのに最も適しているのは、多くの場合パートナー側のエンジニアです。
しかし、すべてのサポート対応のためにクロスアカウント AWS Identity and Access Management (IAM) ロールを設定するのは、運用コストが高く、監査のたびに議論が必要になるという課題があります。Amazon SageMaker AI は、自社でホストする Deepgram 音声 AI モデル向けの管理されたデプロイオプションを提供しています。Nova、Flux、Aura-2 の AWS Marketplace リスティングに加え、既存の AWS インフラストラクチャに重ねて使用できる検証済みリファレンスアーキテクチャと Terraform モジュールも用意されています。
Day 2 サポート(運用開始後のサポート)においても同様の運用成熟度を実現するため、Deepgram は IAM 一時的委譲 と連携しました。これは新しい IAM の機能で、パートナーにスコープを限定し、時間制限付きで、顧客が承認した特定のリソースへのアクセス権を付与します。長期の有効期限を持つ認証情報や、設定が必要なクロスアカウントロール、共有シークレットは不要です。
本稿では、Deepgram が IAM 一時的委譲(IAM Temporary Delegation)を採用した理由と、統合の全体像、そして SageMaker AI で Deepgram の音声モデルを実行する顧客にとって何が実現されるかについて解説します。この統合により、SageMaker AI に関するサポートチケットへの初回調査所要時間が「数日」から「数分」に短縮されました。従来は、顧客とパートナー双方のスケジュールを調整して画面共有を行う必要がありましたが、現在は顧客が自社の IAM コンソールでリクエストを承認するだけで済みます。これにより、長年の課題であったクロスアカウントアクセスの問題も解消されます。
Amazon SageMaker AI が提供する、セルフホスト型 Deepgram 向けの管理制御平面
エンタープライズ顧客が Deepgram のセルフホスティングを選択するのは、データ所在地の要件やネットワークの分離、規制への準拠といった一般的な理由によるものです。しかし、そのためにマネージドクラウドサービスとしての運用体制を犠牲にしたくはありません。このギャップを埋めるのが SageMaker AI です。これは、顧客自身の AWS アカウント内で Deepgram の音声モデルを展開し、スケーリングし、監視するための単一のネイティブな AWS 制御プレーンです。
Deepgram は、エンタープライズプラットフォームチームにとって最も重要な軸において、SageMaker AI を主要なデプロイ先として積極的に投資しています。
Deepgram の Nova、Flux、Aura 各モデル(音声認識とテキスト読み上げに対応)が AWS Marketplace にリストアップされ、ワンクリックでの購読と AWS 請求書の一元化が可能になりました。
Amazon VPC によるネットワーク分離や AWS PrivateLink、自動スケーリング、観測機能などを網羅した検証済み参照アーキテクチャも提供されています。
また、プラットフォームチームが既存の AWS インフラ上にレイヤーとして追加できる Terraform モジュールも用意されており、ゼロからデプロイ環境を構築する必要はありません。
これにより、「自力で対応するしかない」という自ホストのイメージは過去のものとなりました。
サポートアクセスの問題
本番環境での運用が長くなれば、必ずサポートが必要になります。モデルが予期せぬ結果を返す、エンドポイントの自動スケーリングが遅い、GPU の利用率がおかしいといった事態です。こうした問題を診断するのに最も適しているのは Deepgram チームのエンジニアですが、彼らが直面する課題は、アクセス権限を持たない顧客の Amazon VPC 内で動作するワークロードを検証する方法にあります。
従来の方法にはすべて欠点がありました。長期有効なクロスアカウント IAM ロールは機能しますが、顧客側ではプロビジョニングや監査、そして期限切れ時の取り消しを忘れないように管理する必要があり、プラットフォームチームも次の監査で責任を問われることを望んでいません。画面共有やログのコピー&ペーストは遅く、ミスが発生しやすく、すべてのコマンドに追跡可能性が求められる規制環境には適していません。顧客に Deepgram の代わりにコマンドを実行してもらう方法は些細な問題には有効ですが、反復作業を要する事象に対応するにはスケーラビリティがありません。
そこで必要だったのは、Deepgram のエンジニアに対して、特定の SageMaker AI エンドポイントへのアクセス権限を限定し、監査可能な形で、一定の時間だけ付与する方法です。承認は顧客がコントロールし、アカウント間の恒久的な信頼関係(スタンディング・トラスト)は不要としました。
前提条件
本記事で説明する統合を利用するには、以下の準備が必要です:
AWS アカウント内に、Amazon SageMaker AI 上で稼働中の Deepgram モデルのデプロイメントが必要です。
また、委任リクエストの確認と承認を行うための IAM パーミッション(iam:GetDelegationRequest および iam:AcceptDelegationRequest)も AWS アカウントに設定されている必要があります。
SageMaker AI エンドポイントが稼働する AWS リージョンでは、CloudTrail トレイルを有効化しておくことが条件です。これにより、委任された API 呼び出しが監査ログとして記録されます。
さらに、Deepgram のサポート契約または、サポートチケットシステムへのアクセス権を含むアカウントがアクティブであることも必須です。
なお、この統合を実行すると、SageMaker AI エンドポイントのホスティング(GPU インスタンス含む)、Amazon CloudWatch ログ、AWS CloudTrail、および関連するネットワークリソースに対して AWS からの課金が発生します。Deepgram モデル自体は 14 日間の無料トライアルを提供していますが、インフラコストはデプロイ開始と同時に発生することに注意してください。
IAM 一時的委任
IAM 一時的委任 は、まさにこの種の課題に対応するために新たに導入された IAM の機能です。
パートナーは、事前に登録されたパラメータ付きの権限テンプレートに対して委譲リクエストを提出します。顧客は自社の IAM コンソールでこれをレビューし、すべてのリソースが Amazon Resource Name (ARN) で明示され、ワイルドカードが含まれていない完全解決済みの権限を確認できます。その後、顧客はリクエストに承認を与えます。
AWS は直ちに、テンプレートにスコープを定め、セッション期間(SessionDuration)で制限し、エンドツーエンドの監査のために顧客の AWS CloudTrail にパートナーのアカウント ID をタグ付けした、短期有効な AWS Security Token Service (AWS STS) 認証情報をパートナーに発行します。
作成する IAM ロールも、回転させる長期有効なキーも不要です。維持すべき他アカウント間の信頼関係もありません。承認はセキュリティチームがすでに想定している場所、つまり顧客自身の IAM コンソール内で行われます。
Deepgram の統合の仕組み
Deepgram のエンジニアは、ワークフロー全体を通じてサポートチケットシステム内で作業を継続します。この統合は Deepgram のサポートチケットシステムに直接組み込まれており、単一の Amazon Simple Notification Service (Amazon SNS) トピックを介して AWS と接続されています。
エンドツーエンドの流れ(各工程には、Engineer、Customer、System のいずれのアクターが実行するかを示すラベルが付いています):
- [エンジニア] カスタマーサポートのチケットで、/delegate_access コマンドを実行してください。
- [顧客] チケット内の統合プロセスから指示が出たら、SageMaker AI エンドポイント ARN を返信してください。これはリクエスト全体の範囲を定義する唯一の情報です。ARN の取得には、aws sagemaker describe-endpoint --endpoint-name <エンドポイント名> --query EndpointArn コマンドを実行します。
- [システム] 統合側は、事前に登録された「DeepgramSageMakerReadOnlyTroubleshooting」権限テンプレートを使用して iam:CreateDelegationRequest を呼び出し、エンドポイントのリージョン、アカウント ID、名前をランタイムパラメータとして渡します。
- [顧客] 統合から送信される IAM コンソールへのダイレクトリンクを開いてください。リンクが動作しない場合は、IAM コンソールに手動でアクセスし、「Access management」→「Delegation requests」の順に進んでください。
- [顧客] IAM コンソール内で、解決済みの完全な権限内容を確認してください。
- [顧客] 「Approve(承認)」を選択して、Amazon CloudWatch ロググループ 1 つと DescribeEndpoint リソース 1 つへのアクセスを付与します。
- [システム] AWS は、交換トークンを Deepgram の Amazon SNS トピックに公開します。
- [システム] 統合側は sts:GetDelegatedAccessToken を使用して、このトークンを STS 認証情報と交換します。
- [システム] 取得した認証情報は、チケット内の内部ディスカッションへ投稿されます。
- [システム] 認証情報は自動的に 12 時間後に失効します。
- [顧客] 失効する前にいつでも、IAM コンソールで該当の委任リクエストを選択し「Revoke(取り消し)」を実行することでアクセスを取り消すことができます。
- [エンジニア] 発行された STS 認証情報を使用して aws sagemaker describe-endpoint --endpoint-name <エンドポイント名> を実行し、エンドポイントへのアクセスを確認してください。
- [エンジニア] 発行された STS 認証情報を使用して aws logs tail /aws/sagemaker/Endpoints/ を実行し、ログへのアクセスを確認してください。
- [顧客] Deepgram のパートナーアカウント ID でタグ付けされた AWS CloudTrail イベントを確認することで、委任されたアクティビティを検証できます。
図 1 — Amazon SageMaker AI 上の Deepgram サポートにおける、エンドツーエンドの IAM 一時的委譲フロー
お客様が体験する流れ
Deepgram のエンジニアがリクエストを開始すると、お客様には承認リンクが届きます。このリンクを開くと、お客様の IAM コンソール内の承認画面に直接移動します。特別なセットアップやロール作成は不要です。

図 2 — 承認画面では、誰がアクセスを求めているのか、その期間、理由、そして権限の平易な要約が表示されます。
詳細な権限内容を確認したい場合は、View JSON(JSON の表示)を選択します。これにより、すべてのアクションとリソース ARN が具体的に記述され、ワイルドカードが含まれない完全な解決済みポリシーを検証できます。権限テンプレートは定義時にパラメータ化されていますが、承認時にはワイルドカードは使用されません。

図 3 — JSON の表示では、完全なスコープを持つ IAM ポリシーが確認できます。
承認画面からは、Deny access(アクセス拒否)、Allow access(アクセス許可)、または Request approval(承認の依頼)の 3 つの選択肢があります。iam:AcceptDelegationRequest 権限を持つレビューアーが「アクセス許可」を選択すると、すぐにアクセスが可能になります。
レビュー担当者にその権限がない場合、承認をリクエストを選択し、事業上の理由を記載して申請します。すると、その要求はアカウント管理者に転送されます。管理者が承認すると、アクセスが可能になります。

*図 4 — 承認権限のないレビュー担当者は、理由を添えて管理者に要求を転送できます*
どちらの経路でも、最終的に同じ確認に至ります。Deepgram に一時的で期限付きのアクセスが付与され、そのすべての行動は顧客の AWS CloudTrail に記録されます。

*図 5 — スコープを限定し監査可能なアクセスが付与されたことの確認*
読み取り専用です。エンドポイントは一つ、リージョンは一つ、アカウントも一つ。有効期限は 12 時間です。委譲された資格情報を使用して行われた API コールは、CloudTrail が有効で関連する API アクションのログ記録が設定されている場合、Deepgram のパートナーアカウント ID でタグ付けされて顧客の AWS CloudTrail に表示されます。
なぜこれが重要なのか
Deepgram の顧客にとって、この統合により SageMaker AI へのサポートチケット対応で「双方の都合の良い時間にスクリーン共有を予約する」という待ち時間を、「IAM コンソールでリクエストを承認する」一瞬の作業に短縮できます。セキュリティ意識の高いエンタープライズ層の購入担当者が求めるのは、継続的な IAM ロールのプロビジョニングに関する議論ではなく、従来の方式よりもはるかに強力な「永続アクセスなし」の姿勢です。
クリーンアップ手順
テスト後に一時的なアクセス権限を解除し、継続的な課金を防ぐには、以下の手順を実行してください。
警告: 以下のクリーンアップ手順では、リソースとログデータが永久に削除されます。処理を開始する前に、保存したいログやデータのバックアップを行ってください。これらの操作は取り消すことができません。
IAM コンソールで「アクセス管理」に移動し、「委譲リクエスト」を開きます。
アクティブな Deepgram の委譲リクエストをすべて取り消してください。
SageMaker AI エンドポイントを削除します。コマンドは以下の通りです。
aws sagemaker delete-endpoint --endpoint-name .
エンドポイント設定も削除します。
aws sagemaker delete-endpoint-config --endpoint-config-name .
モデルを削除します。
aws sagemaker delete-model --model-name .
Amazon CloudWatch のロググループを削除して、保存料金の発生を防ぎます。
aws logs delete-log-group --log-group-name /aws/sagemaker/Endpoints/.
この統合のために AWS CloudTrail トレイルを作成した場合は、CloudTrail コンソールで削除してください。
モデルが不要になった場合は、AWS Marketplace の Deepgram リストから購読を解除します。
クリーンアップが完了したか確認しましょう。aws sagemaker list-endpoints を実行し、Deepgram のエンドポイントが表示されないことを確認します。エンドポイントを削除すればホスティング料金は停止しますが、ロググループを削除するまで保存料金は発生し続けます。
結論
Amazon SageMaker AI で音声 AI をセルフホストすることで、企業が必要とするデータ所在地の維持、ネットワークの分離、コンプライアンス体制を確立できます。AWS IAM Temporary Delegation は、これまでこの体制に伴っていた「2 日目以降のサポート」の課題を解消します。
これらを組み合わせることで、Deepgram のエンジニアは顧客のエンドポイントを数日ではなく数分で診断できるようになります。顧客が自社の IAM コンソールで承認した範囲限定かつ時間制限付きのアクセス権限により、顧客の AWS CloudTrail でエンドツーエンドの監査が可能となります。長期の有効期限を持つ認証情報や、他アカウント間のロール、共有シークレットは不要です。
AWS 上で SaaS パートナーへの IAM Temporary Delegation の展開が進むにつれ、Deepgram と AWS は、セルフホスト型音声 AI における「安全で摩擦の少ないサポート」をデフォルトとするための投資を継続していきます。
始め方
SageMaker AI は、Deepgram の音声 AI モデルをセルフホストでデプロイする際の推奨オプションです。また、「IAM Temporary Delegation」は、導入後の運用(Day Two)の体験を、初期導入時(Day One)と同じレベルに引き上げるための最新のアプローチです。
SageMaker AI で利用可能な Deepgram のすべてのモデル(Nova、Flux、Aura-2 など)には、追加費用なしで 14 日間のトライアルが含まれています。これにより、プラットフォームチームは本番導入を決定する前に、自社の AWS アカウント内で実際にデプロイを試すことができます。
SageMaker AI で Deepgram モデルを実行する場合、エンドポイントのホスティング(GPU インスタンスを含む)、Amazon CloudWatch のログ、および関連するネットワークリソースに対して課金が発生します。Deepgram 側では追加費用なしで 14 日間のトライアルを提供していますが、AWS インフラストラクチャのコストはデプロイ開始と同時に発生します。そのため、SageMaker AI の価格ページを確認し、AWS Cost Explorer を活用して支出を監視することが重要です。
始め方についてや詳細については、Deepgram または AWS の担当者までお問い合わせください。基盤となる機能の詳細については、AWS IAM Temporary Delegation のドキュメントをご覧ください。
執筆者について

Victor Wang
Deepgram の Staff Software Engineer であり、カリフォルニア州サンフランシスコを拠点にエンジニアリング担当バイスプレジデントの技術顧問も務めています。Deepgram 入社前は AWS で Sr. Solutions Architect、Technical Program Manager、Proserve Consultant、Software Developer など複数の役割を経験しました。新しい技術を学ぶことと世界中を旅することが彼の情熱です。これまでに 200 万マイル以上を飛行し、探求の旅をこれからも続けていく予定です。

Andre Gomes
AWS の Frontier AI Solutions Architect として、スタートアップ企業の大規模なトレーニングやリアルタイム推論を担当しています。AWS 上でワークロードをスケールさせ、世界中の顧客にリーチできるよう支援しています。AWS 入社前は UFMG で博士号を取得し、ケンブリッジ大学製造研究所で研究を行いました。その際、大規模言語モデルとニューラルトピックモデリングを活用してロードマップ策定に取り組んでいます。

ダニエル・ウィルジョ
ダニエルは AWS のソリューションアーキテクトで、AI と SaaS スタートアップに注力しています。元スタートアップの CTO として、創業者やエンジニアリングリーダーと協力し、AWS 上で成長とイノベーションを推進することに情熱を注いでいます。仕事以外では、コーヒー片手に散歩をしたり、自然を楽しんだり、新しいアイデアを学ぶことを好んでいます。

カリーム・シッド=モハメッド
カーリーは AWS のプロダクトマネージャーで、SageMaker HyperPod における生成 AI モデルの開発とガバナンスの実現に注力しています。以前は Amazon QuickSight で埋め込み型分析や開発者体験を主導し、AWS Marketplace やアマゾンリテールでもプロダクトマネージャーとして活躍しました。キャリアの初期にはコールセンターテクノロジーや Expedia のローカルエキスパート、広告分野の開発者として働き、マッキンゼーでは経営コンサルタントを務めました。
原文を表示
Enterprises running self-hosted speech AI need fast, auditable support without giving partners long-lived access to their accounts. When a model returns unexpected results or an endpoint misbehaves, the engineer best positioned to diagnose it is often on the partner side. But provisioning cross-account AWS Identity and Access Management (IAM) roles for every support engagement is operationally expensive and a recurring audit conversation. Amazon SageMaker AI delivers a managed deployment option for self-hosted Deepgram speech AI models. It provides AWS Marketplace listings for Nova, Flux, and Aura-2, along with validated reference architectures and Terraform modules that layer onto existing AWS infrastructure. To bring the same operational maturity to day-two support, Deepgram has integrated with IAM temporary delegation, a new IAM capability that grants partners scoped, time-limited, customer-approved access to specific resources, with no long-lived credentials, no cross-account roles to provision, and no shared secrets.
In this post, we cover why Deepgram built on IAM temporary delegation, how the integration works end-to-end, and what it unlocks for customers running Deepgram speech models on SageMaker AI. With this integration, Deepgram has reduced the time for initial investigation on a SageMaker AI support ticket from days to minutes. Previously, initial investigation required scheduling a screen-share across customer and partner calendars. Now a customer approves a request in their own IAM console, which also reduces long-standing cross-account access.
Amazon SageMaker AI delivers a managed control plane for self-hosted Deepgram
Enterprise customers choose to self-host Deepgram for the usual reasons: data residency, network isolation, and regulatory compliance. But they don’t want to trade away the operational posture of a managed cloud service to get there. SageMaker AI is how we close that gap: a single AWS native control plane for deploying, scaling, and observing Deepgram speech models inside the customer’s own account.
Deepgram has invested in SageMaker AI as a first-class deployment target across the axes that matter most to enterprise platform teams:
- AWS Marketplace listings for Deepgram’s Nova, Flux, and Aura spanning speech-to-text (STT) and text-to-speech (TTS) speech models, with one-click subscription and consolidated AWS billing.
- Validated reference architectures covering Amazon Virtual Private Cloud (Amazon VPC) isolation, AWS PrivateLink, auto scaling, and observability.
- Terraform modules that platform teams can layer onto their existing AWS infrastructure instead of building the deployment from scratch.
The result is a self-hosted experience that no longer means “you’re on your own.”
The support access problem
Production deployments eventually need support. A model is returning unexpected results, an endpoint is autoscaling slower than expected, graphics processing unit (GPU) utilization looks off. The engineer best positioned to diagnose it is on the Deepgram team. The challenge they face is how to investigate a workload running inside a customer Amazon VPC they don’t have access to.
The traditional options all have drawbacks. Long-lived cross-account IAM roles work, but customers don’t want to provision, audit, and remember to revoke them, and platform teams don’t want to be the ones answering for them at the next audit. Shared screens and copy-pasted logs are slow, error-prone, and a poor fit for regulated environments where every command needs to be attributable. Asking the customer to run commands on Deepgram’s behalf works for trivial issues, but doesn’t scale for issues that require iteration.
We needed a way to give a Deepgram engineer scoped, auditable access to exactly one SageMaker AI endpoint, for a bounded window, with the customer in control of the approval and no standing trust between accounts.
Prerequisites
To use the integration described in this post, you need the following:
- An active Deepgram model deployment on Amazon SageMaker AI in your AWS account.
- IAM permissions in your AWS account to review and approve delegation requests, including iam:GetDelegationRequest and iam:AcceptDelegationRequest.
- An AWS CloudTrail trail enabled in the AWS Region where the SageMaker AI endpoint runs, so delegated API calls are captured for audit.
- An active Deepgram support contract or account that includes access to the Deepgram support ticketing system.
- Awareness that running this integration incurs AWS charges for SageMaker AI endpoint hosting (including GPU instances), Amazon CloudWatch logs, AWS CloudTrail, and associated networking resources. Deepgram offers a 14-day trial at no additional cost for its models, but AWS infrastructure costs apply from the start of deployment.
IAM temporary delegation
IAM temporary delegation is a new IAM capability built for exactly this shape of problem.
The partner submits a delegation request against a pre-registered, parameterized permission template. The customer reviews it in their own IAM console and sees the fully resolved permissions, with every resource Amazon Resource Name (ARN) spelled out and no wildcards. The customer then approves the request. AWS then issues short-lived AWS Security Token Service (AWS STS) credentials to the partner, scoped to the template, bounded by SessionDuration, and tagged in the customer’s AWS CloudTrail with the partner’s account ID for end-to-end auditability.
No IAM roles to create. No long-lived keys to rotate. No cross-account trust to maintain. Approval lives where security teams already expect it: in the customer’s own IAM console.
How Deepgram’s integration works
Deepgram engineers keep working in the support ticketing system throughout the workflow. The integration is built directly into Deepgram’s support ticketing system, wired to AWS over a single Amazon Simple Notification Service (Amazon SNS) topic.
End-to-end flow (each step is labeled with the actor who performs it, whether Engineer, Customer, or System):
- [Engineer] In the customer support ticket, run the /delegate_access command.
- [Customer] When the integration prompts you in the ticket, reply with the SageMaker AI endpoint ARN, the single piece of information that scopes the entire request. To find the ARN, run aws sagemaker describe-endpoint --endpoint-name --query EndpointArn
- [System] The integration calls iam:CreateDelegationRequest with the pre-registered DeepgramSageMakerReadOnlyTroubleshooting permission template, passing the endpoint’s Region, account ID, and name as runtime parameters.
- [Customer] Open the deep link to your IAM console that the integration sends you. If the link doesn’t work, navigate to IAM console, then Access management, then Delegation requests.
- [Customer] In the IAM console, review the fully resolved permissions.
- [Customer] Choose Approve to grant access to exactly one Amazon CloudWatch log group and one DescribeEndpoint resource.
- [System] AWS publishes an exchange token to Deepgram’s Amazon SNS topic.
- [System] The integration exchanges the token for STS credentials using sts:GetDelegatedAccessToken.
- [System] The integration posts the credentials to the ticket’s internal discussion.
- [System] Credentials expire automatically after twelve hours.
- [Customer] Revoke access at any time before expiration by choosing Revoke on the delegation request in the IAM console.
- [Engineer] Verify endpoint access by running aws sagemaker describe-endpoint --endpoint-name with the issued STS credentials.
- [Engineer] Verify log access by running aws logs tail /aws/sagemaker/Endpoints/ with the issued STS credentials.
- [Customer] Verify delegated activity by reviewing AWS CloudTrail events tagged with Deepgram’s partner account ID.

*Figure 1 — End-to-end IAM Temporary Delegation flow for Deepgram support on Amazon SageMaker AI*
What the customer sees
When a Deepgram engineer initiates a request, the customer receives a link. Opening it brings them straight to the approval screen in their own IAM console. There’s no setup and no role to create.

*Figure 2 — The approval screen shows who is requesting access, for how long, why, and a plain-language summary of the permissions*
Customers who want to see the exact permissions, rather than the summary, choose View JSON to inspect the fully resolved policy, with every action and every resource ARN spelled out and no wildcards. The permission template is parameterized at definition time but not wildcarded at approval time.

*Figure 3 — View JSON shows the complete, fully scoped IAM policy*
From the approval screen, the customer has three choices: Deny access, Allow access, or Request approval. A reviewer who holds the iam:AcceptDelegationRequest permission can choose Allow access, and access begins immediately.
If the reviewer doesn’t have that permission, they choose Request approval, add a business justification, and the request is forwarded to their account administrator. After the administrator approves, access begins.

*Figure 4 — Reviewers without approval rights can forward the request, with justification, to an administrator*
Either path ends at the same confirmation: Deepgram has been granted temporary, time-bounded access, and every action it takes is recorded in the customer’s AWS CloudTrail.

*Figure 5 — Confirmation that scoped, audited access has been granted*
Read-only. One endpoint. One Region. One account. Twelve hours. API calls made using the delegated credentials appear in the customer’s AWS CloudTrail tagged with Deepgram’s partner account ID, when CloudTrail is enabled and configured to log the relevant API actions.
Why this matters
For Deepgram customers, the integration collapses time-to-first-look on a SageMaker AI support ticket from “schedule a screen-share when both sides are free” to “approve a request in the IAM console.” For Deepgram’s security-conscious enterprise buyers, it replaces a recurring IAM role provisioning conversation with a no-standing-access posture that’s strictly stronger than the alternative.
Clean up
To avoid ongoing charges and remove temporary access after testing, complete the following steps:
Warning: The following cleanup steps permanently delete resources and log data. Back up any logs or data that you want to retain before proceeding. These operations cannot be undone.
- In the IAM console, navigate to Access management, then go to Delegation requests.
- Revoke any active Deepgram delegation requests.
- Delete the SageMaker AI endpoint: aws sagemaker delete-endpoint --endpoint-name .
- Delete the endpoint configuration: aws sagemaker delete-endpoint-config --endpoint-config-name .
- Delete the model: aws sagemaker delete-model --model-name .
- Delete the associated Amazon CloudWatch log group to stop log storage charges: aws logs delete-log-group --log-group-name /aws/sagemaker/Endpoints/.
- If you created an AWS CloudTrail trail specifically for this integration, delete it in the CloudTrail console.
- If you no longer need the model, unsubscribe from the Deepgram listing on AWS Marketplace.
- Verify cleanup by confirming that aws sagemaker list-endpoints returns no Deepgram endpoints. Endpoint hosting charges stop once the endpoint is deleted. Log storage charges continue until the log group is removed.
まとめ
Self-hosted speech AI on Amazon SageMaker AI supports the data residency, network isolation, and compliance posture enterprises need. AWS IAM Temporary Delegation helps close the day-two support gap that previously came with that posture. Together, they let Deepgram engineers diagnose customer endpoints in minutes instead of days, with scoped, time-bounded access that customers approve in their own IAM console and audit end-to-end in the customer’s AWS CloudTrail, with no long-lived credentials, no cross-account roles, and no shared secrets. As IAM temporary delegation expands across more SaaS partners on AWS, Deepgram and AWS will keep investing in making secure, low-friction support the default for self-hosted speech AI.
Get started
SageMaker AI is a preferred deployment option for self-hosted Deepgram speech AI models, and IAM Temporary Delegation is the latest investment in making the day-two experience match the day-one experience. Every Deepgram model available on SageMaker AI including Nova, Flux, and Aura-2 comes with a 14-day trial at no additional cost, so platform teams can stand up a real deployment in their own AWS account before committing. Running Deepgram models on SageMaker AI incurs charges for endpoint hosting (including GPU instances), Amazon CloudWatch logs, and associated networking resources. While Deepgram offers a 14-day trial at no additional cost, AWS infrastructure costs apply from the start of deployment, so review the SageMaker AI pricing page and use AWS Cost Explorer to monitor your spending.
To get started or to learn more, reach out to your Deepgram account representative or your AWS account representative. To learn more about the underlying capability, see the AWS IAM Temporary Delegation documentation.
About the authors

Victor Wang
Victor is a Staff Software Engineer at Deepgram and Technical Advisor to the VP of Engineering based in San Francisco, CA. Before joining Deepgram, Victor held multiple roles at AWS including Sr. Solutions Architect, Technical Program Manager, Proserve Consultant, and Software Developer. His passion is learning new technologies and traveling the world. Victor has flown over 2 million miles and plans to continue his eternal journey of exploration.

Andre Gomes
Andre is a Frontier AI Solutions Architect at AWS, working with startups on large-scale training and real time inference. He helps startups scale their workloads on AWS and reach customers worldwide. Before joining AWS, Andre earned his PhD at UFMG, with research conducted at the Institute for Manufacturing, University of Cambridge, applying large language models and neural topic modeling to roadmapping.

Daniel Wirjo
Daniel is a Solutions Architect at AWS, focused on AI and SaaS startups. As a former startup CTO, he enjoys collaborating with founders and engineering leaders to drive growth and innovation on AWS. Outside of work, Daniel enjoys taking walks with a coffee in hand, appreciating nature, and learning new ideas.

Kareem Syed-Mohammed
Kareem is a Product Manager at AWS. He is focuses on enabling GenAI model development and governance on SageMaker HyperPod. Earlier, at Amazon Quick Sight, he led embedded analytics, and developer experience. He has also been with AWS Marketplace and Amazon retail as a Product Manager. Kareem started his career as a developer for call center technologies, Local Expert and Ads for Expedia, and management consultant at McKinsey.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み