AWS GovCloud (US) で NVIDIA Nemotron および OpenAI GPT OSS モデルを Amazon Bedrock で実行可能に
本文の状態
日本語全文を表示中
詳細モードで約17分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
AWS Machine Learning Blog
AWS が、政府機関向けクラウド環境である AWS GovCloud (US) において、NVIDIA の Nemotron と OpenAI の GPT オープンソースモデルを Amazon Bedrock で利用できるようにした。これにより、機密性の高い任務でも最新の AI 能力を安全に活用できる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AWS GovCloud (US) でワークロードを実行する政府機関は、民間部門と同等の AI 機能を必要としています。同時に、任務に不可欠なセキュリティおよびコンプライアンス制御を妥協することはできません。オープンウェイト基盤モデル(FMs)が実験段階からミッションシステムへと移行するにつれ、あらゆるモデル決定を形作る二つの要件があります。第一に、そのモデルは任務が求める能力を提供できなければなりません。第二に、推論環境は機関のセキュリティ、コンプライアンス、およびデータ所在地に関する義務を満たす必要があります。米国政府機関、防衛・情報コミュニティ、そしてそれらを支える契約業者にとって、これらの要件は妥協の余地がありません。高度なオープンウェイトモデルへのアクセスは、インテリジェンス分析、ミッション計画、調達および契約文書のレビュー、セキュリティログ分析、コンプライアンス自動化などの業務に不可欠です。このアクセスを実現する際、機密データをその管理境界の外へ移動させることは許されません。
AWS GovCloud (US) で、米国発の最先端オープンウェイトモデルを導入できることをお知らせします。今回のリリースにより、Amazon Bedrock は OpenAI のオープンウェイト GPT OSS モデル(120B および 20B)と NVIDIA Nemotron モデル(Nano 9B v2, Nano 12B v2, Nano 30B, Super 120B)をサポートするようになりました。これらの新モデルを活用することで、多様で高性能なファウンデーションモデル(FM: Foundation Model)を用いて生成 AI アプリケーションの構築とスケーリングが可能になります。これにより、単一の統一 API を通じて、OpenAI と NVIDIA の最新モデルを他の主要 AI モデルと共に柔軟に利用できます。この統一 API を用いれば、アプリケーションコードを変更することなく、各特定のユースケースに適したモデルを選択して利用することが可能です。
AWS GovCloud (US) は、機密データや規制対象のワークロードをホストするために設計された、隔離された AWS リージョン群です。これらのリージョンは米国国内に物理的に位置し、米国人のみによって管理されています。これにより、顧客は FedRAMP High(暫定運用許可)および DoD クラウドコンピューティングセキュリティ要件ガイド(SRG: Security Requirements Guide)のインパクトレベル 2, 4, 5 といったコンプライアンスフレームワークへの適合を支援されます。さらに、国際武器取引規制(ITAR: International Traffic in Arms Regulations)や刑事司法情報サービス(CJIS: Criminal Justice Information Services)などの追加的なフレームワークにも対応しています。
Amazon Bedrock は、独立したモデルプロバイダーからファウンデーションモデル(FM)にアクセスするための完全マネージド型サービスであり、推論処理はすべて AWS が運営するインフラ上で実行されます。
Amazon Bedrock を利用すると、推論は AWS GovCloud (US) の分離境界内、すなわち米国人が運営する米国国内のインフラ上で実行されます。Amazon Bedrock がデータをどのように取り扱うかについては、Amazon Bedrock におけるデータ保護 をご参照ください。
OpenAI のオープンウェイト GPT OSS モデルと NVIDIA Nemotron オープンウェイトモデルが、AWS GovCloud (US) 内の Amazon Bedrock で利用可能になりました。今回のローンチにより、2 つのオープンウェイトモデルファミリーが AWS GovCloud (US) リージョンに導入されました。具体的には、OpenAI の gpt-oss-120b および gpt-oss-20b、そして NVIDIA Nemotron 3 ファミリー(Nemotron 3 Super 120B と Nemotron 3 Nano モデルを含む)です。これらのモデルを使用すれば、自動セキュリティ制御評価、複数文書インテリジェンスの統合、契約および買収分析、ポリシーコンプライアンスチェックといった、エージェント型アプリケーションやミッションワークフローを構築できます。これらすべてが AWS GovCloud (US) のコンプライアンス境界内で実行されます。
本記事では、AWS GovCloud (US) で現在利用可能なモデルとその機能、データ所在に関する推論オプション、利用可能なサービスティア、および開始方法について解説します。
モデルについて
このセクションでは、AWS GovCloud (US) で新たに利用可能になった 2 つのオープンウェイトモデルファミリーと、それぞれを特徴づける機能を紹介しています。
NVIDIA Nemotron
NVIDIA Nemotron ファミリーは、計算効率と専門的なエージェント AI システムにおける精度を重視して構築された、小型言語モデル (SLM) および大規模言語モデル (LLM) の両方の機能を提供します。NVIDIA は 2 つのモデルについて以下のように説明しています。
- NVIDIA Nemotron 3 Super は、複雑なマルチエージェントワークロード向けの 120B オープンハイブリッド混合専門家 (MoE: Mixture of Experts) モデルで、総パラメータ数は 1,200 億ですが、トークンごとに活性化されるのは約 120 億パラメータのみです。この MoE デザインにより、前世代と比較してコスト効率の高い推論処理で最大 5 倍のスループットを実現し、100 万トークンのコンテキストウィンドウによって、エージェントは長期的な記憶を保持しながら、長く多段階のタスクに集中して取り組むことができます。
- NVIDIA Nemotron 3 Nano は、約 300 億パラメータを持つオープンモデルで、トークンごとに約 30 億パラメータが活性化され、前世代と比較してスループットを最大 4 倍向上させるとともに、推論用トークンの生成量を最大 60% 削減します。100 万トークンのコンテキストウィンドウにより、長時間実行される多段階のエージェントワークフローをサポートします。
AWS GovCloud (US) で利用可能な NVIDIA Nemotron モデルの完全なリストについては、Amazon Bedrock 上の NVIDIA モデルをご参照ください。
OpenAI GPT OSS
OpenAI の GPT OSS モデルは、推論、エージェントタスク、開発者向けタスクに設計されたオープンウェイトのテキスト入力・テキスト出力モデルで、推論に必要な努力レベルを調整可能であり、外部ツールとの統合もサポートしています。本記事では、以下の 2 つの変種に焦点を当てます。
gpt-oss-120bは、OpenAI が生産環境向け・汎用型・高度な推論用途を目的として設計した、パラメータ数 1200 億のオープンウェイトモデルです。
gpt-oss-20bは、低レイテンシやローカル・専門的なユースケースを対象とした、パラメータ数 200 億のモデルです。
両モデルとも 128K トークンのコンテキストウィンドウと最大 16K トークンの出力トークンを提供し、テキスト入力を受け付けてテキスト出力を生成します。重みがオープンであるため、組織は独自にモデルアーキテクチャを評価したり、公開されたモデルカードを確認したり、代表的なワークロードで独自のベンチマークを実行したりできます。政府チームにとっては、この透明性が組織的なリスク評価を支え、顧客のセキュリティチームが展開前にモデルの挙動を評価することを可能にし、米国の多くの政府機関が採用しているゼロトラスト原則とも整合します。
AWS GovCloud (US) で利用可能な OpenAI モデルの完全なリストについては、Amazon Bedrock 上の OpenAI モデルをご参照ください。
コンプライアンス境界内でのサーバーレス推論
Amazon Bedrock 上の NVIDIA Nemotron および GPT OSS モデルは、Amazon Bedrock の次世代推論エンジンによって提供されています。アーキテクチャを理解するためには、エンジンとエンドポイントを区別することが役立ちます。エンジンは、モデル展開アカウントの分離とゼロオペレーターアクセスを設計理念とする基盤となるサービングインフラストラクチャであり、一方、bedrock-mantle エンドポイントは、アプリケーションがリクエストを送信するために呼び出す OpenAI 互換の HTTPS API です。政府機関にとっては、プロビジョニングするインフラストラクチャも管理する GPU もなく、モデル展開に関する専門知識も必要ありません。
次世代推論エンジンは、ゼロオペレーターアクセス設計に基づいて構築されています。AWS、顧客、またはモデルプロバイダーを問わず、いかなるオペレーターも、推論プロンプトや完成結果などの顧客データにアクセスできません。これに AWS GovCloud (US) の分離境界を組み合わせることで、政府チームには強力なデータ保護の基盤が提供されます。技術的な詳細については、Mantle のゼロオペレーターアクセス設計を探る を参照してください。
Amazon Bedrock は、これらのモデルを呼び出すための 2 つのエンドポイントを提供しています。bedrock-mantle エンドポイントは次世代推論エンジン向けの OpenAI 互換 API であり、OpenAI の Python および TypeScript SDK を使用して呼び出すことができます。このエンドポイントでは Chat Completions API と Responses API が利用されます。一方、bedrock-runtime エンドポイントは AWS SDK を介して Converse API と InvokeModel API を使用し、Guardrails などのネイティブな Amazon Bedrock の機能へのアクセスも可能です。両方のコードサンプルは「Getting started」セクションに記載されています。
リージョン別利用状況とデータ所在地
Amazon Bedrock では、推論リクエストが処理される場所について複数のオプションを提供しています。「In-Region(リージョン内)」ではすべてのリクエストを単一のリージョン内に保持し、「Geographic Cross-Region inference(地理的クロスリージョン推論)」では、より高いスループットを実現するために、同じ地理区域内の異なるリージョン間でリクエストをルーティングします。これにより、データはその地理的境界域内に留まります。AWS GovCloud (US) における NVIDIA Nemotron および GPT OSS モデルの場合、利用可能なオプションは以下の通りです:
- In-Region inference は us-gov-west-1 (AWS GovCloud (US-West)) で利用可能です。
- Geo cross-Region inference は、us-gov-west-1 と us-gov-east-1 の間でリクエストをルーティングする専用 AWS GovCloud (US) 跨リージョン推論 ID を通じて利用可能です。トラフィックは AWS GovCloud (US) の境界内に留まりながら、両リージョンにわたる耐障害性を獲得できます。
これらのモデルに関するすべての推論は、AWS GovCloud (US) の境界内で行われます。世界中の商用 AWS リージョン間でリクエストをルーティングするグローバル跨リージョン推論は、AWS GovCloud (US) では利用できません。要件に応じて、単一リージョンと Geo 跨リージョンのいずれかを選択できます。
サービスティア
Amazon Bedrock は、さまざまなワークロード要件に合わせるために複数の サービスティア を提供しています。3 つのモデルすべてにおいて、Standard (標準)、Priority (優先)、Flex (フレックス) の各ティアがサポートされています。
サービスティア
説明
サポート状況
Standard
トークン課金方式でコミットメント不要
はい
Priority
レイテンシに敏感なトラフィック向けにスループット向上
はい
Flex
柔軟かつ時間制約のないワークロード向けに低コストアクセス
はい
Reserved
期間コミットメントによる専用スループット
現在利用不可
デフォルトでは、リクエストはオンデマンド推論(Standard タイア)を使用し、事前に容量を予約することなくトークン数に応じて課金されます。レイテンシが重要な顧客向けワークロードの場合、個別のリクエストを優先度が高い Tier(Priority tier)にルーティングできます。時間制約のないモデル評価やバッチ要約などの作業には、コストを抑えられる Flex タイアが用意されています。スケーリングのガイドラインや本番環境でのレート制限への対応方法については、スケーリングとスループットのベストプラクティス および「はじめに」セクションをご参照ください。
AWS GovCloud (US) での始め方
このセクションでは、推奨される bedrock-mantle エンドポイントからモデルを呼び出す手順を説明します。例では、リージョン内推論が利用可能な us-gov-west-1 リージョンを使用しています。
コンソールプレイグラウンド
- AWS GovCloud (US) アカウントの Amazon Bedrock コンソールに移動します。
- 左側のメニューにある「Test」セクションから「Playground」を選択します。
- 「モデルの選択」をクリックします。
- カテゴリリストからプロバイダー(NVIDIA または OpenAI)を選択し、その後モデルを選択します(例:NVIDIA Nemotron 3 Super または gpt-oss-120b)。
- 「適用」をクリックしてモデルを読み込みます。
- モデルのテスト用のプロンプトを入力します。
bedrock-mantle エンドポイントの使用(推奨)
これらのモデルを使用するには、Amazon Bedrock モデルを呼び出す権限を持つ AWS GovCloud (US) 内の AWS アカウントが必要です。bedrock-mantle エンドポイントを利用する場合は、Amazon Bedrock API キー または標準的な AWS 認証情報が必要です。以下にサンプルポリシーを示します:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "BedrockMantleInference",
"Effect": "Allow",
"Action": [
"bedrock-mantle:CreateInference",
"bedrock-mantle:Get*",
"bedrock-mantle:List*"
],
"Resource": "arn:aws-us-gov:bedrock-mantle:us-gov-west-1:111122223333:project/*"
},
{
"Sid": "BedrockMantleCallWithBearerToken",
"Effect": "Allow",
"Action": "bedrock-mantle:CallWithBearerToken",
"Resource": "*"
}
]
}
翻訳全文
111122223333 を AWS アカウント ID に置き換え、リージョンをあなたが使用する AWS GovCloud (US) リージョンにスコープしてください。本記事のコード例は Bedrock API キーを使用して認証を行いますが、これには bedrock-mantle:CallWithBearerToken の権限が必要です。このアクションは、2 番目のステートメントに示されているように "Resource": "*" にスコープする必要があります。Amazon Bedrock API キーを生成または使用できるアイデンティティを制御するには、Amazon Bedrock API キーの生成および使用に関する権限の管理 を参照してください。組織を承認されたモデルのみを使用するように制限する場合は、サービスコントロールポリシー (SCP) を使用してください。
以下の例では OpenAI Python SDK を使用して bedrock-mantle エンドポイントを呼び出します。本番環境のワークロードでは、自動的に期限切れ(最大 12 時間)し、生成された IAM ロールの権限を継承する短期間 API キーを使用してください。
import boto3
from openai import OpenAI
# AWS Secrets Manager から Bedrock API キーを取得
secrets_client = boto3.client("secretsmanager", region_name="us-gov-west-1")
api_key = secrets_client.get_secret_value(SecretId="bedrock-api-key")["SecretString"]
client = OpenAI(
# ベース URL に AWS GovCloud (US) リージョンを使用、例:us-gov-west-1
base_url="https://bedrock-mantle.us-gov-west-1.api.aws/v1",
api_key=api_key,
)response = client.chat.completions.create(
model="openai.gpt-oss-120b",
messages=[
{"role": "user", "content": "Explain the benefits of open-weight models for regulated workloads."}
],
reasoning_effort="medium", # low | medium | high
max_completion_tokens=512,
)
print(response.choices[0].message.content)
注釈: これらの例では、Bedrock API キーをAWS Secrets Managerから取得しています。ローカル開発環境では、代わりに環境変数からキーを読み取ることも可能ですが、本番環境ではこのパターンは避けてください。AWS Secrets Manager または他のシークレットストアを使用してください。
NVIDIA Nemotron 3 Super 120B を呼び出す場合は、model パラメータを nvidia.nemotron-super-3-120b に変更し、reasoning_effort パラメータを削除してください(推論努力度の制御は GPT OSS に固有の機能です)。それ以外のコード変更は不要です。
推論努力度の制御
GPT OSS モデルは、調整可能な推論努力度を備えた推論モデルです。Chat Completions 呼び出しで reasoning_effort パラメータを low(低)、medium(中)、high(高)のいずれかに設定することで、応答レイテンシと推論深度のトレードオフを制御できます。大量かつレイテンシに敏感なトラフィックには low を、複雑な多段階推論やエージェントによる計画には high を使用してください。推論モデルにおいては、応答長を制限するために max_completion_tokens を優先して使用してください(古い max_tokens フィールドも引き続き受け付けられます)。
レスポンス API の利用
チャット完了機能に加えて、GPT OSS モデルはレスポンス API もサポートしています。これは OpenAI が提供する推論スタイルの対話のためのインターフェースです。メッセージ配列ではなく、単一の入力を受け取ります。NVIDIA Nemotron 3 Super 120B はレスポンス API をサポートしていません。そのモデルには Chat Completions(チャット完了)、Converse(会話)、または Invoke(呼び出し)を使用してください。
import boto3
from openai import OpenAI
AWS Secrets Manager から Bedrock API キーを取得する
secrets_client = boto3.client("secretsmanager", region_name="us-gov-west-1")
api_key = secrets_client.get_secret_value(SecretId="bedrock-api-key")["SecretString"]
client = OpenAI(
base_url="https://bedrock-mantle.us-gov-west-1.api.aws/v1",
api_key=api_key,
)
response = client.responses.create(
model="openai.gpt-oss-120b",
input="規制対象ワークロードにおけるオープンウェイトモデルの利点を説明してください。",
)
print(response)
ストリーミングレスポンス
トークンを生成するたびにユーザーに表示したいチャットやエージェントユースケースでは、stream=True を設定します。応答は、逐次的なデルタイベントのイテレータになります:
stream = client.chat.completions.create(
model="openai.gpt-oss-120b",
messages=[
{"role": "user", "content": "エキスパート混合アーキテクチャの短い要約を書いてください。"}
],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
print()
bedrock-runtime エンドポイントにおいて、同等の機能には bedrock:InvokeModelWithResponseStream 権限が必要ですが、後ほど示す最小限のポリシーですでに付与されています。
ツール呼び出し
NVIDIA Nemotron および GPT OSS オープンウェイトモデルはエージェントワークフロー向けに設計されており、ツール呼び出しシナリオにおいて実用的な機能を有しています。ツール呼び出しワークフローでは、モデルが呼び出すことができる関数(ツール)を定義し、モデルはユーザーの要求に基づいて呼び出しのタイミングを判断します。その後、アプリケーション側でその関数を実行して結果を取得し、それをモデルに返すことで、最終的な回答に組み込ませます。
以下の例はこのパターンをエンドツーエンドで示しています。get_weather というツールを定義し、ユーザーメッセージを送信し、モデルがツール呼び出しを要求するのを待ち、モックデータを用いて関数を実行した上で結果を戻し、モデルが自然言語による回答を生成できるようにします。
import json
import boto3
from openai import OpenAI
# AWS Secrets Manager から Bedrock API キーを取得
secrets_client = boto3.client("secretsmanager", region_name="us-gov-west-1")
api_key = secrets_client.get_secret_value(SecretId="bedrock-api-key")["SecretString"]
client = OpenAI(
base_url="https://bedrock-mantle.us-gov-west-1.api.aws/v1",
api原文を表示
Government agencies running workloads in AWS GovCloud (US) need AI capabilities that keep pace with the commercial sector. At the same time, they can’t compromise the security and compliance controls their missions require. As open-weight foundation models (FMs) move from experimentation into mission systems, two requirements shape every model decision. First, the model must deliver the capability the mission demands. Second, the inference environment must satisfy the agency’s security, compliance, and data residency obligations. For U.S. government agencies, the defense and intelligence community and the contractors that serve them, these requirements are non-negotiable. Access to advanced open-weight models is essential for work such as intelligence analysis, mission planning, acquisition and contract document review, security log analysis, and compliance automation. This access must not require moving sensitive data outside the boundary that governs it.
We’re excited to introduce US-based frontier open-weight models in AWS GovCloud (US). With this release, Amazon Bedrock now supports OpenAI’s open-weight GPT OSS models (120B and 20B) and NVIDIA Nemotron (Nano 9B v2, Nano 12B v2, Nano 30B, Super 120B) models. With these new models, you can build and scale generative AI applications with diverse, high-performance FMs. This offers the flexibility to use OpenAI’s and NVIDIA’s latest models alongside other leading AI models through a single, unified API. You can use this unified API to select the right model for each specific use case without changing your application code.
AWS GovCloud (US) provides an isolated set of AWS Regions designed to host sensitive data and regulated workloads. Regions are physically located in the United States and administered exclusively by U.S. citizens. They help customers meet compliance frameworks including FedRAMP High (Provisional Authority to Operate) and DoD Cloud Computing Security Requirements Guide (SRG) Impact Levels 2, 4, and 5. Additional frameworks include International Traffic in Arms Regulations (ITAR) and Criminal Justice Information Services (CJIS).
Amazon Bedrock is a fully managed service for accessing FMs from independent model providers, with inference running entirely on AWS-operated infrastructure.
With Amazon Bedrock, inference runs inside the AWS GovCloud (US) isolation boundary, on infrastructure operated by U.S. citizens on U.S. soil. For details on how Amazon Bedrock handles your data, refer to Data protection in Amazon Bedrock.
OpenAI’s open-weight GPT OSS models and NVIDIA Nemotron open-weight models are now available on Amazon Bedrock in AWS GovCloud (US). This launch delivers two open-weight model families into the AWS GovCloud (US) Regions: OpenAI gpt-oss-120b and gpt-oss-20b, and the NVIDIA Nemotron 3 family, including Nemotron 3 Super 120B alongside the Nemotron 3 Nano models. With these models, you can build agentic applications and mission workflows such as automated security control assessments, multi-document intelligence synthesis, contract and acquisition analysis, and policy compliance checking. All of this runs within the AWS GovCloud (US) compliance boundary.
In this post, we cover the models currently available in AWS GovCloud (US) and their capabilities, the inference options for data residency, the available service tiers and how to get started.
About the models
This section introduces the two open-weight model families now available in AWS GovCloud (US) and the capabilities that set each apart.
NVIDIA Nemotron
The NVIDIA Nemotron family delivers both small language model (SLM) and large language model (LLM) capabilities, built for compute efficiency and accuracy in specialized agentic AI systems. NVIDIA describes the two models as follows:
- NVIDIA Nemotron 3 Super is a 120B open hybrid mixture-of-experts (MoE) model for complex multi-agent workloads with 120 billion total parameters that activates only 12 billion parameters per token. This MoE design delivers up to 5 times higher throughput than the previous generation for cost-efficient inference, and its 1-million-token context window gives agents the long-term memory to stay focused across long, multi-step tasks.
- NVIDIA Nemotron 3 Nano is a 30-billion-parameter open model that activates approximately 3 billion parameters per token, delivering 4 times higher throughput than the previous generation and reducing reasoning-token generation by up to 60 percent. Its 1-million-token context window supports long-running, multi-step agent workflows.
For the full list of NVIDIA Nemotron models available in AWS GovCloud (US), refer to NVIDIA models on Amazon Bedrock.
OpenAI GPT OSS
OpenAI’s GPT OSS models are open-weight, text-to-text models designed for reasoning, agentic, and developer tasks, with adjustable reasoning effort and support for external tool integration. This post focuses on two variants:
gpt-oss-120b is OpenAI’s 120-billion parameter open-weight model, designed for production, general-purpose, and high-reasoning use cases.
gpt-oss-20b is the 20-billion parameter model, designed for lower latency and local or specialized use cases.
Both models provide a 128K-token context window and up to 16K output tokens, and both accept text input and generate text output. Because the weights are open, organizations can independently evaluate the model architecture, review the published model card, and run their own benchmarks on representative workloads. For government teams, this transparency supports organizational risk assessments, enables customer security teams to evaluate model behavior before deployment, and aligns with the zero-trust principles many U.S. government agencies are adopting.
For the full list of OpenAI models available in AWS GovCloud (US), refer to OpenAI models on Amazon Bedrock.
Serverless inference inside your compliance boundary
NVIDIA Nemotron and GPT OSS models on Amazon Bedrock are served by the next-generation inference engine in Amazon Bedrock. To understand the architecture, it helps to distinguish between the engine and the endpoint: the engine is the underlying serving infrastructure, designed with Model Deployment Account isolation and zero operator access, while the bedrock-mantle endpoint is the OpenAI-compatible HTTPS API that applications call to send requests to the engine. For agencies, there’s no infrastructure to provision, no GPUs to manage, and no model-deployment expertise required.
The next-generation inference engine is built on a zero operator access design. No operator, whether from AWS, the customer, or a model provider, can access customer data, such as inference prompts or completions. Combined with the AWS GovCloud (US) isolation boundary, this gives government teams a strong data-protection foundation. For the technical details, refer to Exploring the zero operator access design of Mantle.
Amazon Bedrock provides two endpoints for invoking these models. The bedrock-mantle endpoint is the OpenAI-compatible API for the next-generation inference engine, so you can call it with the OpenAI Python and TypeScript SDKs. It uses the Chat Completions and Responses APIs. The bedrock-runtime endpoint uses the Converse and InvokeModel APIs through the AWS SDK, with access to native Amazon Bedrock features such as Guardrails. Code samples for both are in the Getting started section.
Regional availability and data residency
Amazon Bedrock offers multiple options for where your inference requests are processed. *In-Region* keeps every request within a single Region, and *Geographic Cross-Region inference* routes requests across Regions within a geography for higher throughput, so your data stays within that geographic boundary. For NVIDIA Nemotron and GPT OSS models in AWS GovCloud (US), the options are as follows:
- In-Region inference is available in us-gov-west-1 (AWS GovCloud (US-West)).
- Geo cross-Region inference is available through a dedicated AWS GovCloud (US) cross-Region inference ID that routes requests across us-gov-west-1 and us-gov-east-1. Traffic stays within the AWS GovCloud (US) boundary, while you gain resilience across both Regions.
All inference for these models stays within the AWS GovCloud (US) boundary. Global cross-Region inference, which routes requests across commercial AWS Regions worldwide, isn’t available in AWS GovCloud (US). You can choose between single-Region and Geo cross-Region based on your requirements.
Service tiers
Amazon Bedrock offers multiple service tiers to match different workload requirements. For all three models, the Standard, Priority, and Flex tiers are supported.
| Service tier | Description | Supported |
|---|---|---|
| Standard | Pay-per-token access with no commitment | Yes |
| Priority | Higher throughput for latency-sensitive traffic | Yes |
| Flex | Lower-cost access for flexible, non-time-sensitive workloads | Yes |
| Reserved | Dedicated throughput with a term commitment | Not currently available |
By default, requests use on-demand inference on the Standard tier, where you pay per token without reserving capacity in advance. For latency-sensitive, customer-facing workloads, you can route individual requests to the Priority tier. For non-time-sensitive work such as model evaluations or batch summarization, the Flex tier offers a lower-cost option. For scaling guidance and how to handle throttling at production volume, refer to Scaling and throughput best practices and the Getting started section.
Getting started in AWS GovCloud (US)
This section walks through invoking the models, starting with the recommended bedrock-mantle endpoint. The examples use the us-gov-west-1 Region, where in-Region inference is available.
Console playground
- Navigate to the Amazon Bedrock console in your AWS GovCloud (US) account.
- Choose Playground from the left menu under the Test section.
- Choose Select model.
- Choose the provider (NVIDIA or OpenAI) from the category list, then select the model (for example, NVIDIA Nemotron 3 Super or 120B gpt-oss-120b).
- Choose Apply to load the model.
- Enter a prompt to test the model.
Using the bedrock-mantle endpoint (recommended)
To use these models, you need an AWS account in AWS GovCloud (US) with permissions to invoke Amazon Bedrock models. For the bedrock-mantle endpoint, you need an Amazon Bedrock API key or standard AWS credentials. The following is a sample policy:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "BedrockMantleInference",
"Effect": "Allow",
"Action": [
"bedrock-mantle:CreateInference",
"bedrock-mantle:Get*",
"bedrock-mantle:List*"
],
"Resource": "arn:aws-us-gov:bedrock-mantle:us-gov-west-1:111122223333:project/*"
},
{
"Sid": "BedrockMantleCallWithBearerToken",
"Effect": "Allow",
"Action": "bedrock-mantle:CallWithBearerToken",
"Resource": "*"
}
]
}Replace 111122223333 with your AWS account ID and scope the Region to the AWS GovCloud (US) Regions you use. The code examples in this post authenticate with a Bedrock API key, which requires bedrock-mantle:CallWithBearerToken. This action must be scoped to "Resource": "*", as shown in the second statement. To control which identities can generate or use Amazon Bedrock API keys, refer to Control permissions for generating and using Amazon Bedrock API keys. To restrict your organization to approved models only, use a service control policy (SCP).
The following example uses the OpenAI Python SDK to call the bedrock-mantle endpoint. For production workloads, use short-term API keys, which expire automatically (maximum 12 hours) and inherit the permissions of the IAM role that generated them.
import boto3
from openai import OpenAI
# Retrieve the Bedrock API key from AWS Secrets Manager
secrets_client = boto3.client("secretsmanager", region_name="us-gov-west-1")
api_key = secrets_client.get_secret_value(SecretId="bedrock-api-key")["SecretString"]
client = OpenAI(
# Use the AWS GovCloud (US) Region in the base URL, e.g. us-gov-west-1
base_url="https://bedrock-mantle.us-gov-west-1.api.aws/v1",
api_key=api_key,
)
response = client.chat.completions.create(
model="openai.gpt-oss-120b",
messages=[
{"role": "user", "content": "Explain the benefits of open-weight models for regulated workloads."}
],
reasoning_effort="medium", # low | medium | high
max_completion_tokens=512,
)
print(response.choices[0].message.content)Note: These examples retrieve the Bedrock API key from AWS Secrets Manager. For local development, you can instead read the key from an environment variable, but avoid that pattern in production. Use AWS Secrets Manager or another secrets store.
To call NVIDIA Nemotron 3 Super 120B instead, change the model parameter to nvidia.nemotron-super-3-120b and remove the reasoning_effort parameter (reasoning effort control is specific to GPT OSS). No other code changes are required.
Controlling reasoning effort
GPT OSS models are reasoning models that expose an adjustable reasoning effort. Set the reasoning_effort parameter on the Chat Completions call to low, medium, or high to trade response latency against reasoning depth. Use low for high-volume, latency-sensitive traffic, and high for complex, multi-step reasoning or agentic planning. For reasoning models, prefer max_completion_tokens to bound the response length (the older max_tokens field is still accepted).
Using the Responses API
In addition to Chat Completions, GPT OSS models support the Responses API, OpenAI’s interface for reasoning-style interactions. It takes a single input rather than a messages array. NVIDIA Nemotron 3 Super 120B doesn’t support the Responses API. Use Chat Completions, Converse, or Invoke for that model.
import boto3
from openai import OpenAI
# Retrieve the Bedrock API key from AWS Secrets Manager
secrets_client = boto3.client("secretsmanager", region_name="us-gov-west-1")
api_key = secrets_client.get_secret_value(SecretId="bedrock-api-key")["SecretString"]
client = OpenAI(
base_url="https://bedrock-mantle.us-gov-west-1.api.aws/v1",
api_key=api_key,
)
response = client.responses.create(
model="openai.gpt-oss-120b",
input="Explain the benefits of open-weight models for regulated workloads.",
)
print(response)Streaming responses
For chat and agent use cases where you want to surface tokens to the user as they are generated, set stream=True. The response becomes an iterator of incremental delta events:
stream = client.chat.completions.create(
model="openai.gpt-oss-120b",
messages=[
{"role": "user", "content": "Write a short summary of mixture-of-experts architectures."}
],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
print()On the bedrock-runtime endpoint, the equivalent capability requires the bedrock:InvokeModelWithResponseStream permission, which the minimum policy shown later already grants.
Tool calling
NVIDIA Nemotron and GPT OSS open-weight models are designed for agentic workflows, making them actionable for tool-calling scenarios. In a tool-calling workflow, you define functions (tools) that the model can invoke, the model decides when to call them based on the user’s request, and your application runs the function and returns the result for the model to incorporate into its final response.
The following example demonstrates this pattern end to end. We define a get_weather tool, send a user message, let the model request the tool call, run the function with mock data, and pass the result back so the model can generate a natural-language answer.
import json
import boto3
from openai import OpenAI
Retrieve the Bedrock API key from AWS Secrets Manager
secrets_client = boto3.client("secretsmanager", region_name="us-gov-west-1")
api_key = secrets_client.get_secret_value(SecretId="bedrock-api-key")["SecretString"]
client = OpenAI(
base_url="https://bedrock-mantle.us-gov-west-1.api.aws/v1",
api
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み