Amazon Bedrock に OpenAI GPT-5.6 系列が追加
OpenAI は、Amazon Bedrock 上で GPT-5.6 シリーズの Sol、Terra、Luna の 3 つの新モデルを一般提供し、開発者が既存の OpenAI API を利用しながら AWS のセキュリティとコスト管理機能を活用できるようになった。
AIニュース価値スコアβ
主要ニュースAI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 検索具体性
- 100
- 重複の少なさ
- 100
- 日本での有用性
- 25
記事は「Sol」「Terra」「Luna」という具体的な新モデル名、バージョン番号(GPT-5.6)、および利用可能な API エンドポイントや仕様を詳細に報じており、AI モデルの最新動向として極めて関連性が高い。ただし、日本企業への直接的な影響や日本語固有の情報はないため、日本の文脈での価値は限定的である。
キーポイント
GPT-5.6 シリーズの登場と役割分担
Sol(推論・コーディング)、Terra(汎用バランス)、Luna(高速・低コスト)という 3 つのモデルが、それぞれ異なるワークロードに最適化されて提供される。
AWS Bedrock への統合と利点
OpenAI の Frontier モデルを AWS のインフラ上で実行可能にし、地域ごとのデータ処理、セキュリティ、コスト管理の制御権をユーザーに提供する。
API と価格体系の統一
OpenAI Responses API を経由してアクセスでき、価格は OpenAI 公式レートと同等で、AWS の既存コミットメント(容量割引)も適用される。
認証と設定の柔軟性
環境変数から短期キーを取得する簡易な方法と、本番環境向けに自動更新機能を持つプロバイダーやAWS Secrets Managerを利用する方法が提供されています。
推論強度の制御
複雑なタスクに対して GPT-5.6 モデルが使用する推論トークンの量(none から max まで)を明示的に設定でき、コストとレイテンシを調整できます。
ツール呼び出し機能
モデルが定義されたツール(例:天気取得)を要求し、アプリケーション側で実行して結果を返すことで、より高度なタスク処理が可能になります。
ツール呼び出しの多段階プロセス
モデルが関数呼び出しを生成した後、その結果をコンテキストに追加して再度リクエストすることで、最終的な回答を生成するフローを示しています。
重要な引用
Developers building agentic coding, long-horizon reasoning, and high-volume inference workloads want frontier models they can call through familiar APIs, without operating separate model infrastructure.
Sol is the flagship reasoning model, Terra balances performance and cost for everyday production work, and Luna is optimized for fast, low-cost inference
Pricing matches OpenAI first-party rates, and usage counts toward your existing AWS commitments.
GPT-5.6 models can spend additional reasoning tokens on complex, multi-step tasks before answering, which improves results but increases latency and cost.
To move an existing OpenAI SDK application to GPT-5.6 on Amazon Bedrock, update the base URL and the model ID.
# Step 2: Carry the model's output (including any reasoning items) into the next turn.
影響分析・編集コメントを表示
影響分析
この発表は、OpenAI と AWS の戦略的提携をさらに深化させ、企業が OpenAI の最先端モデルを自社のクラウドインフラのセキュリティ・コンプライアンス要件に完全に適合させて利用することを可能にする画期的な一歩です。開発者は複雑な独自インフラ構築の手間を省きつつ、大規模推論やエージェント処理といった高度なユースケースにも即座に対応できるようになります。
編集コメント
GPT-5.6 という最新バージョンのモデルファミリーが、AWS の堅牢なインフラ上で提供されることで、エンタープライズレベルでの AI 導入におけるセキュリティとコスト管理の課題が大幅に緩和されます。特に Sol モデルによる自律型コーディングエージェントや長文推論への対応は、開発生産性の向上において重要な転換点となるでしょう。
本記事は、OpenAI の Chris Dickens と共著です。
エージェントによるコーディング、長期的な推論、大規模な推論ワークロードに取り組む開発者は、別個のモデルインフラを運用することなく、馴染み深い API を通じて最先端モデルを利用できることを望んでいます。OpenAI の「GPT-5.6 Sol」「Terra」「Luna」が、Amazon Bedrock で一般利用可能になりました。この 3 つのモデルは、自律型コーディングエージェントや長期的な推論から、高ボリュームかつ遅延に敏感な推論までをカバーしており、AWS のセキュリティ、リージョンごとのデータ処理、コスト管理機能を備えています。
これら 3 つのモデルは、すべて Bedrock の「mantle」エンドポイントにある OpenAI Responses API を通じてアクセスできます。Sol は中核となる推論モデルであり、Terra は日常の生産環境向けに性能とコストのバランスを最適化し、Luna は高速かつ低コストな推論に特化しています。これにより、各ワークロードに合わせて機能とコストを適切に調整することが可能です。料金は OpenAI 公式のレートに準拠しており、利用量は既存の AWS コミットメントにもカウントされます。
本記事では、モデルの選択方法から、Responses API を通じた最初の推論実行、プロンプトキャッシングによるコスト削減とキャッシュされたトークンの計測方法、OpenAI Codex エージェントとの連携、そしてクォータとスケーリングの計画までを解説します。なお、送信したプロンプトや生成結果は、モデルの学習には使用されず、モデルプロバイダーとも共有されません。
Amazon Bedrock で利用可能な GPT-5.6 ファミリー
OpenAI が導入した新しい命名規則により、GPT-5.6 では「数字」が世代を、「Sol」「Terra」「Luna」という名称がそれぞれ独立して進化できる能力のレベルを示します。Amazon Bedrock 上で利用できる各モデルの主要仕様は以下の通りです。
| モデル | モデル ID | 主な用途 | AWS リージョン |
|---|---|---|---|
| Sol | openai.gpt-5.6-sol | オートメーションによるコーディング、セキュリティ調査、科学分析、多段階の深い推論 | US East (N. Virginia), US East (Ohio) |
| Terra | openai.gpt-5.6-terra | 推論力、パフォーマンス、コストのバランスが取れた汎用生産ワークロード | US East (N. Virginia), US East (Ohio), US West (Oregon) |
| Luna | openai.gpt-5.6-luna | クラス分類、要約、ルーティングなど、大量処理かつ低遅延が求められるワークロード | US East (N. Virginia), US East (Ohio), US West (Oregon) |
これら 3 つのモデルはすべて、テキストと画像の入力、テキスト出力に対応しており、コンテキストウィンドウは 272K トークンまで拡張可能です。また、Responses API もサポートしています。
さらに、推論の強度を「なし」「低」「中」「高」「超高」「最大」の 6 レベルから選択できるため、API の実装を変更することなく、用途に応じてモデルを切り替えることが可能です。
bedrock-mantle エンドポイント経由での GPT-5.6 アクセス
GPT-5.6 モデルは、OpenAI Responses API を bedrock-mantle エンドポイントを通じて利用できます。ベース URL は https://bedrock-mantle.{region}.api.aws であり、Responses API は /openai/v1/responses で提供されます。{region} には us-east-1 などの対応する AWS リージョン名を指定してください。この openai/v1 パスは OpenAI モデルに固有のものです。また、本エンドポイントは OpenAI の Python および TypeScript SDK と互換性があります。
既存の OpenAI SDK アプリケーションを Amazon Bedrock で実行するには、OpenAI のベース URL を bedrock-mantle エンドポイントに置き換え、対応する Amazon Bedrock のモデル ID を使用し、Amazon Bedrock API キー または AWS 認証情報で認証を行ってください。
セキュリティとデータ取り扱い
すべてのモデル呼び出しは、お客様の AWS Identity and Access Management (IAM) ポリシーの下で実行され、仮想プライベートクラウド(VPC)内で完結し、AWS CloudTrail にログとして記録されます。リージョン内推論により、リクエストは指定した AWS リージョン内に留まるため、データ所在地に関する要件を満たすチームにとって有効です。
GPT-5.6 Sol、Terra、Luna は OpenAI が提供するサードパーティ製モデルで、Amazon Bedrock で利用可能です。これらのモデルの利用には OpenAI の利用規約 が適用されます。
OpenAI モデルを利用する場合、分類システムによって検出された不正トラフィックは、自動的なオフラインの 悪用検知 のために最大 30 日間保持されます。保存される入力と出力は AWS が保管・処理しますが、ユーザーがオプトインしない限り、モデルプロバイダーとは共有されません。データの保持設定は データ保持 モードを通じて制御できます。
Amazon Bedrock で GPT-5.6 を使い始める
Amazon Bedrock で GPT-5.6 を利用するには、以下の手順を実行してください。
事前準備
GPT-5.6 モデルを利用するには、bedrock-mantle エンドポイントで推論実行権限を持つ AWS アカウントが必要です。この権限付与の一例として、IAM プリンシパルに AWS マネージドポリシー「AmazonBedrockMantleInferenceAccess」をアタッチする方法があります。これにより、本記事の例に必要な読み取りおよび推論作成アクセス権(bedrock-mantle:CreateInference や bedrock-mantle:CallWithBearerToken など)が付与されます。
OpenAI Python SDK をバージョン 2.45.0 以降でインストールしてください:
pip install "openai>=2.45.0"認証には API キーを使用します。OpenAI SDK の認証には以下の 2 つの方法があります。
AWS 認証情報から自動更新される短期キー
OpenAI SDK に標準搭載されている BedrockOpenAI クライアントは、AWS 認証情報から短期トークンを生成し、リクエスト実行前に自動的に更新するプロバイダーを受け取ります。本記事のサンプルコードではこのクライアントを使用しています。
from aws_bedrock_token_generator import provide_token
from openai import BedrockOpenAI
region = "us-east-1"
client = BedrockOpenAI(
aws_region=region,
bedrock_token_provider=lambda: provide_token(region=region),
)環境変数から取得する短期キー
AWS_BEARER_TOKEN_BEDROCK 環境変数にトークンを設定し、クライアントに渡す方法もあります。ただしこの方式では自動更新されないため、最長 12 時間で期限切れになります。本番環境では、自動更新機能を利用するか、AWS Secrets Manager にトークンを保存することをお勧めします。
import os
from openai import OpenAI
client = OpenAI(
base_url="https://bedrock-mantle.us-east-1.api.aws/openai/v1",
api_key=os.environ["AWS_BEARER_TOKEN_BEDROCK"],
)Responses API を使った最初の推論実行
前項で設定した BedrockOpenAI クライアントを使って、GPT-5.6 Terra モデルを Responses API で呼び出してみましょう。Responses API は入力を 1 つのフィールドにまとめる形式で、生成されたテキストは output_text フィールドとして返されます。
response = client.responses.create(
model="openai.gpt-5.6-terra",
input="Explain the benefits of prompt caching for agentic workloads.",
max_output_tokens=512,
store=False,
)
print(response.output_text)Amazon Bedrock で GPT-5.6 を利用するには、既存の OpenAI SDK アプリケーションからベース URL とモデル ID を更新するだけで済みます。
推論コストの制御
GPT-5.6 モデルは、複雑で多段階を要するタスクに対して回答前に追加の推論トークンを消費できます。これにより精度が向上しますが、レイテンシとコストが増加します。この挙動は reasoning パラメータでレベルを設定することで制御可能です。Sol、Terra、Luna は「none(なし)」から「low」「medium」「high」「xhigh」「max」まで 6 レベルをサポートしています。タスクの難易度に合わせて適切なレベルを選択してください。
response = client.responses.create(
model="openai.gpt-5.6-sol",
input="A train leaves at 3 PM at 60 km/h. Another leaves an hour later at "
"90 km/h from the same station. When does the second catch up?",
reasoning={"effort": "high"},
)
print(response.output_text) ツールの呼び出し
GPT-5.6 はツール呼び出し機能をサポートしており、定義したツールの実行をモデル側からリクエストし、その結果を活用してタスクを完了させることが可能です。以下の例はクライアントサイドでのツール呼び出しを示しています。ここではアプリケーションがツールを実行し、その結果をモデルに返す仕組みです。
まず get_weather というツールを定義します。その後、1 ラウンドのやり取り(リクエストとレスポンス)が完了します。具体的には、モデルがツールの使用を要求し、アプリケーション側で実行して結果を返し、最後にモデルが最終回答を生成するという流れです。
import json以下のコードは、Amazon Bedrock で OpenAI の GPT-5.6 Sol、Terra、Luna を利用してツール呼び出しを実行するサンプルです。
まずはツール定義を準備します。ここでは「get_weather」という関数を定義し、指定された場所の現在の天気を取得するためのパラメータとして「location(都市と国)」と「unit(摂氏または華氏)」を設定しています。
次に、ユーザーからのリクエストを送信するステップです。この例では「Seattle の天気はどうですか?」という質問をモデルに投げかけます。
その後、モデルが返した応答(推論プロセスを含む場合もあります)を取得し、次の対話ラウンドへ引き継ぎます。これにより、ツール呼び出しの結果や追加の推論情報を継続的に処理することが可能になります。
この一連の流れを踏まえることで、GPT-5.6 シリーズの高度な機能を活用した柔軟なアプリケーション開発が可能となります。
ステップ3:各要求された関数を実行し、その結果を追加する
for item in response.output:
if item.type == "function_call":
args = json.loads(item.arguments)
result = {
"location": args["location"],
"temperature": 64,
"condition": "Partly cloudy",
}
input_list.append(
{
"type": "function_call_output",
"call_id": item.call_id,
"output": json.dumps(result),
}
)
ステップ4:ツールの結果を反映した最終レスポンスをモデルに要求する
final_response = client.responses.create(
model="openai.gpt-5.6-terra",
input=input_list,
tools=tools,
)
print(final_response.output_text)
GPT-5.6 モデルは回答前に推論を行うため、モデルの出力アイテム(推論結果を含む場合がある)を次のリクエストに渡す必要があります。前述の例では、response.output を入力リストに追加することでこれを実現しています。
本番環境でのデプロイには、Amazon Bedrock Guardrails を使用し、ユースケースや責任ある AI のポリシーに合わせてカスタマイズしたセーフガードを実装してください。
コンソールで GPT-5.6 を試す
GPT-5.6 モデルは次世代の推論エンジン上で動作しており、新しい Amazon Bedrock コンソール体験 が提供されています。これは bedrock-mantle エンドポイントおよび OpenAI 互換・Anthropic 互換 API に最適化されたものです。
この体験はプロジェクトベースの形式で行われます。まずプロジェクトを作成し、モデルを割り当てて API キーを設定します。その後、アプリケーションコードを書く前に複数のモデルを並列で評価できます。以下の手順に従えば、コンソールから離れることなく GPT-5.6 を試すことができます。
- モデルが利用可能なリージョン(例:US East (N. Virginia))で新しい Amazon Bedrock コンソールを開きます。既存のコンソールを利用している場合は、「新しい Bedrock コンソールを試す」を選択してください。
- プロジェクトを作成するか、既存のものを開きます。
- モデルカタログで利用可能な GPT、Claude、オープンウェイトモデルを確認します。最大 3 つのモデルを並列比較でき、機能やマルチモーダル対応、コンテキストウィンドウの大きさ、価格帯、リージョンごとの利用可否などをチェックできます。
- プロジェクトに追加する GPT-5.6 モデルを選択します。
- 評価を開始し、プロンプトを入力してモデルの回答を確認します。同じプロンプトに対する最大 3 つのモデルの回答を比較することも可能です。
以下のスクリーンショットは、新しい Amazon Bedrock コンソール内のモデルカタログを示しています。ここでは GPT-5.6 を含むさまざまなモデルを検索・比較できます。

GPT-5.6 のプロンプトキャッシングでコストを削減
エージェント型や多段階のワークロードでは、呼び出し間でコンテキストの多くが重複します。システム指示、ツールの定義、参照ファイルなどは変わらず、最新の入力のみが変化することが一般的です。GPT-5.6 は Amazon Bedrock 上でプロンプトキャッシングを 2 つのモードでサポートしています。
デフォルトで有効になっている「インプリシット(暗黙的)キャッシング」では、条件を満たすリクエストはコード変更なしで自動的にキャッシュされます。一方、「エクスペリット(明示的)キャッシング」を使えば、キャッシュの区切り点を指定して、プロンプトのどの部分をキャッシュするかを精密に制御できます。
どちらのモードでも、リクエスト量が増えるにつれて共有コンテキストを繰り返し処理するコストを削減できます。
キャッシュ区切り点(cache breakpoint)を設定すると、再利用可能なプロンプトプレフィックスの終わりを指定できます。そのプレフィックスを共有する後続のリクエストでは、Amazon Bedrock が既に処理済みのコンテキストを再利用するため、各呼び出しで支払うのは新しい作業分のみになります。
キャッシュされた入力は、未キャッシュの入力トークンと比較して 90% の割引料金で課金されます。また、キャッシュに書き込まれたトークンは、未キャッシュ入力料率の 1.25 倍で課金されます。現在の料率については Amazon Bedrock の価格ページ をご確認ください。
キャッシュされたコンテンツは、少なくとも 30 分間は再利用可能となり、単一のエージェント実行で発生する呼び出しのバーストに対応できる長さです。各区切り点には最低 1,024 トークンのプレフィックスが必要です。また、1 つのリクエストあたり最大 4 つのキャッシュチェックポイントを設定できます。
もしプレフィックスが最小値より短い場合、リクエスト自体は成功しますが、何もキャッシュされず、cached_tokens はゼロになります。
キャッシュブレークポイントによる明示的キャッシング
再利用可能なセクションの末尾に prompt_cache_breakpoint を追加し、prompt_cache_options を明示モード(explicit mode)に設定することで、プレフィックスをキャッシュできます。リクエスト間で一貫した prompt_cache_key を使用すれば、同じキャッシュにルーティングされ、マッチングの信頼性が向上します。
以下の例では、システム指示がキャッシュされます。ユーザーからの質問はブレークポイントの後に来るため、内容が変わってもキャッシュされたプレフィックスが無効化されることはありません。この例では、キャッシュへの書き込みと読み込みを示すために、同じリクエストを 2 回送信しています。
# 実際のシステム指示や参照コンテンツに置き換えてください。
# キャッシュされるプレフィックスは少なくとも 1,024 トークン必要です。それ未満だとキャッシュされません。
system_prompt = "You are a technical support agent for Example Corp. ...(1,024+ tokens)..."def ask(question):
return client.responses.create(
model="openai.gpt-5.6-terra",
prompt_cache_key="support-agent:system-prompt-v1",
prompt_cache_options={"mode": "explicit"},
input=[
{
"type": "message",
"role": "developer",
"content": [
{
"type": "input_text",
"text": system_prompt,
# ここまでの内容をすべてキャッシュに保存します(システム指示部分)。
"prompt_cache_breakpoint": {"mode": "explicit"},
}
],
},
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": question}],
},
],
)
初回呼び出し:プレフィックスをキャッシュに書き込みます。
first = ask("シングルサインオンを設定するにはどうすればよいですか?")
print("write:", first.usage.input_tokens_details.cache_write_tokens)
同じプレフィックスとキャッシュキーで2回目の呼び出し:キャッシュからプレフィックスを読み取ります。
second = ask("パスワードをリセットする方法は?")
print("read: ", second.usage.input_tokens_details.cached_tokens)
print(second.output_text)
ここで示されている system_prompt は簡略化されたプレースホルダーです。実際のコンテンツに置き換え、少なくとも 1,024 トークン分の長さを確保してください。十分な長さのプレフィックス(先頭部分)を設ければ、最初の呼び出しで cache_write_tokens がゼロ以外の値となり、2 回目の呼び出しでは cached_tokens が検出されます。これは、先頭部分が再利用されたことを確認するものです。
明示モードは、先頭部分が大きく安定しており、キャッシュされる内容を完全に制御したいエージェントループに適しています。
暗黙的キャッシング(Implicit Caching)
prompt_cache_options を設定しない場合、GPT-5.6 はデフォルトの「暗黙的キャッシング」モードを使用します。Amazon Bedrock は最新のメッセージに自動的にキャッシュのブレイクポイントを設け、ユーザーが追加した明示的なブレイクポイントも尊重します。これにより、入力構造を変更することなく、安定したプロンプトの先頭部分をリクエスト間で再利用できます。これがキャッシングの恩恵を最も早く得られる方法です。
静的なコンテンツ(システム指示、ツール定義、参照ドキュメントなど)はプロンプトの先頭に配置し、変化する内容は末尾に配置します。関連するリクエストに対して一貫した prompt_cache_key を設定すれば、エンドポイントは一致する処理済み先頭部分を自動的に再利用します。
以下の例では暗黙的キャッシングを使用しています。prompt_cache_options もブレイクポイントも設定せず、前回の例で使った system_prompt(実際のコンテンツに置き換え、少なくとも 1,024 トークン分の長さを確保してください)を、一貫した prompt_cache_key とともに再利用します。
response = client.responses.create(
model="openai.gpt-5.6-terra",
prompt_cache_key="support-agent:kb-v1",
input=[
{
"type": "message",
"role": "developer",
# Static content first so it forms a stable, cacheable prefix.
"content": [{"type": "input_text", "text": system_prompt}],
},
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "How do I configure single sign-on?"}],
},
],
)
print(response.output_text)
The trade-off is control. In implicit mode, you don't decide exactly where the cacheable boundary falls. Implicit caching suits chat and Retrieval Augmented Generation (RAG) workloads with a naturally stable prefix,
原文を表示
*This post is co-written with Chris Dickens from OpenAI.*
Developers building agentic coding, long-horizon reasoning, and high-volume inference workloads want frontier models they can call through familiar APIs, without operating separate model infrastructure. OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock. The three models cover workloads from autonomous coding agents and long-horizon reasoning to high-volume, latency-sensitive inference, and they run with the security, Regional processing, and cost controls of AWS.
You access all three models through the OpenAI Responses API on the bedrock-mantle endpoint. Sol is the flagship reasoning model, Terra balances performance and cost for everyday production work, and Luna is optimized for fast, low-cost inference, so you can right-size capability and cost for each workload. Pricing matches OpenAI first-party rates, and usage counts toward your existing AWS commitments.
In this post, we show you how to select a model, run your first inference through the Responses API, reduce cost with prompt caching and measure cached-token usage, connect the OpenAI Codex coding agent, and plan for quotas and scaling. Your prompts and completions are not used to train any models, and are not shared with the model provider.
The GPT-5.6 family on Amazon Bedrock
GPT-5.6 introduces a naming system from OpenAI where the number identifies the generation and the names Sol, Terra, and Luna identify durable capability tiers that can advance on their own cadence. The following table summarizes the key specifications for each model on Amazon Bedrock.
Model
Model ID
Best suited for
AWS Regions
Sol
openai.gpt-5.6-sol
Autonomous coding, security research, scientific analysis, and deep multi-step reasoning
US East (N. Virginia), US East (Ohio)
Terra
openai.gpt-5.6-terra
General-purpose production workloads that balance reasoning, performance, and cost
US East (N. Virginia), US East (Ohio), US West (Oregon)
Luna
openai.gpt-5.6-luna
High-volume, latency-sensitive workloads such as classification, summarization, and routing
US East (N. Virginia), US East (Ohio), US West (Oregon)
All three models support text and image input, text output, a 272K-token context window, and the Responses API. They also support none, low, medium, high, xhigh, and max reasoning effort, so you can switch models without changing your API integration.
Accessing GPT-5.6 through the bedrock-mantle endpoint
You access GPT-5.6 models through the OpenAI Responses API on the bedrock-mantle endpoint. The base URL is https://bedrock-mantle.{region}.api.aws, and the Responses API is served at /openai/v1/responses. Replace {region} with a supported AWS Region, such as us-east-1. This openai/v1 path is specific to the OpenAI models. The endpoint works with the OpenAI Python and TypeScript SDKs. To run an existing OpenAI SDK application on Amazon Bedrock, replace the OpenAI base URL with the bedrock-mantle endpoint, use the corresponding Amazon Bedrock model ID, and authenticate with an Amazon Bedrock API key or AWS credentials.
Security and data handling
Every model call runs under your AWS Identity and Access Management (IAM) policies, inside your virtual private cloud (VPC), and is logged in AWS CloudTrail. In-Region inference keeps requests within the AWS Region you specify, which helps teams meet data-residency requirements.
GPT-5.6 Sol, Terra, and Luna are third-party models from OpenAI, made available on Amazon Bedrock and subject to the OpenAI terms. For these OpenAI models, classifier-flagged traffic is retained for up to 30 days for automated offline abuse detection. Retained inputs and outputs are stored and processed by AWS and are not shared with the model provider unless you opt in. You control retention configuration through data retention mode.
Get started with GPT-5.6 on Amazon Bedrock
Complete the following steps to start using GPT-5.6 on Amazon Bedrock.
Prerequisites
To use GPT-5.6 models, you need an AWS account with permissions to run inference on the bedrock-mantle endpoint. One way to grant these is to attach the AWS managed policy AmazonBedrockMantleInferenceAccess to your IAM principal. It grants the read and inference-creation access the examples in this post need, including bedrock-mantle:CreateInference and bedrock-mantle:CallWithBearerToken.
Install the OpenAI Python SDK, version 2.45.0 or later:
pip install "openai>=2.45.0"You authenticate with an API key. There are two options for authenticating the OpenAI SDK.
- Auto-refreshing short-term key: The OpenAI SDK’s native BedrockOpenAI client takes a token provider that generates a short-term key from your AWS credentials and refreshes it before each request. The examples in this post use this client.
from aws_bedrock_token_generator import provide_token
from openai import BedrockOpenAI
region = "us-east-1"
client = BedrockOpenAI(
aws_region=region,
bedrock_token_provider=lambda: provide_token(region=region),
)- Short-term key from an environment variable: Set the key on AWS_BEARER_TOKEN_BEDROCK and pass it to the client. Because this key isn’t refreshed, it expires after at most 12 hours. For production, use the auto-refreshing option or store the key in AWS Secrets Manager.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://bedrock-mantle.us-east-1.api.aws/openai/v1",
api_key=os.environ["AWS_BEARER_TOKEN_BEDROCK"],
)Run your first inference with the Responses API
Using the BedrockOpenAI client from the previous step, call GPT-5.6 Terra through the Responses API. The Responses API uses a single input field and returns the generated text in output_text.
response = client.responses.create(
model="openai.gpt-5.6-terra",
input="Explain the benefits of prompt caching for agentic workloads.",
max_output_tokens=512,
store=False,
)
print(response.output_text)To move an existing OpenAI SDK application to GPT-5.6 on Amazon Bedrock, update the base URL and the model ID.
Control reasoning effort
GPT-5.6 models can spend additional reasoning tokens on complex, multi-step tasks before answering, which improves results but increases latency and cost. Set the level with the reasoning parameter. Sol, Terra, and Luna support none, low, medium, high, xhigh, and max. Match the level to the task.
response = client.responses.create(
model="openai.gpt-5.6-sol",
input="A train leaves at 3 PM at 60 km/h. Another leaves an hour later at "
"90 km/h from the same station. When does the second catch up?",
reasoning={"effort": "high"},
)
print(response.output_text)Call tools
GPT-5.6 supports tool calling, which lets the model request tools you define and use their results to complete a request. The following example demonstrates client-side tool calling, where your application runs the tool and returns the result to the model. It defines a get_weather tool and completes one round-trip. The model requests the tool, your application runs it and returns the result, and the model produces the final answer.
import json
tools = [
{
"type": "function",
"name": "get_weather",
"description": "Get the current weather for a given location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City and country (for example, Seattle, US)",
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit",
},
},
"required": ["location"],
},
}
]
# Step 1: Send the user request with the tool definition.
input_list = [{"role": "user", "content": "What's the weather like in Seattle?"}]
response = client.responses.create(
model="openai.gpt-5.6-terra",
input=input_list,
tools=tools,
)
# Step 2: Carry the model's output (including any reasoning items) into the next turn.
input_list += response.output
# Step 3: Run each requested function and append its result.
for item in response.output:
if item.type == "function_call":
args = json.loads(item.arguments)
result = {
"location": args["location"],
"temperature": 64,
"condition": "Partly cloudy",
}
input_list.append(
{
"type": "function_call_output",
"call_id": item.call_id,
"output": json.dumps(result),
}
)
# Step 4: Ask the model for the final response, incorporating the tool result.
final_response = client.responses.create(
model="openai.gpt-5.6-terra",
input=input_list,
tools=tools,
)
print(final_response.output_text)Because GPT-5.6 models reason before responding, pass the model’s output items (which can include reasoning) back in the next request, as the preceding example does when it appends response.output to the input list.
For production deployments, use Amazon Bedrock Guardrails to implement safeguards customized to your use cases and responsible AI policies.
Try GPT-5.6 on the console
GPT-5.6 models run on the next-generation inference engine, which has a new Amazon Bedrock console experience optimized for the bedrock-mantle endpoint and its OpenAI-compatible and Anthropic-compatible APIs.
This experience is project-based. You create a project, assign models, configure API keys, and evaluate models side by side before writing application code. Complete the following steps to try GPT-5.6 without leaving the console:
- Open the new Amazon Bedrock console in a Region where the models are available, such as US East (N. Virginia). If you are in the existing console, choose Try the new Bedrock console.
- Create a project, or open an existing one.
- In the model catalog, review the available GPT, Claude, and open-weight models. You can compare up to three models side by side on capabilities, modalities, context window, pricing, and Regional availability.
- Choose a GPT-5.6 model to add it to your project.
- Start an evaluation, enter a prompt, and review the model’s response. You can select up to three models to compare responses to the same prompt.
The following screenshot shows the model catalog in the new Amazon Bedrock console, where you can browse and compare GPT-5.6 and other models.

Reduce cost with GPT-5.6 prompt caching
Agentic and multi-step workloads repeat much of their context between calls. System instructions, tool definitions, and reference files often stay the same while only the latest input changes. GPT-5.6 supports prompt caching in two modes on Amazon Bedrock. Implicit caching is on by default, so eligible requests are cached automatically without code changes. Explicit caching lets you mark cache breakpoints for precise control over which parts of a prompt are cached. In both modes, prompt caching reduces the cost of repeatedly processing shared context as request volume grows.
With a cache breakpoint, you mark the end of a reusable prompt prefix. On a subsequent request that shares that prefix, Amazon Bedrock reuses the processed context, and each call pays full price only for the new work. Cached input is billed at a 90% discount compared to uncached input tokens, and tokens written to cache are billed at 1.25 times the uncached input rate. For current rates, see the Amazon Bedrock pricing page. Cached content stays available for reuse for at least 30 minutes, long enough to cover the burst of calls a single agent run generates. Each breakpoint requires a prefix of at least 1,024 tokens, and you can set up to four cache checkpoints per request. If a prefix is shorter than the minimum, the request still succeeds, but nothing is cached and cached_tokens stays zero.
Explicit caching with cache breakpoints
To cache a prefix, add a prompt_cache_breakpoint to the content block that ends the reusable section, and set prompt_cache_options to explicit mode. Setting a consistent prompt_cache_key across requests routes them to the same cache and improves match reliability. In the following example, the system instruction is cached, and the user’s question comes after the breakpoint so it can change without invalidating the cached prefix. The same request is sent twice to show a cache write followed by a cache read.
# Replace with your real system instructions and reference content.
# The cached prefix must be at least 1,024 tokens, or nothing is cached.
system_prompt = "You are a technical support agent for Example Corp. ...(1,024+ tokens)..."
def ask(question):
return client.responses.create(
model="openai.gpt-5.6-terra",
prompt_cache_key="support-agent:system-prompt-v1",
prompt_cache_options={"mode": "explicit"},
input=[
{
"type": "message",
"role": "developer",
"content": [
{
"type": "input_text",
"text": system_prompt,
# Cache everything up to this breakpoint (the system instruction).
"prompt_cache_breakpoint": {"mode": "explicit"},
}
],
},
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": question}],
},
],
)
# First call: writes the prefix to cache.
first = ask("How do I configure single sign-on?")
print("write:", first.usage.input_tokens_details.cache_write_tokens)
# Second call with the same prefix and cache key: read the prefix from cache.
second = ask("How do I reset a password?")
print("read: ", second.usage.input_tokens_details.cached_tokens)
print(second.output_text)The system_prompt here is an abbreviated placeholder. Replace it with real content of at least 1,024 tokens. With a large enough prefix, the first call reports a nonzero cache_write_tokens and the second a nonzero cached_tokens, confirming the prefix was reused. Explicit mode is a good fit for agentic loops with a large, stable prefix, where you want full control over what gets cached.
Implicit caching
If you don’t set prompt_cache_options, GPT-5.6 uses implicit caching, the default mode. Amazon Bedrock places an automatic cache breakpoint on the latest message and honors any explicit breakpoints you add, so a stable prompt prefix can be reused across requests without any change to your input structure. This is the quickest way to benefit from caching. Keep static content (system instructions, tool definitions, reference documents) at the front of the prompt and variable content at the end, set a consistent prompt_cache_key for related requests, and the endpoint reuses the processed prefix when it matches.
The following example uses implicit caching. It sets no prompt_cache_options and no breakpoint, and reuses the system_prompt from the previous example (replace it with real content of at least 1,024 tokens) with a consistent prompt_cache_key:
response = client.responses.create(
model="openai.gpt-5.6-terra",
prompt_cache_key="support-agent:kb-v1",
input=[
{
"type": "message",
"role": "developer",
# Static content first so it forms a stable, cacheable prefix.
"content": [{"type": "input_text", "text": system_prompt}],
},
{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "How do I configure single sign-on?"}],
},
],
)
print(response.output_text)The trade-off is control. In implicit mode, you don’t decide exactly where the cacheable boundary falls. Implicit caching suits chat and Retrieval Augmented Generation (RAG) workloads with a naturally stable prefix,
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み