Seedream 5.0のプロンプト方法
ByteDanceのSeedream 5.0は高品質な画像を生成し、詳細まで鮮明である。本記事では、同モデルへの効果的なプロンプト入力方法と得られる美的成果について解説している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
ByteDance の Seedream シリーズは目覚ましい進化を遂げています。私たちはこのモデルに数多くのプロンプトを試してきました。その結果をご紹介します。
美しさについて
本題に入る前に、実際に生成される画像の質感についてお話ししましょう。Seedream 5.0 は、拡大してもディテールが崩れないほど本物の美しい出力を生み出します。
**
側面を向く若い男性のポートレートで、背景をぼかす浅い被写界深度を持つカラーフィルム風の画像です。視線は彼の目に集中し、微細な粒状感と色調から高 ISO フィルムを使用していることが示唆されます。また、広角レンズによるボケ効果が生み出すモーションブラーが、カクテル写真のような自然でドキュメンタリースタイルを強化しています。

このモデルは写真撮影の言語を深いレベルで理解しています。特定のフィルム種、レンズ特性、ライティング設定を指定すると、まるでその機材から直接撮られたかのようなリアリティある画像を返してくれます。
夕暮れ時の東京の路地裏に立つ女性。濡れた舗道にネオンサインが反射する様子を捉えた一枚です。使用されたのは期限切れの Kodak Portra 800 で、露出を 2 ストップ分押し上げました。ラーメン屋からのタングステン照明が彼女の顔に温かいオレンジ色を落とし、一方、ネオン光は髪に冷たいシアン色のハイライトを浮かび上がらせます。粒状感は視認可能で、光源の周囲にはハレーションが見られ、黒レベルはわずかに持ち上げられています。彼女は歩行中の瞬間を捉えられており、2 つの色調の世界の狭間で止まっているかのようです。

ポートレートだけではありません。風景、静物、建築写真——どのジャンルも、単なる汎用的な処理ではなく、意図的な品味を感じさせるレベルでモデルが扱います。
氷河の川が火山性の黒い砂と出会うアイスランドの空中写真。抽象的な分岐パターンを作り出し、それは静脈や稲妻のように見えます。黄金時刻に 3000 フィート上空から撮影され、水路は黒曜石のような砂に対してターコイズブルーに輝いています。スケールを判断することは不可能です——毛細血管の顕微鏡画像か、デルタの衛星写真のどちらにも見えます。大型カメラによる極度のシャープネスで、地平線はありません。

粗い石灰岩の表面に置かれた半分のザクロの静物画。高い窓から差し込む午後の光の一本の光束によって照らされています。ルネサンス期のキアロスクーロ(明暗法)ライティング——種子は深い影に対してルビーのように輝いています。数粒の種子が石の上を転がり、小さな赤い跡を残しています。雰囲気はカラヴァッジョと現代のフードエディトリアルの間にあるようなものです。

薄明かりの頃、風雨にさらされた木製の桟橋で網を繕う老人。彼の両手が焦点であり、古びた傷跡と日焼けで黒ずんだ肌は、淡い青色のナイロンメッシュの中で熟練した精度をもって動いている。背後には海が、鉛色からローズゴールドへと柔らかなグラデーションを描いている。ライカ M に 50mm Summilux を開放で使用して撮影され、ボケは港の明かりを完璧な円形に変えている。この画像には、セバスティアーノ・サルガドのポートレートに通じる静かなる尊厳がある。

雨上がりの静かな水たまりに映る、ブルータリズムのコンクリート建築。そこにはルシュタック像のような対称性が生まれている。建物の幾何学的なファサード——繰り返される長方形の窓と無骨なコンクリート——が抽象的なパターンへと二重化している。赤い傘をさした一人の人物がその縁を歩いている。モノクロームのシーンにおいて、それが唯一の色である。曇り空、均一で拡散された光、端にはチルトシフトレンズ効果を用いた建築写真。

例に基づく編集
これは説明が最も難しく、かつ使用して最も楽しい機能です。
複雑な編集を言葉で記述しようとするのではなく、モデルに目的を示します。Before/After のペア——画像 1 と画像 2——と、さらに一つの画像を与えます。モデルは最初の二つの画像間で何が変わったかを推測し、その変換を第三の画像に適用します。
例を示しましょう。まず、無地の白いセラミックマグカップから始め、その後にモデルに対して、日本語の金継ぎ(きんつぎ)による金粉で割れ目を修復した姿を画像として見せます(言葉での説明は行いません)。次に、全く異なるオブジェクトであるセラミックの花瓶を与え、同じ変換を適用するよう指示します。
画像 1 から画像 2 への変化を参照し、同じ操作を画像 3 に適用してください
画像 1

画像 2

画像 3

結果

モデルは、マグカップのペアから「金継ぎのパターンで金粉を埋め込んだ割れ目を追加する」という処理を学習し、それを花瓶にも同じように適用しました。私たちは金継ぎがどのような見た目をするかを言葉で説明する必要さえありませんでした。
この手法は、あらゆる種類の変換に有効です:
- 素材の置換:あるオブジェクトで「木製→大理石」を示し、別のオブジェクトに適用する
- シーンの変更:ある写真で「昼→夜」を示し、全く異なる場所にも適用する
- スタイル転送(スタイルトランスファー):写真を木版画に変換した例を示し、その芸術的な変換を新しいシーンにも適用する
スタイル転送
ここでの強みは、「あの正確な広重の色彩パレット」を説明する適切な言葉を探し出す必要がないことです。それを示すだけでよいのです。
このシーンを、広重風の伝統的な日本の浮世絵木版画として再想像してください——平面的な遠近法、太い輪郭線、藍・朱・赭の限られた色彩パレット。
入力

出力

色彩調整を、以下の条件に合わせて変換してください——彩度の高いティール色の影、温かみのあるアンバー色のハイライト、ソフトな拡散効果。
入力

出力

論理的推論
ほとんどの画像生成モデルは、プロンプトを単なるキーワードの集合体として扱います。しかし Seedream 5.0 は、あなたが何を求めているのかを実際に推論します。
ルーズ・ゴールドバーグ・マシン:大理石が木製の斜面を転がり落ち、一列に並んだドミノに衝突する。最後のドミノが紐を引き、その紐が水差しを傾け、注がれた水が天秤の小さなカップを満たす。カップが重さで沈み、レバーを引いて小さな真鍮のベルを鳴らす。すべての構成要素は物理的に正しい影を落とす。大理石は転がり途中である。マシンはグリッド線が見える製図台の上に置かれている。クロスハッチングスタイルで、特許図面が生き生きと再現されたような表現。

物理的な物体を機械的レベルで理解する能力も拡張されています:
黒いベルベットの上に分解され、展開図のように配置されたアンティークのポケットウォッチ。すべての歯車、バネ、脱進機車輪、バランスコック、宝石軸受が、組み立てられた機構内で本来あるべき位置に対して正しく配置されており、すべてが見えます。主バネは部分的に巻き戻されています。各部品には銅板書体(copperplate script)の小さなラベルで識別されています。博物館保存写真スタイル。

画像入力を用いた多段階推論
モデルに 2 つの画像と複雑な指示を与えると、多段階の操作を推論して実行できます。ここでは、混合された花束と 3 つの空の花瓶を与え、花の種類ごとに仕分けるよう依頼します:
Image 1 の花を品種別に分類し、Image 2 の 3 つの花瓶にそれぞれ別々に配置してください。バラは最初の花瓶へ、ヒマワリは二番目の花瓶へ、ラベンダーは三番目の花瓶へ。
Image 1

Image 2

結果

モデルは各花の種類を特定し、それらをグループ化した上で、単一のプロンプトから正しい容器に分配する必要があります。
試してみたい他の面白いことをご紹介します:
成長した姿がどうなるか生成してみましょう。
入力

出力

左から右へ並べられた 4 つの段階におけるオオカバマダラの完全な変態:縞模様の幼虫がミルクウィード(*Asclepias*)の葉を食べる様子、金色の斑点を持つエメラルドグリーンのさなぎ、翼が一部見える状態で殻が割れ始めたさなぎ、完全に羽化した広げた翼を持つ蝶。各段階にはエレガントなセリフ体のラベルで注釈が付けられています。17 世紀の博物学者マリア・シビッラ・メリアン(Maria Sibylla Merian)の自然画のスタイルでレンダリングされていますが、現代的な科学的正確さを備えています。

正確な指示の遵守
Seedream 5.0 の指示従順性は、以前のバージョンに比べて明らかに厳密になっています。「青いジャケット」と言えば、紫色でも水色でもない、まさに青いジャケットが生成されます。空間関係、数量、具体的な詳細を指定した場合も、モデルはそれらを尊重します。
これは、多くの特定の要件を持つ複雑な構成において特に重要です:
シニアソフトウェアエンジニアの、写真のようにリアルな雑然としたオフィスデスク。開かれた MacBook Pro の画面には、黒背景に緑色の文字でコードがスクロールするターミナルが表示されています。セラミック製のマグカップにはモノスペースフォントで「console.log('coffee')」と書かれており、湯気がカーブを描いて立ち上っています。開かれた O'Reilly 社の書籍には、「Frontend」「Backend」「DevOps」とラベル付けされた3つの重なり合う円で構成されるベン図が描かれています。黄色(To Do)、オレンジ(In Progress)、緑(Done)の付箋がカンバンボードを形成しており、それぞれに小さな手書きのタスクが記されています。カスタムキーキャップを搭載したメカニカルキーボードと、Slack の通知を表示しているスマートフォンがあります。コンクリート製の頭蓋骨型の鉢植えに入った小さな多肉植物も置かれています。右側から黄金時刻の光が斜めに差し込み、長い影を落としています。

これは、マグカップの文字、書籍内のラベル付き図表、色分けされた付箋、頭蓋骨型の鉢など、10 以上の具体的な要件を含むプロンプトです。このモデルはそれらすべてを正確に追跡します。
このモデルは視覚的な手がかりも処理できます。画像に矢印、バウンディングボックス、または特定の領域を示す色付きの領域がある場合、それらをプロンプトで参照することができます:
スプレーペイントのマーカーに従ってロフトを装飾してください。壁にある赤い長方形の場所に、大きな抽象表現主義の絵画を置きます。床にある青い長方形の場所に、ミッドセンチュリーモダンの革製ソファを配置します。天井にある黄色い円の場所に、真鍮製のグローブペンダントライトを吊り下げます。スプレーペイントのマーカーは取り除いてください。産業的な雰囲気を保ってください。
入力 — スプレーペイントマーカーで描かれたロフト

出力

ドメイン知識
Seedream 5.0 は、複数の専門分野にわたって深く組み込まれた知識を有しています。これは単なる「建物がどのように見えるかを知る」を超えたものであり、技術的なコンテンツの構造と慣習を理解しています。
床面図スケッチを入力すると、空間レイアウトを尊重したリアルなビジュアライゼーションを生成します:
この床面図に基づき、インテリアのフォトリアリスティックなレンダリングを生成してください。温かみのある檜(ひのき)色のトーンを持つ日本風のミニマリストハウスで、天井まで届くガラス越しに見える中央の中庭ガーデンには赤いカエデの木が一本あり、畳敷き、日差しが木漏れ日となって差し込む縁側(えんがわ)のウッドデッキを備えています。引き戸の障子(しょうじ)は一部開いています。レイアウトは正確に一致させてください。
入力 — 建築家のスケッチ

出力

このモデルは、通常、専門知識と数時間にわたる慎重な作業を必要とする正確な科学イラストレーションを生成することができます:
サンゴ礁生態系の詳細な断面図、科学イラストレーションスタイル。下部:火山岩の玄武岩基盤。中部:化石層が確認できる炭酸カルシウム製のサンゴ礁構造。上部のサンゴ礁:ラベル付きの種々多様な生物が生息する繁栄した生態系 — 鹿角サンゴ、脳サンゴ、イソギンチャクに潜むカクレクマノミ、岩の隙間にいるウツボ、採食中のオウム魚、その上空を泳ぐウミガメ。水面より上部:ヤシの木が生い茂る小さな熱帯島。左端には水深を示す目盛り。水彩塗りつぶしを施した精密な線画、エレガントなセリフ体のラベル。ネイチャー誌のスタイルで出版。

食品の写真をモデルに提示し、栄養情報で注釈をつけるよう依頼してください:
このメゼ(前菜)盛り合わせの各料理を特定し、それぞれの隣にエレガントな手書きカリグラフィ風の注釈カードを追加します。各カードには料理名と 100g あたりのカロリー数を記載します。カードはクリーム色の紙地にブルゴーニュ色のインクで印刷。
入力画像

出力画像

テキストレンダリング
Seedream はバージョン 3.0 からテキストレンダリングに強みを持っており、5.0 でもその伝統を継承しています。画像内でレンダリングしたいテキストには二重引用符「」を囲んで使用すると、最も良い結果が得られます:
架空のジャズフェスティバルのための大型タイポグラフィポスター。上部には太字のコンデンスドサンセリフ体で「BLUE NOTE SESSIONS」を深いネイビー色で配置。その下にはエレガントなスクリプト体で「Summer 2026 — Central Park, New York.」と記す。ポスターには4組のパフォーマーがリストアップされている:「Miles Ahead Quintet / Saturday 8PM」、「Sarah Chen Trio / Saturday 10PM」、「The Monk Revival / Sunday 7PM」、「Coltrane Legacy Orchestra / Sunday 9PM」。右端には様式化された金色のサックスのシルエットが縦に走っている。背景は真夜中の青から温かみのあるアンバーへのグラデーション。下部には小文字大で「Tickets at bluenote.nyc — All ages welcome.」と記す。

これは複数の書体、混合した大文字・小文字、句読点、特定の演奏者名、そしてグラデーション背景を正確にレンダリングしたポスターです。このモデルは多言語テキストの処理も得意です。中国語、日本語、韓国語、またはその他のスクリプトを正確にレンダリングする必要がある場合でも、Seedream 5.0 は対応可能です。
マルチ画像生成
Seedream 4.5と同様に、5.0 モデルは一度に複数の関連する画像を生成できます。「シリーズ」や「セット」と指定するか、数を明示することで、一貫したスタイルとキャラクターの連続性を保った画像を生成します。
シネマティックな 2x2 のストーリーボードグリッド。パネル 1:廃墟となった宇宙ステーションの内部、孤独な宇宙飛行士が割れた船体パネルから蔓が生い茂る廊下を浮遊し、発光するキノコが青く輝いている。パネル 2:彼女は霜降りガラスの向こう側で脈打つ緑色の光を放つシールされた実験室の扉を発見し、手動解放レバーに手を伸ばす。パネル 3:扉が開き、繁栄した庭園が現れる。無重力の中で小さな木が育ち、根はクラゲのように外側に螺旋状に広がっており、蝶々は飛行中の瞬間で凍りついている。パネル 4:ヘルメットのバイザー越しの彼女の顔のクローズアップ。涙は小さな球体として浮遊し、その瞳には庭園が映し出されている。一貫したキャラクターデザイン、アノマフィックレンズフレア、リドリー・スコット色調。
「ALTITUDE」という名物のコーヒーロースター向けの包括的なブランドアイデンティティのフラットレイ。ダークスレートの上に配置されたマットブラックのコーヒーバッグ(エンボス加工されたゴールド箔で「ALTITUDE」のワードマークと標高コンターラインロゴが施されている)、名刺(表裏)、コンターロゴがエッチングされたセラミックのペーパードリップドレッパー、ブランド化されたクラフト紙のシール、スクリーンプリントされたロゴが入ったリネンのトートバッグ、レーザー彫刻でブランド名が施された金属製の旅行用タンブラー、「Brewing Guide」と題された小冊子。カラーパレットは黒、ゴールド、自然なクラフト色。上からの製品写真撮影、均一なスタジオ照明。
ストーリーボード、ブランドアイデンティティパッケージ、絵文字セット、そして一貫性のあるコレクションが必要なあらゆるシナリオにおいて、これは非常に優れた手法です。
API を使用した始め方
JavaScript と Replicate API を使用して Seedream 5.0 を実行する方法は以下の通りです:
そして Python では:
プロンプト作成のヒント
テストを通じて学んだいくつかのポイントをご紹介します:
- キーワードリストではなく、自然言語を使用してください。「Monet の油絵スタイルで、日傘をさした裕福なドレスを着た少女が、並木道を歩いている」というプロンプトは、「少女、傘、並木道、油絵の質感」といったキーワードリストよりも優れた結果をもたらします。
- テキスト描画には二重引用符を使用してください。画像に特定のテキストを含めたい場合は、それを二重引用符で囲みます:タイトルが「Seedream 5.0」であるポスターをデザインする。
- 維持したい要素について具体的に指示してください。編集を行う際は、変更してほしくない部分をモデルに明確に伝えます:「帽子を王冠に変えるが、ポーズと表情はそのままにする」。
- 複雑な編集には視覚的マーカーを使用してください。入力画像に矢印、枠線、または色付きの領域を描画し、どこを変更すべきかを正確に示します。
- 使用ケースを指定してください。「ゲーム会社のロゴをデザインする」とモデルに伝えることは、単に視覚要素を説明するよりも優れた結果をもたらします。
- 例に基づく編集を行う場合は、言葉で説明するのではなく見せることが重要です。言葉で変換方法を説明するのが難しい場合、変更前後のサンプルペアを提供してください。
原文を表示
ByteDance’s Seedream line has been on a tear. We spent a bunch of time throwing prompts at it. Here’s what we found.
Aesthetics
Before we get into the meat, let’s talk about how the images actually look. Seedream 5.0 produces genuinely beautiful output — the kind of images where you zoom in and the details hold up.
A color film-inspired portrait of a young man looking to the side with a shallow depth of field that blurs the surrounding elements, drawing attention to his eye. The fine grain and cast suggest a high ISO film stock, while the wide aperture lens creates a motion blur effect, enhancing the candid and natural documentary style.

The model understands photographic language at a deep level. You can reference specific film stocks, lens characteristics, and lighting setups, and it responds with images that feel like they came from that exact equipment.
A woman standing in a Tokyo alleyway at dusk, neon signs reflecting off wet pavement. Shot on expired Kodak Portra 800, pushed two stops. The tungsten light from a ramen shop spills warm orange across her face while the neon casts cool cyan highlights on her hair. Visible grain, halation around the light sources, slightly lifted blacks. She’s mid-step, caught between two worlds of color.

It’s not just portraits. Landscapes, still lifes, architectural photography — the model handles all of them with a level of taste that feels intentional rather than generic.
Aerial photograph of Iceland’s glacial rivers meeting volcanic black sand, creating abstract branching patterns that look like veins or lightning. Taken from 3000 feet during golden hour, the water channels glow turquoise against the obsidian sand. The scale is impossible to determine — it could be a microscope image of capillaries or a satellite photo of a delta. Large format camera, extreme sharpness, no horizon line.

Still life of a half-eaten pomegranate on a rough limestone surface, lit by a single shaft of afternoon light from a high window. Renaissance chiaroscuro lighting — the seeds glisten like rubies against the deep shadow. A few seeds have rolled across the stone, leaving tiny red trails. The mood is somewhere between Caravaggio and a modern food editorial.

An elderly fisherman mending nets on a weathered wooden dock at dawn. His hands are the focal point — scarred, sun-darkened, moving with practiced precision through the pale blue nylon mesh. Behind him, the sea is a soft gradient from pewter to rose gold. Shot on a Leica M with a 50mm Summilux wide open, the bokeh turns the harbor lights into perfect circles. The image has the quiet dignity of a Sebastião Salgado portrait.

A brutalist concrete building reflected in a perfectly still puddle after rain, creating a Rorschach-like symmetry. The building’s geometric facade — repeating rectangular windows and raw concrete — doubles into an abstract pattern. A single figure with a red umbrella walks along the edge, the only color in an otherwise monochrome scene. Overcast sky, flat diffused light, architectural photography with a tilt-shift lens effect on the edges.

Example-based editing
This is the feature that’s hardest to explain and most fun to use.
Instead of trying to describe a complex edit in words, you show the model what you want. Give it a before/after pair — Image 1 and Image 2 — and then a third image. The model figures out what changed between the first two and applies the same transformation to the third.
Here’s an example. We start with a plain white ceramic mug, then show the model what that mug looks like with Japanese kintsugi gold-crack repair (without any textual description). Then we give it a completely different object — a ceramic vase — and ask it to apply the same transformation:
Reference the change from Image 1 to Image 2, apply the same operation to Image 3
Image 1

Image 2

Image 3

Result

The model learned “add gold-filled cracks in a kintsugi pattern” from the mug pair, then applied the same treatment to the vase — without us ever having to describe what kintsugi looks like in words.
This works for all kinds of transformations:
- Material swaps: Show wood → marble on one object, apply to another
- Scene changes: Show day → night in one photo, apply to a completely different location
- Style transfers: Show a photograph converted to a woodblock print, apply the same artistic transformation to new scenes
Style transfer
The power here is that you don’t need to figure out the right words to describe “that exact Hiroshige color palette.” You show it.
Reimagine this scene as a traditional Japanese Ukiyo-e woodblock print in the style of Hiroshige — flat perspective, bold outlines, limited color palette of indigo, vermillion, and ochre.
Input

Output

Transform the color grading to match the following — saturated teal shadows, warm amber highlights, soft diffusion.
Input

Output

Logical reasoning
Most image models treat your prompt as a bag of keywords. Seedream 5.0 actually reasons through what you’re asking.
A Rube Goldberg machine: a marble rolls down a wooden ramp, hits a row of dominoes, the last domino pulls a string that tips a watering can, the water fills a small cup on a balance scale, which lowers and pulls a lever that rings a tiny brass bell. Every component casts physically correct shadows. The marble is mid-roll. The machine sits on a drafting table with visible grid paper underneath. Cross-hatching style, like a patent drawing brought to life.

It extends to understanding physical objects at a mechanical level:
An antique pocket watch disassembled and laid out on black velvet in an exploded-view arrangement. Every gear, spring, escapement wheel, balance cock, and jewel bearing is visible and correctly positioned relative to where it would sit in the assembled movement. The mainspring is partially uncoiled. Tiny labels in copperplate script identify each component. Museum conservation photography style.

Multi-step reasoning with image inputs
Give the model two images and a complex instruction, and it can reason through multi-step operations. Here we give it a mixed bouquet and three empty vases, and ask it to sort the flowers by type:
Classify the flowers from Image 1 by variety and arrange them separately into the three vases from Image 2. Roses in the first vase, sunflowers in the second, lavender in the third.
Image 1

Image 2

Result

The model has to identify each flower type, group them, then distribute them into the correct containers…all from a single prompt.
Check out some more cool things you can try:
Generate what they will look like when they grow up.
Input

Output

The complete metamorphosis of a monarch butterfly in four stages arranged left to right: a striped caterpillar eating a milkweed leaf, a jade-green chrysalis with gold spots, the chrysalis cracking open with wings partially visible, the fully emerged butterfly with wings spread. Each stage annotated with elegant serif labels. Rendered in the style of Maria Sibylla Merian’s 17th-century naturalist illustrations but with modern scientific accuracy.

Precise instruction following
Seedream 5.0’s instruction following is noticeably tighter than previous versions. When you say “blue jacket,” you get a blue jacket — not purple, not teal. When you specify spatial relationships, quantities, or specific details, the model respects them.
This matters most for complex compositions with lots of specific requirements:
A photorealistic cluttered office desk of a senior software engineer. An open MacBook Pro displays a terminal with green-on-black code scrolling. A ceramic mug reads “console.log(‘coffee’)” in monospace font, steam curling up. An open O’Reilly book shows a Venn diagram of three overlapping circles labeled ‘Frontend’, ‘Backend’, ‘DevOps’. Three Post-it notes form a kanban board: yellow (To Do), orange (In Progress), green (Done), each with tiny handwritten tasks. A mechanical keyboard with custom keycaps, a smartphone showing a Slack notification. A small succulent in a concrete pot shaped like a skull. Golden hour light rakes across from the right, casting long shadows.

That’s a prompt with over a dozen specific requirements — text on the mug, labeled diagram in the book, color-coded Post-it notes, a skull-shaped pot. The model tracks all of them.
The model can also handle visual cues. If your image has arrows, bounding boxes, or colored regions marking specific areas, you can reference them in your prompt:
Furnish this loft according to the spray-painted markers. Place a large abstract expressionist painting where the red rectangle is on the wall. Place a mid-century modern leather sofa where the blue rectangle is on the floor. Hang a brass globe pendant light where the yellow circle is on the ceiling. Remove the spray paint markers. Keep the industrial character.
Input — loft with spray-painted markers

Output

Domain knowledge
Seedream 5.0 has deep built-in knowledge across several professional fields. This goes beyond “knowing what a building looks like” — it understands the structure and conventions of technical content.
Feed it a floor plan sketch and it generates realistic visualizations that respect the spatial layout:
Based on this floor plan, generate a photorealistic rendering of the interior. A Japanese-inspired minimalist house with warm hinoki wood tones, a central courtyard garden visible through floor-to-ceiling glass with a single red maple tree, tatami flooring, an engawa wooden veranda with dappled afternoon light. Sliding shoji screens partially open. Match the layout exactly.
Input — architect’s sketch

Output

The model can generate accurate scientific illustrations that would normally require specialized knowledge and hours of careful work:
A detailed cross-section diagram of a coral reef ecosystem, scientific illustration style. Below: volcanic basalt foundation. Middle: calcium carbonate reef structure with visible fossil layers. Upper reef: a thriving ecosystem with labeled species — staghorn coral, brain coral, clownfish in an anemone, a moray eel in a crevice, parrotfish grazing, sea turtle swimming above. Above the waterline: a small tropical island with palm trees. Depth markers on the left edge. Precise linework with watercolor fills, labels in elegant serif type. Published in Nature magazine style.

Give the model a photo of food and ask it to annotate with nutritional information:
Identify each dish in this mezze spread and add elegant hand-lettered calligraphy annotation cards next to each one, showing the dish name and calorie count per 100g. Cards on cream-colored stock with burgundy ink.
Input

Output

Text rendering
Seedream has been strong at text rendering since version 3.0, and 5.0 continues that tradition. Use double quotation marks around text you want rendered in the image for best results:
A large-format typographic poster for a fictional jazz festival. At the top in bold condensed sans-serif: “BLUE NOTE SESSIONS” in deep navy. Below in elegant script: “Summer 2026 — Central Park, New York.” The poster lists four performers: “Miles Ahead Quintet / Saturday 8PM” “Sarah Chen Trio / Saturday 10PM” “The Monk Revival / Sunday 7PM” “Coltrane Legacy Orchestra / Sunday 9PM”. A stylized golden saxophone silhouette runs vertically along the right edge. Background is a gradient from midnight blue to warm amber. At the bottom in small caps: “Tickets at bluenote.nyc — All ages welcome.”

That’s a poster with multiple typefaces, mixed case, punctuation, specific performer names, and a gradient background — all rendered accurately. The model handles multilingual text well too. If you need Chinese, Japanese, Korean, or other scripts rendered accurately, Seedream 5.0 can do it.
Multi-image generation
Like Seedream 4.5, the 5.0 model can generate multiple related images in one go. Ask for “a series” or “a set” or specify a number, and it produces images with consistent style and character continuity.
A cinematic 2x2 storyboard grid. Panel 1: Interior of a derelict space station, a lone astronaut floats through a corridor where vines have grown through cracked hull panels, bioluminescent fungi glow blue. Panel 2: She discovers a sealed laboratory door with a pulsing green light behind frosted glass, reaching for the manual release lever. Panel 3: The door opens to reveal a thriving garden, a small tree growing in zero gravity with roots spiraling outward like a jellyfish, butterflies frozen mid-flight. Panel 4: Close-up of her face through the helmet visor, tears floating as small spheres, the garden reflected in her eyes. Consistent character design, anamorphic lens flare, Ridley Scott color palette.

A comprehensive brand identity flat-lay for a specialty coffee roaster called “ALTITUDE”. Arranged on dark slate: matte black coffee bags with the “ALTITUDE” wordmark in embossed gold foil and an elevation contour line logo, business cards (front and back), a ceramic pour-over dripper with the contour logo etched in, branded kraft paper stickers, a linen tote bag with screen-printed logo, a metal travel tumbler with laser-engraved branding, and a booklet titled “Brewing Guide.” Palette of black, gold, and natural kraft. Overhead product photography, even studio lighting.

This is great for storyboarding, brand identity packages, emoji sets, and any scenario where you need a cohesive collection.
Getting started with the API
Here’s how to run Seedream 5.0 using JavaScript and the Replicate API:
And in Python:
Prompting tips
A few things we learned while testing:
- Use natural language, not keyword lists. “A girl in a lavish dress walking under a parasol along a tree-lined path, in the style of a Monet oil painting” works better than “girl, umbrella, tree-lined street, oil painting texture.”
- Use double quotes for text rendering. If you want specific text in your image, wrap it in double quotation marks: Design a poster with the title "Seedream 5.0".
- Be specific about what to keep. When editing, tell the model what shouldn’t change: “Replace the hat with a crown, keeping the pose and expression unchanged.”
- Use visual markers for complex edits. Draw arrows, boxes, or colored regions on your input image to indicate exactly where changes should happen.
- Specify your use case. Telling the model “Design a logo for a gaming company” gives better results than just describing the visual elements.
- For example-based editing, show don’t tell. When the transformation is hard to describe in words, provide a before/after example pair.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み