Google publishes three separate prompting documents for Nano Banana Pro, and they do not agree on what to call the parts of a prompt. Most guides online quietly pick one version and present it as the only one, which is a small loss until you go looking for a template and find three.
This page lines all three up, gives you the framework Google repeats most often, and then does what none of the official pages do: shows the same prompt before and after a rewrite.
What the three official guides actually say
| Official page | Elements it names | Strongest at |
|---|---|---|
| The seven-tips post on blog.google | Subject, composition, action, location, style, editing instructions | Camera, lighting, text and the role of each reference image |
| The Cloud “ultimate prompting guide” | Subject, action, location/context, composition, style | Per-task formulas, including edits and reference images |
| The DeepMind prompt page | Style, subject, setting, action, composition | Editing an existing image, one element at a time |
The naming does not line up
Subject, action and style appear on all three lists. Composition appears on all three and moves position; location becomes “setting” on the DeepMind page. The difference is cosmetic, and the practical reading is identical: name the subject, say what it does, place it, frame it, set the style.
One gap matters more than the wording. Every official list asks for editing instructions, and that is the slot beginners leave empty.
The rules that carry every prompt
Google’s Cloud guide states the operating principle plainly: start a prompt with a strong verb that tells the model the primary operation you want.
Two more rules come from the same page and both save rounds. Be specific, because details are what the model reasons with. Frame positively — the guide’s own example is that “empty street” lands and “no cars” usually does not.
Positive framing beats negatives
Watch what changes when you delete a prohibition and describe the scene instead.
Weaker: a coffee cup on a table, no clutter, no people, no text
Better: a single ceramic coffee cup on a bare oak table, morning light from the left, empty room behind it
The Cloud guide describes the loop as conversational: refine with follow-up prompts rather than starting over. Change one thing per round, so you learn which word caused the shift.
Framework one: text to image

The formula Google publishes for generation without reference images is [Subject] + [Action] + [Location/context] + [Composition] + [Style]. Read it as five slots you fill in order, not five sentences to write.
Subject: a ceramic coffee cup, unglazed base
Action: resting still, steam rising in a thin line
Location: bare oak table in an empty room
Composition: centred, slightly above eye level, negative space to the right
Style: editorial product photography, 50mm, soft key light from the upper left
Notice where the specificity sits: not in adjectives, but in the named object, the named surface and the named light direction. Those are what the model can act on.
Framework two: prompts built on reference images
When you supply images, Google’s formula changes shape to [Reference images] + [Relationship instruction] + [New scenario]. The relationship instruction is the whole trick, because it tells the model what each input is for.
The Cloud guide’s own example attaches a napkin sketch as the structure and a fabric sample as the texture, then asks for a high-fidelity 3D armchair render in a sun-drenched minimalist living room. Structure from one image, surface from another, and a new scene described in words only.
Label every input. An unlabelled second image leaves the model guessing which one to copy, and guessing is the most common cause of a plausible but wrong result.
Framework three: editing prompts
Google’s editing approach is what the guide calls semantic masking: you define the mask with words instead of drawing it. The instruction the page prints in bold is to be explicit about what to keep exactly the same.
Keep the cup exactly as it is: same shape, same glaze colour, same handle, same two chips on the rim. Change only the setting. Place it on a matte grey studio surface against a seamless backdrop, lit from the upper left with one large softbox. 85mm, f/8, neutral white balance, no props, no text, no watermark.
The official example for this pattern is as short as it gets — “remove the man from the photo”. Short works when the change is unambiguous and the rest is obviously untouched. It stops working the moment the model has to decide what “the rest” means.
Prompting text, fonts and other languages

Text is where this model separates itself, and Google documents the handling rules. Put the words you want rendered in quotation marks. Name the typography, describing it as a bold sans-serif or neon cursive signage rather than hoping the model guesses. Google states the model supports multilingual text generation in over 10 languages, and the translation workflow is to ask for the localisation rather than for a fresh image.
There is a text-first trick in the same guide worth stealing: work out the wording in conversation first, then ask for an image with that text. Generating the sentence and painting it in one step splits your attention across two problems.
Worked example: five prompts rewritten
The rewrite below is the same request twice. The first version is what people type; the second fills the slots the official guides name.
Vague: make this product photo look professional
Buildable: Keep the bottle exactly as it is, including the label text and the condensation on the glass. Change only the setting and light. Place it on a wet dark slate surface with a seamless charcoal backdrop, lit by one large softbox from the upper left plus a soft rim light on the right edge, shadow falling gently right. 85mm lens, f/8, editorial product photography, no props, no added text.
Two more rewrites follow the same two moves.
- “Make it more cinematic” becomes the actual variable: golden-hour backlighting creating long shadows, cinematic colour grading with muted teal tones.
- “Remove the background” becomes the keep-list first: keep the subject, its proportions and its edge detail; replace only the backdrop with a seamless light grey.
The pattern is hard to miss once you see it twice. The vague prompt names a feeling. The buildable prompt names a fixed object, a changed variable, and a reason the two are separable.
What Nano Banana Pro still gets wrong
Google’s own model page lists what the model still struggles with: small faces, accurate spelling and fine details. Its prompting page adds that even short prompts generate high-quality images, but control comes from adding detail bit by bit.
Treat that as a checklist. If your image depends on a face at thumbnail size, a long string of text or a fine repeating pattern, look at the full-size result before you publish it.
Cost is the short answer here: the official prompting pages carry no per-image price and no free quota, so this article prints no figure. For the workflow around the prompt, the getting-started guide walks through a first image, and the API guide shows where a prompt goes in a request.
Frequently asked questions
Which framework should I use? The one Google publishes for your task. Text-to-image, reference images and editing each have their own formula, and they differ in the middle rather than at the ends. The topic hub collects the model pages in one place.
Do I have to write the prompt in English? Google documents text generation in over 10 languages and treats localisation as a first-class task. Specificity matters more than the language.
Are negative prompts supported? The Cloud guide recommends positive framing and gives “empty street” over “no cars”. That is guidance, not a documented negative-prompt syntax, and the official pages publish no such syntax.
How many reference images can I include? The official announcement says up to 14 images in a composition, varying by surface. The API guide shows where that input goes.
Should I compare it with GPT Image before writing prompts? If the choice is still open, the model comparison sorts the two by task, which is cheaper than writing prompts for the wrong model.
Fill the five slots, name what must not change, then move one element per round. That is the whole method, and it is the part the official pages leave you to assemble yourself.