Google Nano Banana Pro is the image model Google built on Gemini 3, and the fastest way to understand it is by what it changes about your habits: you describe a scene instead of stacking keywords, and you edit by describing only the difference.
This guide covers where the model runs today, how a first image comes together, and one worked edit with the prompt in full. Everything factual comes from pages Google publishes; where Google has published nothing, this article says so.
What Nano Banana Pro is
Google’s DeepMind model page describes the model as built on Gemini 3, for creating and editing images with what it calls studio-quality precision and control. The official prompting post calls it Google’s most advanced image model to date.
The earlier Nano Banana is a different model, and the Google pages we verified print no side-by-side comparison of the two, so treat those tables elsewhere as their authors’ work. What Google does publish is a capability list.
| What Google publishes | What it means at your desk |
|---|---|
| Built on Gemini 3 | More world knowledge, so a described object can be plausible |
| Output at 1K, 2K or 4K | A draft and a printable file from the same prompt |
| Up to 14 input images in a composition, varying by surface | A product, a model and a background at once |
| Five characters and fourteen objects held consistent per job | A small cast handled well, not a crowd |
Where you can use it today
Google’s announcement says the model is available now in the Gemini app and starting to roll out in AI Studio, Vertex and more. The surface you are on decides what you get, so start where the model already is.
If you would rather call it from code, the Nano Banana Pro API guide covers the key, the request body and what comes back.
One caution that saves money. Plenty of sites offer Nano Banana Pro on a credit plan. Those prices belong to those sites, not to Google, and the model behind them may not be the one described on Google’s pages.
Your first image, step by step

The first run confirms the loop; it is not the run where you produce your best image.
- Write the scene as one sentence. Subject, then action, then location.
- Say what must not change. One clause is the difference between an edit and a reroll.
- Set the format first. Aspect ratio and resolution decide whether the picture is usable.
- Read the result honestly. Note which element drifted, not just whether you like it.
- Change one thing per round. Rewriting everything hides which word caused the shift.
Three rounds of one change each beat one round of a longer prompt.
Worked example: a flat product photo into a studio shot
The input is a phone photo of a ceramic mug on a kitchen table, hard window light from the left, clutter behind it. The goal is a version clean enough for a product page, without reshooting.
Keep the mug exactly as it is: same shape, same glaze colour, same handle, same two chips on the rim. Change only the setting. Place it on a matte grey studio surface against a seamless grey backdrop. Light it from the upper left with one large softbox, add a soft rim light on the right edge, and let the shadow fall gently to the right. Shot on an 85mm lens at f/8, product photography, neutral white balance, no props, no text, no watermark.
The result holds up because of how the prompt is split. The object survives untouched, glaze and chips included, because the prompt names them as fixed. Background, light and shadow are the only things that move.
Case two takes the same file further. Once the studio version looks right, ask for a lifestyle variant and leave the object alone: “Keep the mug, its proportions and its glaze identical to the reference. Replace the backdrop with a light oak tabletop and morning light from a window on the right. Same 85mm framing, no text, no watermark.” Two rounds, one object that never needs describing again.
Look at the result at full size before publishing. Google’s model page lists what it still struggles with: small faces, accurate spelling, fine details. A convincing thumbnail and a legible label at 4K are different tests.
A prompt structure that stays editable
The prompts above share a shape, and that shape is what keeps a result reproducible.
The six elements Google names
Google’s prompting post lists subject, composition, action, location, style and editing instructions. Read it as a checklist, not a rigid template: subject and action say what the image is about, composition and location place it, style sets the surface, editing instructions protect what is already correct. Most failed prompts are strong on subject and silent on that last item.
The advanced controls worth adding
Five more controls do the rest: aspect ratio, lens and lighting, text you want rendered verbatim inside quotation marks, factual constraints, and a stated role for each reference image. That last one matters most when you pass several, because an unlabelled second image leaves the model guessing which one to copy.
Subject: a ceramic mug, unglazed base, two chips on the rim
Composition: centred, slightly above eye level, negative space to the right
Action: resting still on the surface
Location: matte grey studio surface, seamless backdrop
Style: clean product photography, 85mm, soft key from upper left
Editing instructions: keep shape, glaze and chips identical to the reference
Text (verbatim): "HANDMADE"
Aspect ratio: 4:5 Resolution: 2K
Reference image 1: the product, keep the object unchanged
Reference image 2: background style only, do not copy its subject
Limits, cost and what Google has not published
Google describes the limits on its own pages. The DeepMind page says the model can still struggle with small faces, accurate spelling and fine details, and notes that output carries a SynthID watermark. On text, the published benchmark puts the single-line error rate mostly under 10%: a good number, not a promise about your poster.
Cost is where the honest answer is short. The official pages we verified carry no per-image price, no free quota and no subscription tiers, so this article prints no figure. Check the official Gemini API pricing page before you budget, and read any credit price on an aggregator site as that site’s own rate.
If you are choosing between models rather than learning one, the Nano Banana vs GPT Image comparison sorts them by task.
Frequently asked questions
Do I have to pay to try it? The Google pages we verified state no free quota, so no reliable number exists here. The Gemini app is the shortest path to a first image, and a credit offer on a third-party site is that site’s own plan.
Why does the text in my image come out misspelled? Spelling is a documented weak point. Keep rendered text short, put it in quotation marks, and make it large enough to survive at the size you display.
How many reference images can I pass? Up to 14 in a composition according to the official announcement, and the exact number varies by surface. The API guide shows where that input goes in a request.
Should I check the output before publishing? Yes. Generation is the first step, not the last, especially for anything carrying a face, a brand name or a price.
Learn the loop once, one change per round and a fixed object, and the model stops being a slot machine.