FLUX.2 comes from Black Forest Labs and Nano Banana Pro is Google’s Gemini 3 Pro Image. Both vendors publish model pages, and both answer some of the same questions: how many reference images the model accepts, what resolution it returns, what it does with text. Those are the questions this page puts side by side.
Only claims the vendors have published appear below, and no per-image price is quoted for either model, because a price one vendor has not published is not a price this page can state.
The two models, as each vendor describes them
Black Forest Labs documents FLUX.2 as “production-grade AI image generation and editing” with 4MP photorealistic output and multi-reference control, and as a family rather than one checkpoint. The overview names five tiers: [max] for top-tier quality with grounding search, [pro] for production at scale, [flex] for typography and small details, [klein] for sub-second inference with open weights, and [dev] as the local-only open-weight tier.
Google’s page opens with one sentence: “Nano Banana Pro (Gemini 3 Pro Image) is an image generation and editing model,” built on Gemini 3, and lists the surfaces it runs on — the Gemini app, Google AI Studio, the Gemini API and the Gemini Enterprise Agent Platform.
What each side publishes about itself
Most differences below follow from that split.
Side by side on the attributes both vendors publish
| Attribute | FLUX.2 (Black Forest Labs) | Nano Banana Pro / Gemini 3 Pro Image (Google) |
|---|---|---|
| Output resolution | Up to 4MP; sides in multiples of 16, minimum side 64px, up to 2MP recommended | Output at 1k, 2k or 4k resolution |
| Reference images | Up to 8 via the API and 10 in the playground for [max], [pro] and [flex]; up to 4 for [klein]; recommended 6 for [dev] | Up to 14 images in a composition, varying by surface |
| Character consistency | Multi-reference documented for identity across scenes | Up to five characters and fourteen objects |
| Text rendering | “Complex typography, UI mockups now work reliably in production” | “State-of-the-art text rendering in over multiple languages” |
| Downloadable weights | [klein] 4B under Apache 2.0, [klein] 9B under a non-commercial licence, [dev] open weights local-only | None on the pages referenced here |
| Provenance marking | Not documented on the pages referenced here | All images watermarked with SynthID technology |
| Price basis | Published per output megapixel; [pro] text-to-image from $0.03 | No per-image price published on the pages referenced here |
| Head-to-head benchmark | Not published | Not published |
How to read that table
Three kinds of row sit in it: rows both vendors publish, rows only one publishes, and the benchmark row neither publishes. Anyone handing you a winner is quoting a number neither vendor signed.
Reference images and consistency
Black Forest Labs labels the ceilings per tier and gives the arithmetic: the [pro] budget is 9MP shared between input and output, so a 1MP output leaves room for 8 references and a 2MP output for 7. Google’s prompting guide asks for the same discipline in its own words: “When using uploaded images, clearly define the role of each.” Its consistency figure has a different shape — up to five characters and fourteen objects across one workflow.
Both treat an unlabelled reference image as a mistake, and neither publishes how often that consistency holds in practice.
Text rendering and typography
Typography is where both vendors make their loudest claim, and the claims are not the same kind of claim.
Black Forest Labs puts text at the centre of the model page and documents a tier for it: [flex] is described as specialized for typography and for keeping small details. Google documents “state-of-the-art text rendering in over multiple languages” with a localisation workflow, and its own example of translating the text on three cans into Korean.
One asymmetry is worth stating plainly. Black Forest Labs publishes no accuracy figure for text rendering in any language. Google’s model page does publish its own single-line error-rate chart, but the FLUX entry in it is FLUX.1 Kontext [max] — the previous generation — so it is not a FLUX.2 figure, and it is Google’s own measurement.
Open weights, deployment and provenance
This is the axis where the two documentation sets diverge most, and it is usually compressed into the word “open”.
- [klein] 4B is fully open under Apache 2.0, and [klein] 9B ships under the FLUX Non-Commercial License.
- [dev] is documented as open weights for local development and research, and is not offered on the public API.
- [max], [pro] and [flex] are hosted tiers; the FLUX.2 topic hub and the FLUX.2 API getting-started guide cover those routes.
- Google publishes no downloadable weights on the pages referenced here, and marks every image with SynthID; Black Forest Labs documents no equivalent marking.
Open weights is not the same as free to use: the licence decides what a commercial deployment may do.
Worked example: one brief, written for each model
The brief is the same on both sides: put a product into a scene and keep the label exactly as supplied. Each vendor publishes its own example.
Black Forest Labs publishes this one on the model page, with the result in the figure above.

The jar in image 1 is filled with capsules exactly same as image 2 with the exact logo
The published result is a studio shot of the jar in which the capsules carry the supplied logo unchanged. Notice what is missing: no style clause and no lighting clause, because both come from the reference images. The whole instruction is two inputs and the relationship between them.
Google publishes the other example on its prompting page.
translate all the English text on the three yellow and blue cans into Korean, while keeping everything else the same
The result beside it keeps the cans and their layout and changes the language on the labels. The closing clause is the interesting part.
What the two published results show
Both examples are reference-driven, and both spell out the constraint that must not move — “exactly same” on one side, “keeping everything else the same” on the other. That is the shared technique, and it is where the published material stops: neither vendor’s example runs the other model’s prompt, so neither supports a claim about which one preserves a label better.
What neither vendor publishes
- Neither publishes a head-to-head benchmark or a preference score for these two models.
- Black Forest Labs publishes no text-rendering accuracy figure by language; Google’s chart compares against the previous FLUX generation.
- Black Forest Labs documents no provenance watermark; Google publishes no downloadable weights.
- Neither publishes a per-image price for the other’s model. Third-party resellers set their own numbers, and those belong to the reseller.
The FLUX.2 topic hub collects the Black Forest Labs pages, and the Nano Banana Pro hub collects Google’s.
Frequently asked questions
Which model is better? The published documentation does not answer that. One is a tiered family with open-weight options; the other is a hosted model with world knowledge and SynthID marking.
Can I run either one on my own hardware? FLUX.2 yes, for two tiers: [klein] under Apache 2.0 or a non-commercial licence, and [dev] with open weights and no hosted API. Google publishes no downloadable weights.
Do both accept reference images? Yes. FLUX.2 documents up to 8 through the API and 10 in the playground for [max], [pro] and [flex], 4 for [klein], and 6 recommended for [dev]. Google documents up to 14 images in a composition.
Where should I start? The FLUX.2 getting-started guide builds a first image; the Nano Banana Pro hub starts on Google’s side.
Compare the rows both vendors filled in, read the single-vendor rows as directions rather than verdicts, and decide on your own constraints.