Google’s prompting page for Veo lays out seven elements to write against, and the whole method fits in one sentence from that page: “The more detail you add, the more control you’ll have over the final output.”

There is a catch worth stating early, because most guides skip it. That page is titled for Veo 3, not Veo 3.1. Google’s published description of Veo 3.1’s prompting is one clause — stronger prompt adherence — so everything below treats the seven elements as the structure Google documents and marks the rest as inference.

What Google actually publishes about prompting

The Veo 3.1 announcement is generous about capability and thin about prompting. It describes richer audio, more narrative control and enhanced realism, and adds that the model builds on Veo 3 with stronger prompt adherence and improved audiovisual quality when turning images into videos.

That is the entire official statement on how Veo 3.1 responds to prompts. No formula, no template, no word limit. The separate prompt guide carries the structure, and the model page carries the capability list: 1080p and 4K output, camera and motion controls, scene extension, and character consistency through reference images.

The seven elements Google names

Shot framing and motion. Style. Lighting. Character descriptions. Location. Action. Dialogue. The guide frames framing and motion as a question rather than a command — how you want the output framed and how the camera should move during the shot — which is a useful hint that camera direction is a first-class input, not decoration.

Character descriptions come with a worked contrast: “a woman in her twenties with wavy brown hair and light freckles” produces more specific results than “a brown-haired woman”. Same subject, different amount of information the model can act on.

The two elements the community version drops

Most third-party guides compress the seven into five blocks: subject, camera, lighting, audio, technical. That mapping is close but it merges two things Google keeps separate. Dialogue gets its own element on the official list, because Veo can generate speech and you can either give characters a topic or give them the exact words. Audio gets its own section too, and the guide says you can put audio cues inline or in a separate section.

Keep those two separate in your own template and you stop writing prompts where the camera is perfect and nobody speaks.

Official seven, community five: what the difference is

Community blockOfficial element it maps toWhat changes if you keep them apart
SubjectCharacter descriptionsNamed physical traits beat relationship words
CameraShot framing and motionFraming and movement are one element, not two
LightingLightingDrop this block and the shot reads flat
AudioDialogue, plus audio cuesSpeech and ambience are separate decisions
TechnicalNo official elementLength and resolution come from the surface, not the prompt

The last row is the honest one. Google publishes no prompt-level parameter block, so treat duration and resolution as settings rather than prompt text.

Worked example: a text-to-video prompt, line by line

A collage of official Veo 3.1 output, including a dancer in a studio, a candle scene and a rider crossing a field
Official Google Veo examples Google

Google’s own sample prompt on the guide page is a good benchmark for density. It describes a medium shot of an old sailor whose knitted blue hat casts a shadow over his eyes, with a thick grey beard obscuring his chin, holding a pipe and gesturing toward the churning grey sea beyond the railing — and then he speaks, in quotation marks.

Read what is doing the work. A shot size, a named garment, a named colour, a specific shadow, a specific gesture, a described sea. Every one of those is something a camera can see.

Framing and motion: medium shot, slow push in, camera at chest height
Style: documentary naturalism, handheld, 35mm
Lighting: overcast daylight from the left, no fill, deep shadow under the hat
Character: a man in his late sixties, knitted blue sailor hat, thick grey beard
Location: wooden trawler deck, churning grey sea behind the railing
Action: he gestures with a pipe toward the water, then steadies himself on the rail
Dialogue: "This ocean, it's a force, a wild, untamed might."

Seven lines, seven elements, nothing that the frame cannot show.

Case two: keeping a character and adding dialogue

The second prompt type is the one that produces the most retries: a recurring character across several clips. The model page lists character consistency through reference images as a supported control, so the prompt’s job changes from describing the person to describing what to do with them.

Use the attached reference image for the character. Keep her face, hair and coat identical to the reference. Change only the setting: an empty platform at dawn, light mist, no other people. Framing: wide shot, she walks toward camera and stops at the platform edge. No dialogue.

Two things make this version work. The reference image carries identity, so the prompt does not spend words re-describing a face it cannot improve on. And the closing instruction says what the shot does not contain, which keeps the model from filling the platform.

Three fixes for the prompts that fail most

The camera wanders. Name the framing and the movement as one instruction, and give the movement a direction and a rate rather than an adjective. “Slow push in” constrains the shot; “cinematic” does not.

The scene looks clean but empty. Look for a prose prompt that described mood and skipped the physical world. Add the specific objects the camera should see — a surface, a garment, a light source.

When audio does not appear

The guide’s advice is to define the sounds you want explicitly, matching audio to visuals, either inline or in a separate section. Silence usually means the prompt described only the picture.

One caution that belongs in any guide to dialogue: Google’s model page states that creating videos with natural and consistent spoken audio, particularly for shorter speech segments, remains an area of active development. Write the line, then check the result rather than assuming it.

What Google has not published

Four gaps, all of which you will see filled in confidently elsewhere.

No prompt length or token limit is published. No negative-prompt syntax is documented — the guide discusses what to include, not what to forbid. The order in which elements should appear is unspecified. And generation length is not a prompt parameter: the benchmark footnote on the model page simply states that Veo videos are eight seconds long, and Flow’s Extend is described as reaching a minute or more.

None of that blocks a good prompt. It does mean a confident template claiming an official word count or a negative-prompt field is describing its author’s workflow.

Frequently asked questions

Is there an official Veo 3.1 prompt guide? The published prompt guide is titled for Veo 3. Veo 3.1’s own documentation describes stronger prompt adherence and leaves the structure to that page.

Should I write prompts in English? Google publishes the guide and its examples in English, and publishes no language rule. Specificity matters more than the language you write it in.

How long can one generation be? The model page’s benchmark footnote says Veo videos are eight seconds long, and the Flow announcement describes Extend as creating longer videos, even lasting a minute or more. No full menu of length options is published.

My character changes between clips. What do I fix? Move identity into a reference image and state that it must stay identical, the way the second prompt above does. The getting-started guide walks through the same idea for a first clip.

How much does this cost per clip? Neither the model page nor the announcement carries a price. The API and pricing guide covers the numbers Google does publish.

Where can I use Veo 3.1? Google names Flow, the Gemini app, the Gemini API and Vertex AI. For prompt-writing practice in a text model first, the Gemini prompting guide covers the general technique.

Write all seven elements once, keep speech and ambience as separate decisions, and change a single line when a shot misses. The topic hub collects the pages this guide draws from.