Google Veo 3.1 is the video model Google built on Veo 3, and the quickest way to understand it is by what it changes: you brief a shot the way you would brief a camera operator, and you ask for the sound in the same prompt.
This guide covers where it runs, how a first clip comes together, one worked prompt, and what Google has not published.
What Veo 3.1 is
Google’s model page presents Veo 3.1 as its leading video generation model, built for filmmakers and storytellers, and says it handles text-to-video, image-to-video and text-to-audio-plus-video. The October 2025 announcement adds the key detail: it builds on Veo 3 with stronger prompt adherence and better audiovisual quality on image-to-video.
Veo 3 is a different model, and Google’s pages carry no side-by-side table of the two; what they publish is the list below.
| Capability | What the official pages say |
|---|---|
| 1080p and 4K output | A cut-ready file and a large-screen master from one clip |
| Native audio | Dialogue, ambience and effects generated with the video |
| Ingredients to Video | Reference images hold a character or a style, now with audio |
| Frames to Video | A starting and an ending image, bridged by a transition |
| Extend | Continues from the last second of the clip, a minute or more |
| Insert and Remove | Adds or removes an element; Remove was still “soon” in 2025 |
Where you can use Veo 3.1 today
Google’s model page lists six entry points: the Gemini app, Flow, Google Vids, AI Studio, the Gemini API and the Gemini Enterprise Agent Platform. A January 2026 update adds YouTube Shorts.
| Surface | What the official pages attach to it |
|---|---|
| Gemini app | Veo 3.1, improved Ingredients to Video, native portrait output |
| Flow | The full creative set, including Extend and Insert, plus 1080p/4K |
| Google Vids | Veo inside a work-video editor, Ingredients to Video rolling out |
| Gemini API and Vertex AI | The same model for code, with 1080p and 4K upscaling |
| YouTube Shorts | Ingredients to Video with native 9:16 for short-form |
To call it from code instead, the Veo 3.1 API guide covers how a request reaches the model.
One caution that saves money: plenty of sites resell Veo 3.1 on a credit plan, and those prices are theirs, not Google’s.
Your first video, step by step

The first run confirms the loop; it is not the run for your best clip.
- Write the shot as one sentence. Subject, action, then where it happens.
- Say what the camera does. A locked-off frame and a slow push are different asks.
- Ask for the sound in the same prompt. Dialogue, ambience and effects all belong there.
- Name what must not change. That clause separates an edit from a reroll.
- Change one thing per round. Rewriting everything hides the word that caused the shift.
Three rounds of one change each beat one longer prompt. New to the app? The Gemini getting-started guide covers it.
Worked example: a vertical clip where the maker stays the same
The job is a short clip for a mobile feed: a maker turns a ceramic mug toward the lens and says one line. The hands and the mug have to match the stills, which is what Ingredients to Video is for. Google announced native 9:16 here in January 2026, so vertical framing is an official option, not a crop.
Ingredients: image one is the maker, keep her face, hair and apron unchanged; image two is the mug, keep the glaze and the two chips on the rim.
Shot: a slow push-in from a medium shot to a close-up, camera at chest height, slight handheld drift, shallow depth of field.
Action: she picks up the mug, turns it a quarter turn toward the lens, and sets it back down.
Dialogue: she looks at the lens and says, "Every one of these comes out a little different."
Ambience: quiet workshop room tone, a kettle ticking somewhere behind her.
SFX: the dry knock of ceramic on wood as the mug lands.
Light: one large window on the left, soft and cool, plus a warm lamp behind her to separate her from the wall.
Format: native 9:16, 1080p, no text, no watermark.
Each block answers a different question. The ingredient lines lock identity before anything moves. Dialogue, ambience and effects are separate asks, so a missing knock does not take the line with it.
Case two moves one block: “Keep the maker, the mug and the lighting as they are. Replace the workshop with a plain concrete wall and morning light from a high window. Same push-in, same line, same sound.”
The ingredient images matter as much as the prompt, and Google’s tip is to build them with Nano Banana Pro first — the Nano Banana Pro guide covers that.
Watch the result at full size before publishing.
A prompt structure that holds up
The two cases share a skeleton, which is what makes a good result repeatable rather than lucky.
The blocks worth writing every time
Google generates picture and audio together natively, so a prompt that describes only pictures answers half the request. The blocks below come from the official capability descriptions; for the picture half, the Gemini prompting guide covers the craft.
Subject and action: who or what is on screen, and what changes during the clip
Camera: framing, movement, and where the camera sits
Light: source, direction and colour
Dialogue: the speaker, then the line inside quotation marks
Ambience and SFX: the room first, then the sound tied to one action
Constraints: what must stay identical to the reference
Format: aspect ratio, resolution, and what to leave out
Where the audio block changes the take
Dialogue is the block most first prompts skip, and it carries a published caveat. Google states that natural and consistent spoken audio, particularly for shorter speech, remains an area of active development. Keep lines short, name the speaker, and use quotation marks.
Limits, cost and what Google has not published
Google names its limits on its own pages: spoken audio is the one it repeats, upscaling to 1080p and 4K is confined to Flow, the Gemini API and Vertex AI, and output carries a SynthID watermark. Insert shipped in Flow; Remove was still “soon” in October 2025, when neither was in the API.
Cost is where the honest answer is short: Google’s pages list no per-second price, no free quota and no subscription tier, so this article prints no figure.
Frequently asked questions
Do I have to pay to try it? The pages we checked publish no free quota and no price, so no reliable figure exists. The Gemini app is the shortest path to a first clip.
How long can one clip be? Google’s benchmark footnotes use eight-second clips and publish no product-wide duration. Extend continues from the last second and can reach a minute or more.
Can I make vertical video? For Ingredients to Video, yes: Google announced native 9:16 in January 2026.
Why is my dialogue audio uneven? Spoken audio for short segments is the limitation Google names.
Learn the loop once — one change per round, a named speaker, a locked reference — and the model stops being a slot machine.