Generate video with Veo 3.1.
Google Veo 3.1 is DeepMind’s current video generation model. It turns a written prompt, a single image or a set of reference images into a short clip, and generates the dialogue, sound effects and ambient audio in the same pass as the picture.
Every capability, benchmark summary and access route on this page comes from Google DeepMind’s own Veo model page, and the showcase stills are the demo frames Google publishes on that page.
What Google’s own Veo demos render
These stills are the frames Google publishes beside its own demo clips in the model page’s showcase section, and each prompt below is the one Google wrote for that clip. They are the vendor’s own demonstrations rather than anything generated for this site.
The demo stills below belong to Google and are reproduced with credit, from the Veo model page. Google DeepMind’s Veo model page
Veo 3.1 in Google’s own words
Google introduces Veo 3.1 as its leading video generation model, “designed to empower filmmakers and storytellers”, and builds the page around a single claim: that video and audio now arrive together. The showcase, the control sections and the benchmark summaries below all come from that page.
Video and audio in one pass
Dialogue, sound effects and ambient noise are generated natively rather than dubbed on afterwards. The page’s own line for it is “video, meet audio”, and several of its demo prompts write the spoken part into the prompt itself.
Control at the level of the shot
Past a text prompt, the model accepts reference images for a scene, a character or an object, a style reference, a first and a last frame, and camera instructions such as zoom in or move right.
Edits rather than restarts
Google documents outpainting an existing frame, adding an object to footage and removing one from it, extending a clip from its final second, and driving a character’s performance with your own body, face and voice.
Resolutions aimed at an edit
Output is documented at 1080p and 4K. Google presents 1080p as the resolution for material that has to cut against other footage, and 4K as the one for texture and detail.
The controls Google documents
Google groups the model’s control surface under headings of its own and demonstrates each one with a clip. Read together they answer the practical question — what can I actually steer — and every item below is a heading on the model page rather than an interpretation of one.
Ingredients to video
Supply reference images of a scene, a character or an object and the shot is built around them. The page’s own prompts for this section ask for a commercial, a film trailer and a music video from the same set of references.
Match your style
A style reference carries the look, so an aesthetic that is hard to put into words — a paper diorama, a painting, a particular grade — can be handed over as an image instead of described.
Keep your characters consistent
Character references hold a subject’s appearance steady across separate shots, which is the piece that makes a multi-shot sequence possible rather than a set of unrelated clips.
Extend your scene
A clip can be continued from its last second while keeping the look and the sound consistent; the page shows prompts chained this way into one longer piece.
Camera controls
Framing and movement are set explicitly rather than asked for in prose. The page names four: move back, zoom in, move up and move right.
First and last frame
Two images are enough to define a transition, with the model filling in the motion between them.
Outpainting, adding and removing objects
Outpainting extends the frame outward so one clip can fit a different screen shape, while separate controls add an object to a scene or take one out, accounting for scale, interaction and shadow.
Character and motion controls
A performance can be driven by your own body, face and voice, and an object’s movement can be defined by drawing the path you want it to take.
1080p and 4K output
Google documents two output resolutions and says which is for what: 1080p for a sharper, cleaner file to edit with, 4K where texture and detail matter more than weight.
What Google publishes about Veo 3.1
Only facts Google states on its own pages appear here. Where the model page gives no figure — a maximum clip length, a per-second price — this table does not invent one.
- Developer
- Google DeepMind
- Current model
- Veo 3.1
- Category
- Video generation
- Output
- Video with natively generated audio
- Output resolutions
- 1080p and 4K
- Documented inputs
- Text prompts, images, reference images, video
- Documented controls
- Ingredients, style reference, character consistency, scene extension, camera controls, first and last frame, outpainting, object add and remove, character and motion controls
- Watermarking
- SynthID
- Consumer surfaces
- Gemini, Google Flow, Google Vids
- Developer surfaces
- Google AI Studio, Gemini API
- Benchmarks last updated
- October 2025
- Stated limitation
- Google says natural and consistent spoken audio, particularly for shorter speech segments, remains an area of active development.
Documented access routes
Veo 3.1 is a hosted model, so every route is an account on someone else’s service. Google names five surfaces on the model page, and each row below links to the one Google itself points at.
Gemini
The consumer assistant, where Veo appears as a named video surface.
OpenGoogle Flow
Google’s AI filmmaking tool, built for assembling clips into scenes and stories.
OpenGoogle Vids
Video creation inside Google’s work suite, aimed at workplace output rather than film.
OpenGoogle AI Studio
The browser path from prompt to API key, with a Veo studio app of its own.
OpenGemini API
The developer route, documented on ai.google.dev with worked prompt examples.
OpenVeo 3.1 questions people search for
What is Veo 3.1?
Google DeepMind’s current video generation model. It produces short clips with dialogue, sound effects and ambient audio generated in the same pass as the picture, from a text prompt, a single image or a set of reference images.
Does Veo generate audio as well as video?
Yes, and it is the model’s headline claim. Google’s own wording is “video, meet audio”, the showcase section is built around it, and several of the demo prompts write the spoken line into the prompt rather than leaving it to be dubbed later.
What resolution does Veo 3.1 output?
Google documents 1080p and 4K. Its page presents 1080p as the sharper, cleaner option for editing and 4K as the one for texture and detail.
Can Veo 3.1 edit a video I already have?
It can work on footage you supply. The documented controls include outpainting to extend the frame, adding an object, removing an object, extending a clip from its last second, and setting the first and last frame of a transition.
How long are Veo 3.1 clips?
The model page does not state a single maximum duration. It documents scene extension as the way to build longer pieces, and the footnotes to its head-to-head comparisons describe the evaluation clips as 6 or 8 seconds long.
Is Veo output watermarked?
Yes. Google says videos made with Veo are marked with SynthID, its watermarking and detection technology for AI-generated content, and that outputs also go through safety evaluations and memorisation checks before release.
Where can I use Veo 3.1?
Google names five surfaces: Gemini, Google Flow, Google Vids, Google AI Studio and the Gemini API. There is no self-hosted or open-weights route.
Latest Google Veo 3.1 articles
Browse every published Google Veo 3.1 article, from introductions to practical guides and developer documentation.
-
Veo 3.1 vs Sora 2: What Each Vendor Still Confirms
Google documents Veo 3.1 at eight seconds, 1080p and 4K with native audio. OpenAI says the Sora product is no longer available. Built only on vendor pages.
Read article -
Veo 3.1 vs Kling 3.0: What Each Vendor Publishes
Google documents Veo 3.1 at eight seconds, 1080p and 4K. Kuaishou documents Kling 3.0 at three to fifteen seconds and publishes a per-second credit rate.
Read article -
Veo 3.1 Prompt Guide: The Seven Elements Google Names
Google documents seven prompt elements for Veo and one line about Veo 3.1 itself. Here they are, with worked prompts and the parts Google leaves unpublished.
Read article -
Veo 3.1 Price and Free Access: The Official Numbers
Google publishes three AI plans, monthly Flow Credits and per-unit Vertex AI rates for Veo 3.1. Here is each number with its source, and what is missing.
Read article -
How to Use Veo 3.1
Where Veo 3.1 runs today, how a first clip comes together, one worked prompt with camera, light and audio direction, and the gaps Google leaves.
Read article -
Veo 3.1 API Pricing and Vertex AI Access
Veo 3.1 API pricing from the official pages, plus working curl and Python requests against the Gemini API and Vertex AI model ids.
Read article