Adobe 为 Firefly 视频只公布了一套提示词公式,一行就能写完:镜头类型描述 + 角色 + 动作 + 地点 + 美学。

官方提示词页面上剩下的内容,都是在展开这五个槽位。这篇保留官方顺序,补上提示词框旁边那套官方相机词典,并标出文档没写到的位置。

官方公布的五槽位结构

公式来自 Adobe 的提示词指南,同样的结构也出现在 2024 年那篇 beta 发布博客里,所以它不是一次性的建议。

槽位Adobe 提的问题Adobe 自己给的例子
镜头类型描述相机的视角是什么?怎么运动?a close-up shot with a slow zoom-in.
角色描述角色是谁?长什么样?穿什么?情绪如何?a large polar bear with bright white fur looking pensive.
动作角色在场景里做什么?the polar bear is walking softly toward a hole it opened in the ice.
地点角色在哪里?天气如何?地形如何?the location is barren and snowy; gray clouds move slowly in the distance.
美学什么类型的镜头?氛围是什么?景深如何?cinematic, 35mm film, highly detailed, shallow depth of field, bokeh.

Adobe 对密度的要求写得很直白:尽可能多用词,把灯光、摄影、调色、情绪和美学风格说具体。

每个槽位填什么

镜头类型描述

先说景别,再说运动。官方例子把尺寸和运镜配在一起,这正是这个槽位的意义。

角色描述

是谁、长什么样、穿什么、什么情绪。官方那只北极熊用十几个词回答了四个问题,这就是这个槽位要的密度。

动作

一个具体动词加上节奏。官方例子用了 walking softly but confidently,还补了一个目的地,模型因此有地方可去,而不只是端着一种情绪。

地点

要天气和地形,不要一个地名。官方例子先给 barren and snowy,再让天空动起来。

美学

技术层面的收尾:介质、细节程度、景深。官方例子是一串逗号分隔的词,官方博客里重复用的也是这个写法。

示例:两段官方提示词逐槽位拆开

下面两段提示词都来自 Adobe,发布在官方提示词页上,分节标题就是镜头名。它们不是拿来填空的模板。

示例一:特写肖像

Cinematic closeup and detailed portrait of a golden retriever dog in a field of sunflowers at golden hour. The lighting is cinematic, gorgeous, and soft, with beautiful, strong backlight and lens flare. The color grade is warm, sun-drenched, and sunlit. The dog is extremely realistic with detailed fur texture. The movement is subtle and soft. The camera doesn't move. There is heavy film grain and textures.
Shot size: Close up shot
Camera angle: None
Motion: Static

Adobe 把它放在 Close-up 分节下,并解释了用意:特写常用来突出某个细节或表情。这段提示词把景别写进了句子里,而不是交给下拉框,所以换成别的分辨率时,这个构图仍然站得住。

示例二:夜晚街道的横移

A lonely street with cobbled streets and Victorian street lamps late at night bathes the scene in the warm yellow glow of the street lamps. A beautiful, vivid full moon adorns the city, and the camera pans slowly from right to left, revealing an abandoned Mansion in the distance.
Shot size: None
Camera angle: Eye level shot
Motion: Move left

Adobe 把它放在 Move 分节下,官方给这一节的用意是让主体产生运动感或焦点感。提示词把横移写进了句子,Motion 下拉框提供对应的取值,模型不必自己去猜。

提示词旁边的相机控件

相机这套说法在 Adobe 不止一个地方出现。网页端是三个下拉框,同样的三个概念在 Firefly Video API 里又变成字段。

Adobe 官方 Firefly Services 音视频页面,展示了同一项生成能力背后的开发者接口
Adobe
分组取值默认
Shot sizeExtreme close up、Close up shot、Medium shot、Long shot、Extreme long shotNone
Camera angleAerial shot、Eye level shot、High angle shot、Low angle shot、Top down shotNone
MotionZoom in、Zoom out、Move left、Move right、Tilt up、Tilt down、Static、HandheldNone

这张表里的 None 不是占位符。Adobe 写明:默认情况下相机没有规定的运动,除非你选了某个 Motion 选项,或者在提示词里描述了相机运动。下拉框是覆盖项,不是必填项。

控件被系统收走的时候

有两个动作会静默关掉相机设置,在怪罪提示词之前值得先知道。

上传一段 Motion reference 视频,Shot size、Camera angle 和 Motion 会被禁用,因为运镜改由那段参考视频提供。加上首帧或尾帧,被牵连的清单更长:Composition 下的 Reference、Motion reference、Shot size、Camera angle 和 Style 都会自动禁用。Style 反过来也一样,选了风格预设就不能再设首尾帧。

绝大多数「为什么这里变灰了」的答案都在这里。而上面这些情况里,提示词的槽位始终可用,这也是把景别写进句子的价值。

长度与主体上限

Adobe 给了两个数字和一个提醒。提示词没有长短的硬限制,但 Firefly 有 1800 词的上限,而且长提示词不一定效果更好。主体方面,超过四个常常会把 Firefly 弄糊涂。

两个数字都不是目标。1800 是天花板,真正会改变阅读方式的是主体数:一行里五个有名字的角色,和一串美学词汇里的五个名词,是两种不同的风险。

迭代,而不是重写

Adobe 的建议是从基础版开始,每跑一次加一层细节。官方页面上用同一个主体、四段提示词演示了这件事。

Full scene, eye level, wide shot, a giant mech in the street, high quality, high details.
Full scene, eye level, wide shot, a giant mech, yellow and orange armor, in desolate streets, wires, LEDs, cybernetic parts, high quality, high details.
Full scene, eye level, wide shot, a giant black and yellow mech, accents of orange neon LEDs, marching, firing lasers, scanning, in the street of a destroyed city, rubble, fires, decayed buildings, desolate, ominous, high quality, great details.
Cinematic action scene, a group of giant mechs is invading the city, they are menacing, giant black and yellow mechs, yellow and orange matte armor, a dystopian future, in the street of a destroyed city, rubble, fires, decayed buildings, desolate, ominous, high quality, high details, volumetric lighting.

最后一段把 eye level 和 wide shot 换成了 cinematic action scene。这就是槽位在起作用:镜头描述从字面指令变成类型,剩下的词自己补上了差额。一次生成跑出满意的结果之后,Seed 选项能在种子、提示词和控件设置都不变的情况下复现相近片段,分步教程就是按这个方式留住成品的。每秒费率负责这些重复生成的账,模型对照解释了换掉 Model 下拉框之后同一段提示词为什么表现不同。

常见问题

换模型之后结构要改吗? 这套公式是 Adobe 针对自研模型给的指引。Adobe 另外说明,可用的视频生成设置会随所选视频模型而不同。

把镜头写进提示词了,还需要下拉框吗? 不需要。Adobe 说默认没有规定的相机运动,除非你选 Motion 选项或在提示词里描述运动。

提示词最长能写多少? Firefly 的上限是 1800 词,Adobe 也提醒长提示词不一定效果更好。

一段提示词能写几个主体? 少于五个。Adobe 提醒超过四个主体常会让 Firefly 困惑。