How to Describe Camera Movement in an AI Video Prompt (Without Adjective Soup)

9 min read

Make this article actionable

Send the article context into Vife Agent and turn it into a plan, checklist, or draft you can keep working on.

Open in Agent

Quick answer

To describe camera movement in an AI video prompt, write it as a physical instruction, not a mood. Name four things in order: (1) the subject and where it sits in frame, (2) the movement type — push in, pull out, lateral truck, small orbit, or locked-off, (3) the speed and distance of that move, and (4) the start and end composition. Then add a short "keep constant" clause listing what must not change: subject appearance, background, light direction.

If a camera operator could not perform your sentence in one take, the prompt is not specific enough yet.

Mid-read shortcut

Turn the useful parts into next steps

Vife Agent can convert this guide into a prioritized workflow with tasks, risks, and reusable prompts.

Create a brief

Why adjectives fail and instructions work

Words like cinematic, dynamic, epic, or sweeping describe how a shot feels. They do not tell a system where the camera is, where it goes, or how fast. When several vague adjectives compete, the result is usually a shot that drifts, zooms unpredictably, or changes the subject's face between frames.

Concrete language fixes this because it constrains the problem. "Slow push in, about one meter over the full clip" is a bounded instruction. "Dramatic camera work" is an invitation to improvise.

A useful mental model: you are writing a shot card for a camera operator who has never seen your storyboard. Everything they need is position, motion, timing, and what must stay put.

The five-part prompt structure

1. Subject and frame

Start with who or what is on screen and where they sit. Use frame-relative positions rather than vague ones.

  • Weak: "a woman in a nice kitchen"
  • Strong: "a woman in a linen apron, centered in the left third of frame, waist-up, kitchen counter visible behind her"

Also state the aspect ratio and framing if it matters to the edit — vertical 9:16, medium shot behaves differently from wide 16:9, full body.

2. Movement type

Pick one primary move. Common, well-understood options:

MovementWhat it doesTypical use
Push in
Camera moves toward subject
Build attention, reveal detail
Pull out
Camera moves away
Reveal context, end a beat
Lateral truck / pan
Camera slides or rotates sideways
Follow action, show a space
Small orbit
Camera arcs partway around subject
Add dimension to a static subject
Locked-off
No camera movement
Product detail, dialogue, text-safe shots

One move per shot is the safest default. Two moves in one short clip often read as a glitch rather than a flourish.

3. Speed and magnitude

Speed needs a unit of comparison. "Slow" alone is subjective; "slow" relative to what?

  • Tie speed to the clip: slow push in, roughly one meter of travel across the full clip
  • Or tie it to the subject: camera moves at the same pace as the subject's walking speed
  • Or use a plain descriptor plus a bound: gentle, no faster than a slow dolly

Magnitude matters as much as speed. A push in of 20 cm and a push in of 3 meters are different shots.

4. Start and end composition

Describe the first frame and the last frame. This is the single highest-value habit in prompt writing, because it turns a movement into a destination.

  • Start: subject small in frame, wide shot, doorway visible on the right
  • End: subject fills the right half of frame, doorway out of view

If you cannot describe the end frame, you probably do not know what the shot is for yet.

5. What must stay constant

Close with an explicit consistency clause. This is where most first drafts lose their subject's face, wardrobe, or lighting.

Keep constant: subject's face and hairstyle, apron color, background wall, window light coming from camera left.

Note that parameter syntax is not portable between models. Some systems accept negative or exclusion flags, others do not, and a flag that works in one tool may be ignored or misread in another. Write consistency requirements as plain sentences so they survive a model change, and check the model selector for what the current tool actually supports.

Fictional example: vague vs. specific

The following is an illustrative example written for this article, not a tested or measured result. It uses a made-up scenario so you can see the rewrite pattern.

Vague version

text
Cinematic dramatic shot of a barista making coffee, dynamic camera, moody lighting, epic feel, 4k

Problems: no subject position, no movement type, no speed, no end frame, no consistency clause. "Dynamic camera" and "cinematic" pull in different directions.

Specific version

text
Vertical 9:16, medium shot. A barista in a dark green apron stands centered in frame behind an espresso machine, steam rising from the left. Camera performs a slow push in, moving about 40 cm toward the subject across the full clip, no faster than a slow dolly. Start frame: waist-up, machine and counter visible. End frame: chest-up, machine cropped out of frame. Keep constant: barista's face and hairstyle, apron color, counter surface, warm window light from camera left.

The second version is longer, but every clause removes a decision the system would otherwise make for you. It also survives being handed to a human camera operator, which is a good test of clarity.

Step-by-step workflow

  1. Write the shot's job in one line. "Show the product label clearly" or "Reveal that the room is empty." If you cannot state the job, the movement is decoration.
  2. Choose one movement type from the table above.
  3. Set the start frame in frame-relative terms.
  4. Set the end frame. Check that start and end are reachable in one continuous move.
  5. Add speed and magnitude with a unit of comparison.
  6. Add the consistency clause.
  7. Read it aloud as a camera instruction. If it sounds like a mood board, rewrite it.
  8. Generate, then review against the checklist below before changing anything else.

Output review checklist

Run these three questions first, because they catch most failures:

  • Can a single camera perform this? A push in plus a full orbit plus a rack focus in a three-second clip is not one shot. Split it into two shots or drop a move.
  • Does the movement fit the duration? A slow push that needs eight seconds will feel rushed in a three-second clip. Either shorten the travel distance or lengthen the clip.
  • Does the movement conflict with consistency requirements? A large orbit around a subject whose face must stay identical is a harder ask than a small arc. A pull out that reveals a new background area introduces new content that must also stay consistent.

Then check the output itself:

  • Subject identity, wardrobe, and hair match across the clip
  • Background elements do not appear, vanish, or rearrange
  • Light direction stays the same
  • The move starts and ends where you specified
  • No unintended second movement (drift, wobble, unrequested zoom)

Treat the first four as subjective judgments — you are comparing the clip to your own shot card. Treat the last one as closer to a measured check: you can scrub frame by frame and confirm whether the end composition matches your stated end frame.

Common mistakes

  • Stacking movements. Two or three moves in one short clip. Pick one.
  • Speed without a unit. "Slow" means nothing without a reference distance or a comparison.
  • No end frame. The system picks one, and it is rarely yours.
  • Consistency clause omitted. Then the face, apron, or lighting drifts.
  • Adjective pile-up. Cinematic, epic, dynamic, sweeping in one line. Each one dilutes the others.
  • Copying parameter syntax across models. Flags are not portable. Describe intent in words.
  • Describing the edit instead of the shot. "Cut to a close-up" is an editing instruction, not a camera move.

FAQ

How many camera moves should one prompt contain? One primary move is the safe default. If the shot needs two, consider splitting it into two shots.

Do I need to specify speed numerically? A number helps, but a comparison works too — "same pace as the subject's walk" or "no faster than a slow dolly." What matters is that speed is bounded.

What if I want no camera movement at all? Say so explicitly: locked-off camera, no movement. Silence is not the same as a static shot.

Where do I put the consistency clause? At the end, as its own sentence. Keeping it separate makes it easy to reuse across every shot in a sequence.

Can I reuse one movement description across a whole project? Yes, and you should. A small library of three or four movement phrases — push in, pull out, lateral truck, locked-off — keeps a sequence visually coherent and makes prompts faster to write.

How do I know if my prompt is specific enough? Hand it to a camera operator. If they would have to ask a question, add the answer to the prompt.

Where this fits in a production workflow

Prompt writing is one step in a longer chain: brief, shot card, prompt, generate, review, assemble. Once your movement language is consistent, you can move through that chain faster and spend your review time on composition instead of on re-explaining the shot.

If you are producing product or marketing footage, the AI product video generator is where these shot cards become clips — you can write the movement description, generate, and compare the result against your stated start and end frames. Availability of specific models and the credit cost of a generation are shown in the live model selector and the quote displayed at generation time, so check there rather than relying on a fixed list. If you are planning volume, pricing explains how credits are consumed.

The habit worth keeping is the shot card itself. It is portable, it works with a human crew, and it makes every review faster because you already wrote down what "correct" looks like.