How to Keep a Subject's Position Stable Across Shots in a Short AI Video
Make this article actionable
Send the article context into Vife Agent and turn it into a plan, checklist, or draft you can keep working on.
Quick answer
Subject drift in a short AI video is usually a description problem, not a model problem. When each shot is written as a fresh scene, the generator has no reason to place your subject where the previous shot left it. The fix is to lock one anchor point in the frame, repeat that anchor in the wording of every shot, and change exactly one other variable between shots.
For a 15-second vertical clip, that means:
- One subject, described the same way every time.
- One anchor position, stated in every shot (for example, "lower-third center").
- One change per shot transition — lighting, background, or distance — never two at once.
- Two shots before three. If two shots hold, add the third.
The rest of this article walks through a concrete three-shot example, a copyable prompt, a review checklist, and the failure patterns that cause the jump.
Turn the useful parts into next steps
Vife Agent can convert this guide into a prioritized workflow with tasks, risks, and reusable prompts.
Why the subject moves when the scene changes
A short AI video is not a continuous camera move. It is a set of separate generations that a viewer's eye stitches together. Each generation is a new composition. If your shot 1 says "a mug on a desk" and your shot 2 says "a close-up of a mug," you have described two different pictures that happen to share an object. Nothing in the wording tells the second generation to preserve the mug's screen position or its relative size.
Three things make this worse:
- Vague position language. "On the desk" is a region, not a point. A mug can sit anywhere in that region and still satisfy the description.
- Changing two variables at once. If shot 2 moves closer and changes the light, you cannot tell which change caused the jump — and the generator has more freedom to recompose.
- Scale drift disguised as framing. "Wide" and "closer" are relative words. Without a stated anchor, "closer" often means "centered and larger," which reads as a jump even when the subject is technically correct.
The practical consequence: you are not asking the tool to track an object across time. You are asking it to produce compositions that look like the same setup. That is a writing task you can control.
The anchor-point method, step by step
Step 1: Fix one anchor point and name it in every shot
Pick a single point in the frame and commit to it. For a vertical 9:16 clip, useful anchors include:
lower-third centercenter of frame, slightly below the midlineleft third, at the height of the desk edge
Write the anchor into every shot description, using the same words. Do not paraphrase. "Lower-third center" in shot 1 and "near the bottom middle" in shot 2 are two different instructions.
Step 2: Keep distance and angle wording identical, and change only one variable
Decide what your three shots actually are. In the mug example:
- Shot 1: wide
- Shot 2: closer
- Shot 3: back to wide
That is already a distance change between shots 1 and 2, and again between 2 and 3. So the only other thing that should move is the background light — cool in shot 1, warm in shot 2, warm in shot 3, or cool-to-warm across the sequence. Do not also change the desk, the camera height, or the surface.
If you want the light to be the single visible change, keep the camera wording frozen and let the light do the work.
Step 3: Write each shot as one sentence
One sentence per shot, with three parts: subject + position + one change. If a shot needs a second sentence, it probably contains a second change.
Step 4: Review by scrubbing the shots side by side
Put the three clips in a row and scrub through them at the same time. Watch the anchor point, not the subject. Your eye will forgive a mug that looks slightly different; it will not forgive a mug that teleports.
Step 5: If it drifts, cut to two shots
Two shots that hold read better than three that jump. Remove the middle shot, or remove the shot with the largest framing change, and re-check. Add the third shot back only after the pair is stable.
Example input: one mug, three shots
Here is a copyable brief. It is an example of the method, not a measured result — treat it as a starting point and adjust the wording to your own subject.
Subject: one plain unbranded ceramic mug, matte off-white, no logo, no text.
Setting: a plain wooden desk, no other objects.
Format: vertical 9:16, 15 seconds total.
Anchor point (repeat in every shot): the mug sits at lower-third center,
its base resting on the desk edge, occupying roughly one quarter of the
frame width.
Shot 1 (wide): A plain unbranded ceramic mug at lower-third center on a
plain wooden desk, seen in a wide shot from desk height, cool daylight
from the left.
Shot 2 (closer): A plain unbranded ceramic mug at lower-third center on a
plain wooden desk, seen in a closer shot from the same desk height and
angle, warm light from the left.
Shot 3 (wide): A plain unbranded ceramic mug at lower-third center on a
plain wooden desk, seen in a wide shot from desk height, warm light from
the left.
Constraints: no text overlays, no screenshots, no invented interface
elements, no hands, no second object. Only the light temperature changes
between shots.Two things to notice. First, the anchor sentence is repeated verbatim. Second, the camera wording — "from desk height," "same angle" — is repeated so the only moving part is the light.
If you are generating these shots from a still image rather than from text alone, the same discipline applies to the source frame: the anchor point has to be visible and unambiguous in the image before you describe it. The image-to-video workflow is where that first frame gets set, and it is worth locking the composition there before you write the shot list.
Output review checklist
Run this against the finished clip, not against the prompt.
| Check | What to look for | Pass condition |
|---|---|---|
Anchor holds | Watch the mug's screen position across all three shots | It does not move more than a small margin — a few percent of frame width at most |
Relative size holds | Compare the mug's width to the frame width in shot 1 and shot 3 | Wide shots match; the closer shot is the only size change |
Subject readable | Can you identify the mug in every shot? | Yes, in all three, without squinting |
One change only | List what differs between shots | Exactly one thing — here, light temperature |
No text overlays | Scan the frame edges and center | No captions, labels, or watermarks |
No screenshots or invented UI | Look for panels, buttons, cursors, or fake app chrome | None present |
No extra objects | Check the desk surface | Only the mug |
The first two rows are subjective judgments made by eye. They are not measurements, and different viewers will draw the line in slightly different places. The last four rows are closer to binary — you can see whether a caption is there or not.
Common mistakes
Rewriting the anchor in new words. "Lower-third center" becomes "bottom middle" becomes "near the front edge." Each paraphrase is a new instruction. Copy and paste the anchor sentence instead.
Changing light and distance in the same transition. This is the most common cause of a jump that looks like a tracking failure but is really a composition change. Freeze one, move the other.
Using "close-up" without a scale reference. "Closer" is relative. If you want the mug to stay the same relative size, say so, or keep the wide framing and change only the light.
Adding a third shot before the pair is stable. Three shots triple the number of transitions you have to inspect. Get two right first.
Describing the subject differently each time. "Ceramic mug" in shot 1 and "coffee cup" in shot 2 invites a different object. Keep the noun phrase identical.
Letting the background carry detail. A busy desk gives the generator more to recompose. A plain surface makes the anchor easier to hold and easier to check.
Assuming a prompt guarantees consistency. It does not. A tight brief reduces the number of ways a shot can be composed; it does not control the model. Expect to regenerate, and expect to cut a shot.
FAQ
Why does my subject jump position even when I use the same prompt?
Because the prompt describes a scene, not a continuity constraint. If the wording allows the subject to sit anywhere in a region, each generation may place it differently. Naming a specific anchor point narrows the range of acceptable compositions.
Should I change the camera angle between shots?
Only if the angle change is the one variable you are testing. If your single change is lighting, keep the angle wording identical. If you want an angle change, freeze the light and the distance instead.
How much drift is acceptable?
There is no universal number. A practical rule for a 15-second vertical clip: if you notice the movement while watching at normal speed rather than while scrubbing, it is too much. Scrubbing side by side is the stricter test.
Can I fix drift in editing instead?
Sometimes — a slight reframe or a short cross-dissolve can hide a small shift. It will not fix a scale change or a subject that has moved to a different part of the frame. Fixing the shot list is faster than fixing the timeline.
Does this work for subjects other than objects?
The method is about screen position, so it applies to a person, a product, or a plant equally well. The harder the subject is to describe in one noun phrase, the more carefully you need to repeat that phrase.
Do I need a paid plan to try this?
That depends on the plan and the current feature set, which change over time. Check the pricing page for what is currently included rather than relying on a general assumption.
What to do next
Write your three shots as three sentences. Copy the anchor phrase into all three. Pick one variable to move. Generate, scrub side by side, and cut to two shots if the anchor wanders. The clip will read as one continuous setup even though it was built from separate generations — and you will have a shot list you can reuse for the next product.

