Image to Video Prompts That Keep Your Reference Image: One Camera Move, Bounded Motion
Make this article actionable
Send the article context into Vife Agent and turn it into a plan, checklist, or draft you can keep working on.
Quick answer
A reliable image to video prompt does three jobs and stops there: it locks the reference image, it names exactly one camera move, and it bounds subject motion with a start state and an end state. Everything else — style adjectives, mood words, extra shots — competes with the image you already supplied and gives the model room to redraw your subject.
A working template:
Preserve the reference image: [subject], [framing], [lighting], [background].
Camera: one [move] — [direction] — [speed].
Motion: [subject] goes from [start state] to [end state] over the clip.
Keep [locked elements] unchanged.If your clip drifts, the fix is almost always subtraction, not more description.
Turn the useful parts into next steps
Vife Agent can convert this guide into a prioritized workflow with tasks, risks, and reusable prompts.
Why reference-image drift happens
Image-to-video models are conditioned on your still, but the conditioning is a strong hint, not a hard constraint. The model still generates every frame, and it resolves ambiguity by inventing. Drift usually comes from one of four sources:
- Competing descriptions. If your prompt describes the subject in different words than the image shows, the model averages the two. "A woman in a red coat" against an image of a woman in a rust-colored jacket produces a jacket that is neither.
- Multiple camera moves. "Slow push in while orbiting and tilting up" is three instructions. Models tend to blend them into a wobble, and the blended motion drags geometry with it.
- Unbounded motion. "She walks through the market" gives the model a whole scene to rebuild. "She turns her head from left to right" gives it one joint to move.
- Duration mismatch. A 10-second clip from a single still has far more frames to fill than a 4-second clip. Longer clips give drift more time to accumulate.
The practical consequence: write prompts that describe change, not appearance. The image already carries appearance.
The workflow: lock, move, bound, review
Step 1 — Write the lock line from the image, not from memory
Open the image and list what you actually see: subject, framing, lighting direction, background elements, and any text or logo. Put those in the first line as things to preserve. Use the same nouns the image suggests rather than synonyms.
Preserve: ceramic mug on a wooden desk, three-quarter view, soft window light
from the left, blurred bookshelf behind, no text.Step 2 — Choose one camera move and name its direction and speed
Pick a single move from a short list and commit to it. Direction and speed matter more than the move's name, because "push in" without a direction or rate is ambiguous.
| Move | What it does to the frame | Good for |
|---|---|---|
Slow push in | Camera moves toward the subject; background compresses | Product reveals, portraits, building tension |
Slow pull out | Camera moves away; more environment enters frame | Establishing context after a detail shot |
Lateral truck (left/right) | Camera slides sideways; parallax separates layers | Showing depth in a scene with foreground objects |
Tilt up or down | Camera rotates vertically in place | Revealing height, architecture, tall subjects |
Static with subject motion | Camera holds; only the subject moves | Dialogue-style shots, subtle product motion |
Avoid combining moves in one clip. If you need a push and then a pan, generate two clips and cut them together — that is a cheaper fix than fighting a blended move.
Step 3 — Bound the subject motion with a start and end state
Describe the subject's motion as a transition between two states you can see in your head. Keep it to one primary action.
- Bounded: "The steam rises and thins; the mug itself does not move."
- Bounded: "She blinks once and turns her head slightly to the right, ending in a three-quarter profile."
- Unbounded: "She reacts to something off-screen and the scene comes alive."
The second pair shows the difference. Bounded motion names a body part, a direction, and an end position. Unbounded motion names an emotion and hopes the model picks a matching action.
Step 4 — Add a short keep-unchanged list
This is where you protect the details that matter most: faces, hands, product labels, logos, text, and background architecture. Keep it to three or four items. A long list dilutes attention.
Keep unchanged: face and hair, mug shape and handle, desk grain, window light direction.Step 5 — Generate, then review against the checklist below
Run the clip, then score it on the review criteria rather than on a gut feeling. Regenerate with one change at a time so you learn which line caused the problem.
Copyable example prompt
This is a written example, not a measured result — treat it as a starting structure and adjust the nouns to match your own reference image.
Preserve the reference image: a matte black wireless earbud case on a pale
concrete surface, centered, shot from slightly above, soft diffused light
from the upper right, shallow depth of field, no text or logos visible.
Camera: one slow push in, moving toward the case, gentle and steady.
Motion: the case lid opens from fully closed to about 45 degrees over the
clip; the case body stays planted on the surface.
Keep unchanged: case shape and finish, concrete texture, light direction,
shadow position under the case.
Duration: short clip. No cuts, no additional camera moves.What makes this prompt work: one move, one bounded action with a numeric end state, a short lock list, and an explicit "no cuts" instruction that removes a whole class of failure.
Output review checklist
Score each item pass or fail before you decide to regenerate. Subjective items are marked; the rest are things you can verify by looking.
- Subject identity — face, silhouette, and proportions match the reference. (subjective, but check side by side)
- Product fidelity — shape, color, and finish match; no invented buttons, seams, or text.
- Camera move — exactly one move is visible, in the direction you asked for, at roughly the speed you asked for.
- Motion bounds — the subject ends where you specified; nothing walks out of frame or transforms.
- Background stability — architecture, horizon lines, and repeated patterns do not warp or crawl.
- Hands and edges — fingers, hair edges, and thin objects do not smear or duplicate.
- Text and logos — any text in the reference is either preserved or absent; it is not garbled into new letters.
- Loopability — if you need a loop, the first and last frames are close enough to cut together.
If three or more items fail, the prompt is usually over-specified. Cut the lock list to three items and remove any motion that is not the primary action.
Common mistakes
Describing the image instead of the change. Restating appearance wastes prompt budget and invites the model to reinterpret. Describe only what moves.
Stacking camera moves. Two moves in one clip usually produce neither. Split into two clips.
Using negative parameters as if they were universal. Some models accept flags like --no for exclusions; others ignore them or treat them as text. Provider-specific parameters are not portable across models, so check the parameter documentation for the specific model you select rather than copying flags between them.
Asking for a long clip from one still. Drift compounds. Generate short, then extend or cut.
Fixing drift by adding adjectives. "Ultra-detailed, photorealistic, 8K" does not restore a face. Removing a competing description does.
Assuming one prompt transfers across models. Models differ in how strongly they weight the reference image and in which motion vocabulary they respond to. A prompt tuned for one model may need its camera line rewritten for another. Availability and credit costs also differ per model — check the live model selector for what is currently offered and the displayed quote for credits before you build a batch around one option.
Where this fits in a larger workflow
Reference-locked clips are usually one step in a sequence: a still you generated or shot, a short motion clip, then assembly with other shots. If you are producing several clips from a set of stills, keep a prompt log with the lock line, camera line, and motion line for each, so a successful result is reproducible and a failed one is diagnosable.
You can run image-to-video generation alongside image editing, audio, and document work in the same workspace — see Vife's image-to-video capability for how the still-to-clip step fits with the rest of a project. If you plan to generate many clips, review current credit costs on the pricing page before committing to a long shot list.
FAQ
How many camera moves should one image to video prompt include?
One. A single named move with a direction and a speed gives the model an unambiguous instruction. If your shot needs a push and then a pan, generate two clips and cut them.
Why does my subject's face change even though I supplied a reference image?
Usually because the prompt describes the subject in words that differ from the image, or because the motion is unbounded enough that the model rebuilds the head to animate it. Shorten the lock line, remove appearance adjectives, and reduce motion to a single action such as a blink or a small head turn.
Can I use --no to exclude unwanted elements?
Only if the model you selected documents that flag. Parameters are model-specific and are not portable; a flag that works in one model may be ignored or read as literal text in another. Check the parameter reference for your chosen model.
Should I write the prompt before or after choosing the model?
After. Motion vocabulary and reference-image weighting vary by model, so write the lock line and motion bound first, then adjust the camera line to match the model's documented phrasing. Confirm the model is currently available in the live model selector rather than assuming from an older tutorial.
How do I stop the background from warping?
Add the background's stable features to the keep-unchanged list — horizon lines, wall edges, repeated patterns — and avoid lateral moves in scenes with strong perspective unless you want parallax. Static camera plus subject motion is the most stable combination.
Is a longer prompt better for image-to-video?
No. Longer prompts add competing instructions. The useful prompt is short: a lock line, one camera move, one bounded motion, and a short keep list. Add detail only when a specific review item fails, and add one line at a time so you know what fixed it.

