A still image can solve the hardest part of a shot: the first frame. Image-to-video generation then asks the model to infer what moves, what stays stable, and how the camera sees the change. Clear inputs and a narrow motion brief make that inference easier to evaluate.
1. Prepare the source frame
Choose an image with a clear subject, intentional composition, and enough resolution for the final crop. Remove accidental borders, duplicate objects, or background details that you do not want interpreted as part of the action. If the subject must remain recognizable, keep its face, silhouette, and key edges visible.
Decide whether the source is a first frame, a visual reference, or both. The distinction matters because each model may use inputs differently. MotionVeo's current video catalog lists source-image and reference-image support for Veo 3.1 Fast, Seedance 2.0, and Wan 2.6, with up to seven references listed for each.
2. Pick one dominant motion
Ask what should visibly change: hair moves in a breeze, a product rotates, a person turns toward the light, or a train crosses the background. One dominant motion gives you a clean test. If the first result is stable, build the next beat as a separate clip.
Useful motion sentence
“The camera stays locked while the paper model slowly unfolds from the center; the front edge lifts first, then the side panels open by the final second.”
This tells the model what the camera does, what the subject does, and how the action progresses.
3. Separate camera and subject language
Write camera movement and subject movement as two different clauses. “Slow push in as the dancer turns left” is easier to inspect than “dynamic cinematic movement.” If you want no camera movement, state “locked camera” or “static framing.” If the subject must stay centered, say where it remains in the frame.
4. Choose 16:9 or 9:16 early
Use 16:9 when the environment and lateral motion matter. Use 9:16 when the subject is tall or the destination is mobile-first. Check that the source image already has enough room for the selected crop; changing the ratio after the fact can cut off the action.
5. Iterate in short loops
- Render a simple motion with the original source frame.
- Watch the full clip for identity drift, edge warping, and unwanted camera movement.
- Change one instruction: motion speed, camera path, or framing.
- Keep the source stable while you test the revised sentence.
Do not judge an image-to-video result from its thumbnail alone. The first frame can look perfect while the transition introduces artifacts in hands, text, thin objects, or repeating patterns.
6. Review the artifact checklist
| Area | Question |
|---|---|
| Subject | Does the main identity and silhouette stay coherent? |
| Motion | Is the requested action visible and paced across the clip? |
| Camera | Does the camera follow the instruction, or drift unexpectedly? |
| Edges | Do hands, text, product edges, hair, and thin lines remain usable? |
| Composition | Does the selected ratio keep the subject inside the intended crop? |
When to use references
Use a reference set when several visual facts matter together: a character wardrobe, product color, environment, or recurring design language. Start with the smallest set that communicates those facts. More references add constraints but do not guarantee a better result, so test deliberately.
Next step
Read the AI video prompt guide for a reusable shot formula, or open the AI video generator hub to compare the current model pages and settings.