motionveo
Guide

AI video generator guide: prompts, images, and formats.

An AI video generator turns a text direction, source image, or reference set into a short moving clip. Better results come from describing what changes over time and checking one output variable at a time.

An AI video generator is most useful when it behaves like a controllable production tool. Give it a visible subject, one primary action, a camera intention, and a destination format. Then review the motion and make a small, deliberate change.

In this guide
  1. Text-to-video or image-to-video?
  2. A seven-part prompt formula
  3. Choose the frame before the render
  4. An eight-second iteration loop
  5. A review checklist

Text-to-video or image-to-video?

Text-to-video is a good fit when the scene is still an idea. Describe the subject and the action, then let the model explore composition. Image-to-video is a better fit when the first frame already contains an important product, character, layout, or visual identity. A clean source frame reduces the number of things the model must invent at once.

MotionVeo's current video catalog lists source-image and reference-image support for Veo 3.1 Fast, Seedance 2.0, and Wan 2.6. Check Studio before rendering because the catalog is live.

A seven-part prompt formula

Build the first prompt from these parts:

  1. Subject: who or what is visible?
  2. Action: what changes during the clip?
  3. Environment: where does the action happen?
  4. Camera: is the camera static, tracking, panning, or pushing in?
  5. Timing: what happens at the start, middle, and end?
  6. Light and look: describe light direction, contrast, color, or texture.
  7. Output constraints: state 16:9 or 9:16 and any important framing boundary.

Example

Weak: “A beautiful cinematic city video.”

Stronger: “A red delivery bicycle crosses a rain-darkened downtown intersection at dusk. The camera tracks left at walking speed, reflections move across the pavement, headlights bloom softly, and the rider exits frame right by the final second. Vertical 9:16 composition with the bicycle kept fully in frame.”

Choose the frame before the render

Aspect ratio is a creative decision, not a finishing step. A 16:9 frame leaves room for a wider environment and is easier to use in presentations or landscape edits. A 9:16 frame protects a tall subject and is usually the safer starting point for mobile-first work. Keep the main subject away from the extreme edges until you know how the selected model handles motion.

Common output choices
DestinationStarting ratioPrompt reminder
Landscape edit16:9Describe the horizontal camera path and leave breathing room at both sides.
Vertical social cut9:16Keep the subject legible in a tall crop and state where it exits the frame.
Storyboard studyChoose the final delivery shapeDo not optimize a frame you will later crop away.

An eight-second iteration loop

Short clips reward a tight loop. First, render the simplest version that can answer one question: does the subject move the right way? Second, inspect camera motion, identity, hands, edges, and pacing. Third, change only the weakest variable. If the camera is correct but the subject drifts, preserve the camera sentence and revise the subject sentence. If the framing is wrong, change the ratio or camera instruction before adding more style words.

Do not treat a single generation as a benchmark. Different prompts can expose different strengths and failure modes, and the current catalog fields do not establish a universal quality winner.

A review checklist

  • Does the first frame contain the intended subject and composition?
  • Is the primary action visible from beginning to end?
  • Does the camera move as described, or introduce unwanted drift?
  • Are faces, hands, text, and product edges stable enough for the use?
  • Does the chosen ratio preserve the subject for the destination?
  • Can you describe the next change in one sentence?

Keep learning

Compare the enabled video models in the AI video generator hub, then read the prompt-writing guide or the image-to-video guide for a narrower workflow.