Techniques & Methods

Image-to-Video in plain English.

Also known as: I2V,animate an image,still to video

The one-sentence version

Generating a video that starts from, or is guided by, an existing still image, giving far more control than a text prompt alone.

Image-to-video takes a still image and animates it: the model treats the image as the first frame (or a reference) and generates motion that follows a text prompt. It is how most professional AI video is actually made, because it solves the biggest problem with text-to-video, which is control. You can design the frame in Midjourney, Photoshop, or a camera, then ask the model only to move it. Variants include first-and-last-frame interpolation, where you supply both ends and the model fills the middle, and reference-guided generation, where an image sets a character's appearance across many clips. Every major video model supports it: Runway, Kling, Veo, Sora 2, and the open-weight Wan and HunyuanVideo. Typical failure modes are the subject drifting from the source image after a few seconds, unwanted camera motion, and objects morphing. Tools address these with motion strength settings, camera presets, and shorter clips chained together.

Read the full guide