Quick answer
Text-to-video models generate a short clip, usually 5 to 20 seconds, from a written prompt, using diffusion-style models trained on video and captions. Sora 2, Veo, Kling, and Runway lead the closed models; Wan and HunyuanVideo are the main open-weight ones. Because a prompt gives limited control, most professional work uses image-to-video: design the frame first, then ask the model to animate it. Consistency across shots, length, and control remain the open problems.
The demos are astonishing and the reality is a craft. Understanding how these models work explains both why the demos look the way they do and why your first ten attempts will not.
How generation works
Video models extend image diffusion into time. They learn from millions of clips paired with descriptions how pixels change frame to frame, and generate by starting from noise and refining toward something that matches the prompt and stays coherent across frames. Newer architectures (diffusion transformers) handle longer clips and better physics; several models now generate synchronised audio too. Each clip costs seconds to minutes of expensive GPU time, which is why the closed services charge per second.
Why image-to-video dominates real work
- Control: a prompt cannot specify a composition; an image can
- Consistency: reusing a character or product image across clips keeps it recognisable
- Workflow: the still can be made in Midjourney, Photoshop, or a camera, then animated
- Editing: first-and-last-frame modes let you define exactly where a shot begins and ends
What still goes wrong
Hands, text, and complex physics; subjects drifting away from the reference after a few seconds; unwanted camera moves; morphing objects. Clips are short, so anything longer is several clips stitched in an editor. And the training data of most models is undisclosed, which matters for commercial use; Moonvalley's Marey is the notable model trained only on licensed footage.
Related reading
Bottom line
Text-to-video is the headline; image-to-video is the tool. Start from a still, keep clips short, and expect to generate several takes per usable shot.



