Core Concepts
Text-to-Video in plain English.
Also known as: T2V,AI video generation,prompt to video
The one-sentence version
Generating a video clip from a written description, using a model trained on video and its captions.
Text-to-video models take a prompt such as "a golden retriever running through autumn leaves, slow motion" and produce a short clip, typically five to twenty seconds, at up to 1080p or 4K. They are the moving-picture cousins of image generators, built on diffusion or diffusion-transformer architectures that learn how frames change over time. Sora 2, Google Veo, Kling, Runway, and Hailuo lead the closed models; Alibaba's Wan and Tencent's HunyuanVideo are the main open-weight options. Quality has improved sharply since 2024: physics, hands, and text are more coherent, and native audio generation arrived in 2025. The remaining limits are length, consistency across shots, and control. Most tools therefore add image-to-video (start from a still), reference images for characters, and camera-control settings, and creators stitch short clips in an editor. Prices run from free tiers with watermarks to per-second credits, and licensing questions about training data remain unresolved for several models.