Quick answer

Five video models changed the field between late July and early October 2026. Seedance 2.5 (ByteDance) and MiniMax H3 both launched on July 31; Alibaba's Wan3.0 followed in August; LTX-2.5 shipped open weights in August; and Kling 4.0 Flash opened to Ultra Yearly subscribers on September 28 ahead of a full October launch. The pattern: 30-second single generations (Seedance, Wan3.0, Kling 4.0), native audio everywhere, 4K on three of them, keyframe control as the headline feature, and a split between closed API models and open weights. Google's Veo 3.1 remains the Western developer baseline with native audio; OpenAI's Sora API shut down during 2026.

A year ago the question was whether a model could make a coherent eight-second clip. This autumn the questions are how long, how controllable, and whether you can run it yourself. Here is each model, then how to choose.

Kling 4.0 (Kuaishou)

Up to 30 seconds in a single generation, 4K with stereo sound, and steering with as many as 10 keyframe images, plus multimodal input combining images, videos, and subjects. 10-bit HDR and two-minute extension are announced as coming. The Flash tier is in early access for Ultra Yearly subscribers; the full model and API are expected in October. Kling's strengths in physics and character consistency carry over.

Wan3.0 (Alibaba)

30-second clips at up to 1080p in one pass, double Wan2.7's ceiling, with the unusual ability to generate from documents, spreadsheets, slides, and web pages. Priced per output second on Alibaba Cloud Model Studio: $0.05 at 480p, $0.10 at 720p, $0.20 at 1080p, so a 30-second 1080p clip is $6. Unlike the open-weight Wan 2.x family, Wan3.0 is closed and API-only.

Seedance 2.5 (ByteDance)

30 seconds of synchronised audio and video at native 4K, twice Seedance 2.0's length. Up to 30 images, 10 videos, and 10 audio clips as input, and timestamp-based direction and revision of individual sections, which moves it from generator toward editing system. Rolling out through Jimeng AI and Doubao Pro, and hosted by partner platforms.

MiniMax H3

A general multimodal video model: text, image, video, and audio in; up to 15 seconds at 2K with native stereo out. Its distinct features are video-to-video motion transfer, multimodal editing, and multi-shot generation, and it is distributed through MiniMax's apps and partners including Luma Agents. Shorter than the others, stronger at transforming footage you already have.

LTX-2.5 (Lightricks)

The open-weight entry: downloadable weights with synchronised audio, native multishot, precise editing, and 4K HDR output, positioned to hold up from first draft to final render. It runs in LTX Studio and through hosted providers at per-second prices comparable to H3. Running the weights yourself needs serious GPUs.

How to choose

  • Longest controllable clip: Kling 4.0 (once the full model lands) or Seedance 2.5
  • Predictable per-second pricing via API today: Wan3.0
  • Transforming existing footage: MiniMax H3, or Runway Aleph in the Western stack
  • Open weights and HDR: LTX-2.5; open weights on consumer GPUs: the older Wan 2.x
  • Western vendor, native audio, included with a subscription: Google Veo
  • All Chinese-hosted models: read the content policy and data terms before commercial use

Bottom line

Thirty seconds with sound is the new baseline, and keyframes are how professionals now direct a shot. Pick Kling 4.0 or Seedance 2.5 for control, Wan3.0 for pricing clarity, H3 for editing footage, LTX-2.5 for open weights. And generate the same storyboard in two of them before you commit a project to either.