China’s technology giant Alibaba has launched a public beta of Wan3.0, an AI video generation model supporting clips of up to 30 seconds and multimodal inputs including text, image, video, audio, web pages, PDFs, and PowerPoint presentations.

In a statement on Monday, Alibaba said users can apply for testing through Alibaba Cloud’s Model Studio and Qwen Cloud platforms.

The 30-second clip length is double the 15-second maximum of Wan2.7-Video, the preceding model, and extends beyond the typical few seconds to 15 seconds produced by mainstream AI video generators. The longer format allows for complex camera movements and continuous unbroken shots. Wan3.0 also introduces an intelligent duration feature that recommends optimal video length based on user prompts, alongside video extension capabilities to expand narrative timelines.

To address visual drifting and distortion common in AI-generated video, Wan3.0 can maintain high-precision visual continuity, rendering realistic human faces with synchronized micro-expressions, producing multilingual voice outputs, and maintaining stable software user interfaces and motion graphics. The model accurately replicates characters, props, audio, spatial layouts, and styles from reference inputs while keeping layouts and audio stable.

Intended use cases include filmmaking, short dramas, social media content, marketing and educational videos for businesses, and simulation video generation for training self-driving car and robotics systems.

Alibaba’s Wan series of visual generation models was first introduced in July 2023 and has since undergone multiple upgrades to improve realism, control, and usability.

China’s Alibaba launches Qwen3.8-Max AI model with 2.4T parameters, 1M token context window