Alibaba has pushed Wan 3.0 into public beta, doubling maximum video output duration to 30 seconds over Wan 2.5 and introducing automatic runtime optimization alongside a clip extension tool. For enterprise workflows, the core upgrade is not raw duration but ingestion: the model ingests native PDF documents, raw web pages, and PPTX slide decks alongside up to ten images, five video snippets, and five audio tracks within a single unified prompt.
By parsing structured documents and rich media simultaneously, Wan 3.0 targets the operational bottleneck in corporate content pipelines—collapsing the manual gap between raw product specifications, training slide decks, and finalized video assets without requiring third-party storyboarding tools. Alibaba claims the architecture mitigates character drift and UI distortion by locking reference geometry and key identities across scenes.
Commercial access is live across the wan.video portal, Alibaba Cloud Model Studio, and Qwen Cloud APIs in Standard and low-latency Prime tiers. Positioning the tool across training, marketing, and robotics simulation represents an operational test for Alibaba: after logging a 75 percent drop in quarterly operating income driven by surging infrastructure spending, the group urgently needs B2B generative video pipelines that generate tangible enterprise software revenue.