Back to blog

Wan 3.0

PA

PoloX AI

Run Alibaba Wan 3.0 on PoloX. An AI video generator for text to video, image to video, and reference to video — up to 30 seconds of 1080p with native audio in one pass.

Capabilities

Up to 30 seconds in one pass

Wan 3.0 generates anywhere from 2 to 30 seconds in a single take, not short clips stitched together afterwards. One continuous pass is what lets camera movement and one-take shot language hold together across a whole scene.

1080p with native audio

Output at 480p, 720p, or 1080p, with 1080p the default. Audio is generated in the same pass as the picture, so dialogue, ambience, and on-screen action land together instead of being dubbed on afterwards.

Full-body motion at speed

Fast, athletic movement is the hardest thing to keep coherent. Wan 3.0 is positioned for limbs, weight, and ground contact that stay readable through a whole take — dance, crowd, and mixed lighting included.

Reference-conditioned shots

Reference-to-video conditions on what you already have: up to 10 images, 5 video clips, and 5 audio tracks in one request. Address them in the prompt as Image1, Video1, or Audio1 to say which file is the character and which is the location.

FAQ

What is Alibaba Wan 3.0?

Wan 3.0 is the next generation of the Wan video model family from Alibaba’s Tongyi Lab, the team behind Wan 2.1 through Wan 2.7. It generates video from text, from a still image, or from reference material, at up to 30 seconds a shot with audio in the same pass.

Which Wan 3.0 tasks are available on PoloX?

Three: text-to-video for generating a shot from a prompt, image-to-video for animating a still, with an optional end frame so you can set where the shot finishes, and reference-to-video for conditioning on material you already have. All three share the same duration, resolution, aspect ratio, and audio controls. Jobs use PoloX credits.

What resolutions, durations, and aspect ratios does Wan 3.0 support?

On PoloX, output runs at 480P, 720P, or 1080P, with 1080P the default. Durations run from 2 to 30 seconds in a single generation. Aspect ratios cover 16:9, 4:3, 1:1, 3:4, and 9:16, or adaptive so the model can pick the ratio that suits the shot. Audio is generated alongside the video by default.

What can I use as a reference with Wan 3.0?

Reference-to-video takes up to 10 reference images (JPEG, PNG, WEBP, or BMP, 20MB each), up to 5 reference video clips (MP4 or MOV, 1–15 seconds each, 15 seconds total), and up to 5 reference audio tracks (MP3 or WAV). You need at least one image or video. Your prompt can address the files as Image1, Video1, or Audio1. PoloX does not currently expose document or web-page inputs.

Try this model on PoloX →