Ad teams
A complete 30-second spot — footage, voiceover, and music — from one brief, iterated take after take.
Write a line, drop a frame, or hand over a stack of references — you get back up to 30 seconds of finished video with dialogue, music, and lip-synced speech already in place.
Production preview
Wan 3.0Clips from Alibaba's official launch demos — a long single take, character performance, and cinematic scale.
A cherub climbs a spiral staircase through the clouds in one continuous shot — the long, coherent take the model is built for.
Generate a long take
A stylized character raids the fridge at night — expression, motion, and lighting stay consistent through the whole scene.
Animate a character
Lantern-lit ruins, a giant stone Buddha, one warrior — a 30-second cinematic straight out of the model.
Stage a cinematic
Wan 3.0 is Alibaba's all-in-one video model: text to video, image to video, reference-driven generation, editing, and extension in a single model. Give it a sentence or a first frame and it renders up to 30 seconds at up to 1080p — and because audio is native, dialogue, sound effects, and background music arrive in the same pass, with speech that moves the character's lips. As of September 2026 it sits at #1 on Artificial Analysis' text-to-video (with audio) leaderboard at Elo 1,242, ahead of Gemini Omni Flash and MiniMax H3 Max.
Pick any whole-second length, with 5 as the default — or let the model choose the length that fits the action.
Start at 480p while you iterate, then render the keeper at 1080p — every clip runs at a steady 30 FPS.
Dialogue, ambient effects, and music are generated with the picture, and spoken lines drive the lips.
16:9, 4:3, 1:1, 3:4, or 9:16 — or leave the frame on adaptive.
Stack images, clips, audio — even a PDF, a deck, or a webpage — and the model builds the scene from all of them.
Swap elements, restyle a shot, rewrite a character's line, or add time before or after a clip, up to 30 seconds total.
Type what happens — or drop a first frame, an end frame, or a stack of references.
Choose 2–30 seconds, 480p to 1080p, and a frame from 16:9 to 9:16 — or leave the length on auto.
The clip comes back with voices, effects, and music in place. Keep it, edit a line, or extend the take.
It covers in one pass what used to take a stack of tools.
A complete 30-second spot — footage, voiceover, and music — from one brief, iterated take after take.
Image to video from one hero still: the product shot opens the scene and the model sets it in motion.
Turn a PDF, a slide deck, or a webpage into a narrated explainer without storyboarding it first.
Vertical 9:16 scenes with dialogue that lip-syncs — with the dialogue already audible — ready for the feed.
Cinematic sequences with consistent characters, staged shot by shot and extended scene by scene.
What changed from the Wan 2.7 API line, per Alibaba's launch announcement.
| This siteWan 3.0The all-in-one generationGenerate | Wan 2.7Previous API generation | |
|---|---|---|
| Longest single pass | Up to 30 seconds | Capped at 15 seconds |
| Model shape | Reference, editing, generation, and extension in one model | Split across separate models |
| Inputs | Text and frames, plus up to 20 reference assets — documents and webpages included | Narrower input set |
Subscriptions load your account with credits — $19.90 loads 1,500: 60 five-second drafts, or 15 ten-second 720p takes.
Enjoy Limited-Time 30% OFF!
2 to 30 seconds in a single pass, with 5 as the default — or hand the length over and the model picks what fits the action. Extending a clip keeps the total under 30 seconds.
Yes — dialogue, ambient effects, and background music render natively in the same pass, and spoken lines drive the character's lip movements.
Yes — image to video starts from one frame, or pin both the first and last frames and the model fills the motion between.
Yes — documents (PDF, Word, Excel, PowerPoint and more, up to 100MB or 50 pages) and webpage links go straight in as references, alongside up to 20 total assets.
No — output is 480p, 720p, or 1080p at a fixed 30 FPS. 1080p is the ceiling; the generator opens at 480p so drafts stay cheap.
No — it's a paid, hosted model. Plans start at $19.90 a month for 1,500 credits; a render costs by the second, so a five-second 480p draft is 25 credits and a ten-second 720p take is 100.
No — this generation is Alibaba's hosted API model. Alibaba hasn't released 3.0 weights — the downloadable open-source models stop at Wan 2.2 — so the 3.0 generation runs through a hosted generator like this one.
Keep the take, rewrite a line, extend the ending — your first scene is one prompt away.