Limited-Time 30% OFF!Get Offer
#1 on Artificial Analysis' text-to-video (with audio) leaderboard — Sep 2026

Wan 3.0: direct a whole scene, voices included

Write a line, drop a frame, or hand over a stack of references — you get back up to 30 seconds of finished video with dialogue, music, and lip-synced speech already in place.

Start frame
End frame
Duration (seconds)5s
2s16s30s

Production preview

Wan 3.0
Official showcase

Inside the official Wan 3.0 demos

Clips from Alibaba's official launch demos — a long single take, character performance, and cinematic scale.

One long, unbroken take

A cherub climbs a spiral staircase through the clouds in one continuous shot — the long, coherent take the model is built for.

Generate a long take
A cherub climbs a spiral marble staircase through the clouds — a long Wan 3.0 single take

Character performance that holds

A stylized character raids the fridge at night — expression, motion, and lighting stay consistent through the whole scene.

Animate a character
Wan 3.0 AI video demo: stylized character opening a fridge at night

Game-cinematic scale

Lantern-lit ruins, a giant stone Buddha, one warrior — a 30-second cinematic straight out of the model.

Stage a cinematic
Wan 3.0 video generation demo: warrior before a giant stone Buddha in lantern-lit ruins

What is Wan 3.0?

Wan 3.0 is Alibaba's all-in-one video model: text to video, image to video, reference-driven generation, editing, and extension in a single model. Give it a sentence or a first frame and it renders up to 30 seconds at up to 1080p — and because audio is native, dialogue, sound effects, and background music arrive in the same pass, with speech that moves the character's lips. As of September 2026 it sits at #1 on Artificial Analysis' text-to-video (with audio) leaderboard at Elo 1,242, ahead of Gemini Omni Flash and MiniMax H3 Max.

Controls

Wan 3.0 gives you every dial the scene needs

  • 2 to 30 seconds, one pass

    Pick any whole-second length, with 5 as the default — or let the model choose the length that fits the action.

  • Up to 1080p at 30 FPS

    Start at 480p while you iterate, then render the keeper at 1080p — every clip runs at a steady 30 FPS.

  • Sound is native

    Dialogue, ambient effects, and music are generated with the picture, and spoken lines drive the lips.

  • Five frame shapes, or adaptive

    16:9, 4:3, 1:1, 3:4, or 9:16 — or leave the frame on adaptive.

  • Up to 20 references

    Stack images, clips, audio — even a PDF, a deck, or a webpage — and the model builds the scene from all of them.

  • Edit and extend

    Swap elements, restyle a shot, rewrite a character's line, or add time before or after a clip, up to 30 seconds total.

Workflow

How to make an AI video with Wan 3.0

  1. 1

    Stage the scene

    Type what happens — or drop a first frame, an end frame, or a stack of references.

  2. 2

    Set the frame

    Choose 2–30 seconds, 480p to 1080p, and a frame from 16:9 to 9:16 — or leave the length on auto.

  3. 3

    Generate with sound on

    The clip comes back with voices, effects, and music in place. Keep it, edit a line, or extend the take.

Use cases

What teams make with Wan 3.0

It covers in one pass what used to take a stack of tools.

  • Ad teams

    A complete 30-second spot — footage, voiceover, and music — from one brief, iterated take after take.

  • Product marketers

    Image to video from one hero still: the product shot opens the scene and the model sets it in motion.

  • Educators & explainers

    Turn a PDF, a slide deck, or a webpage into a narrated explainer without storyboarding it first.

  • Social creators

    Vertical 9:16 scenes with dialogue that lip-syncs — with the dialogue already audible — ready for the feed.

  • Story & game teams

    Cinematic sequences with consistent characters, staged shot by shot and extended scene by scene.

Generation leap

Wan 3.0 vs the previous Wan generation

What changed from the Wan 2.7 API line, per Alibaba's launch announcement.

This siteWan 3.0The all-in-one generationGenerateWan 2.7Previous API generation
Longest single passUp to 30 secondsCapped at 15 seconds
Model shapeReference, editing, generation, and extension in one modelSplit across separate models
InputsText and frames, plus up to 20 reference assets — documents and webpages includedNarrower input set
Pricing

Wan 3.0 pricing, by the credit

Subscriptions load your account with credits — $19.90 loads 1,500: 60 five-second drafts, or 15 ten-second 720p takes.

Limited-time offer

Enjoy Limited-Time 30% OFF!

07DAY00HOUR00MIN00SEC
Loading…
FAQ

Wan 3.0 questions, answered

One pass in. A whole scene out, speaking.

Keep the take, rewrite a line, extend the ending — your first scene is one prompt away.