Skip to main content
We’ve moved.ClipDance.ai
14 days 07:04:43
Unlimited GPT Image 2 & Nano Banana 2 LiteGet Unlimited

Alibaba Wan 3.0 · public beta

Wan 3.0 AI Video GeneratorOne shot. Thirty seconds. Sound included.

Wan 3.0 is Alibaba's omni-modal video model. Feed it a prompt, a photo, a clip, a voice note, a slide deck or a URL, and it returns a single continuous take up to 30 seconds long with audio already mixed in.

Up to 30 secondsNative audioUp to 1080pText to videoImage to videoDocument to video

Made with Wan 3.0

Community Wan 3.0 clips, unedited. Hover to play, click for the prompt.

What it does

Five things Wan 3.0 does that shorter models cannot

Thirty seconds changes what you can tell in one shot. Wan 3.0 gives a scene room to breathe, a camera room to travel, a character room to finish a thought.

Generate a clip
01What it does

A Full 30 Seconds, No Cuts

Wan 3.0 renders up to 30 seconds as one continuous take — double the ceiling of the model before it. Long enough for a real scene instead of a loop.

No stitching, no crossfades hiding a seam — Wan 3.0 renders it in one pass.

02What it does

Feed It Almost Anything

Text, photos, video, audio, a PDF, a slide deck, a web page. Wan 3.0 reads all of them as source material and builds the video from what it finds.

Point it at a product page and get a promo cut.

03What it does

Shaped for the Feed

Vertical for TikTok and Reels, square for marketplaces, wide for YouTube. Pick the ratio before you render so nothing gets cropped later.

Six ratios, or let Wan 3.0 match your source footage.

04What it does

Anime to Photoreal

The same prompt box covers cel-shaded animation, gritty live action, and skin-level realism. Wan 3.0 follows the style you describe instead of forcing a house look.

Say the look you want and it holds.

05What it does

Faces That Stay the Same Face

Drift is what kills long AI clips — a character morphs, a logo warps, a room rearranges itself. Wan 3.0 was built to hold characters, products and spaces steady across the whole take.

Thirty seconds is only useful if second 29 still matches second 1.

The short version

What is Wan 3.0?

Wan 3.0 is the video generation model Alibaba released into public beta in August 2026. It produces single-shot clips of up to 30 seconds at 480p, 720p or 1080p, with a synchronized audio track and multilingual speech, from text or from mixed media you supply.

What separates Wan 3.0 from most video models is the range of things it accepts as input. Alongside prompts and images it reads video, audio, PDFs, presentations and live web pages, which makes it useful for turning material you already have into something watchable. On this site Wan 3.0 runs in the browser and bills in credits — a 5-second 480p clip costs 10 credits, and 1080p costs 38. Failed renders are refunded, and there is no API key or subscription to set up.

Picking between them

Wan 3.0 or Seedance 2.5?

Both reach 30 seconds. Wan 3.0 differs on what you can put in, and on what you pay.

Wan 3.0Seedance 2.5
Max length30 seconds30 seconds
ResolutionUp to 1080pUp to 720p
AudioNative, multilingual speechNative, optional
PDF / slides / web page inputYesNo
Reference imagesUp to 10Up to 30
CostFrom 10 credits per 5sFrom 30 credits per 5s

Rule of thumb: reach for Wan 3.0 when the source is a document, a web page or mixed media, or when 1080p matters. Reach for Seedance 2.5 when you need a large reference stack.

Where it earns its keep

What people build with Wan 3.0

The long single take and the document input make Wan 3.0 worth reaching for on jobs that were not worth starting before.

01What people build with Wan 3.0

Decks That Play Themselves

Hand Wan 3.0 a PDF or a slide deck and get a narrated video back. Useful for course modules, onboarding, and the internal update nobody was going to read.

02What people build with Wan 3.0

Short-Form Drama

Thirty unbroken seconds fits a whole beat — setup, turn, punchline. Wan 3.0 holds the actor's face and the set steady long enough for the scene to land.

Getting started

How to use Wan 3.0

Three steps. No install, no key, no card.

  1. 01

    Bring your source

    Type a prompt, or upload a photo, a clip, a voice note, a document, or paste a web page URL. Wan 3.0 accepts them mixed together.

  2. 02

    Set length and size

    Choose anywhere from 2 to 30 seconds, pick a resolution, and pick a ratio. Wan 3.0 updates the credit cost as you change them.

  3. 03

    Render and download

    Wan 3.0 returns a finished clip with its audio track. Keep it, tweak the prompt and rerun, or post it as is.

Open the generator

Getting better output

Four ways to get more out of Wan 3.0

Small changes to how you write the prompt move a Wan 3.0 result more than any setting does.

Tip

Write a Beginning and an End

Thirty seconds is long enough to need structure. Tell Wan 3.0 where the shot starts, what changes in the middle, and where it lands.

Tip

Say the Sound Out Loud

Audio is generated with the picture, so describe it. Put spoken lines in quotes, name the music mood, name the room tone.

Tip

Draft Cheap, Finish Sharp

Lock the composition at 480p for 10 credits, then rerun the prompt you like at 1080p. Paying 38 credits for a draft you throw away is the expensive way.

Tip

Trim the Document First

A 40-page PDF gives Wan 3.0 too much to choose from. Feed it the two pages that matter and the cut gets tighter.

Questions

Wan 3.0 FAQ

What people ask before their first render.

Wan 3.0 is Alibaba's video generation model, released in public beta in August 2026. It creates single-shot clips of up to 30 seconds with synchronized audio, from a text prompt or from mixed inputs including images, video, audio, PDFs, slide decks and web pages. You can run Wan 3.0 here in the browser without an API key or a subscription.

Anywhere from 2 to 30 seconds, as a single continuous take. That is double the 15-second ceiling of the previous generation, and it is the reason a Wan 3.0 clip can carry a full scene rather than a loop.

Yes, and it is the thing that sets it apart. Pass a document URL or a web page URL and Wan 3.0 reads the content as source material for the video. It works with PDFs, slide decks and live pages, which makes it a fast route from a deck you already wrote to something you can post.

It does, in the same pass as the picture. Wan 3.0 produces dialogue, ambience, effects and music together with the video, including multilingual speech synchronized to the character's mouth. There is no silent draft to score afterwards.

It bills per second of output by resolution. A 5-second clip costs 10 credits at 480p, 19 at 720p, and 38 at 1080p; a full 30-second render at 1080p costs 227. Reference images, video and audio inputs are not charged separately, and failed renders are refunded automatically.

Both reach 30 seconds. Choose Wan 3.0 when your source is a document, a web page or mixed media, or when you want 1080p output. Choose Seedance 2.5 when you need to feed a large stack of reference images. Both run in the same studio, so switching costs you one dropdown.

Your turn

Make something 30 seconds long

A prompt, a photo, or the deck sitting in your downloads folder. Wan 3.0 turns it into a clip with sound in a couple of minutes.