Wan 2.5 marked the point where Alibaba’s Wan series became a serious pick for polished short-form video, and it remains a dependable middle option in the family: newer than the Turbo-era 2.2 models, simpler and cheaper to reason about than 2.6 or 2.7. Its motion and prompt adherence hold up well for everyday production work.
Two modes are available on Flux 3 AI. Text-to-video builds a clip from a written prompt and carries a generate-audio toggle; image-to-video animates an upload and accepts a reference image for subject consistency. Uniquely among the Wan generations on the platform, both 2.5 modes honor negative prompts, so you can exclude artifacts, text overlays, or unwanted elements up front.
Renders come out at 5 or 10 seconds, 720p or 1080p, in 16:9, 9:16, or 1:1 across both modes. If a composition keeps drifting, add the offending elements to the negative prompt and re-roll — it’s usually cheaper than upgrading generations to fix a framing problem.
See what's possible on Flux 3 AI
Sample generations from the platform — write your own prompt to see what Wan 2.5 does with it.
Wan 2.5 vs similar AI video models
Specs from the live catalog — every model below is available on the same account.
| Spec | Wan 2.5 | Wan 2.6 | Wan 2.2 Turbo |
|---|---|---|---|
| Max clip length | 10s | 15s | 8s |
| Durations | 5s, 10s | 5s, 10s, 15s | 4s, 6s, 8s |
| Resolutions | 720p · 1080p | 720p · 1080p | 480p · 720p |
| Aspect ratios | 3 | 3 | 2 |
| Audio | |||
| Image to video | |||
| Reference images | up to 1 | up to 1 | up to 1 |
| Keyframes | |||
| Negative prompt |
Capabilities reflect what each model supports on Flux 3 AI today.
What makes Wan 2.5 stand out
Negative prompts on both modes
The only Wan generation on Flux 3 AI where text and image modes both accept exclusions — fewer retries, cleaner takes.
Audio toggle on T2V
Text-to-video can generate a soundtrack with the clip when you enable the audio switch in the Create studio.
Three aspect ratios everywhere
Both modes frame 16:9, 9:16, and 1:1 — widescreen, vertical, and square from one model without cropping.
Reference image on I2V
Image-to-video takes a reference alongside your starting frame to hold a subject’s appearance steady.
Sensible credit profile
A step below the newest generations in cost, a step above Turbo in polish — the balanced default for routine clips.
Popular use cases
Why creators run Wan 2.5 on Flux 3 AI
- Negative-prompt retakes stay cheap when your credits live in one pool across every model.
- Put Wan 2.5 output next to Hailuo or Seedance takes in the same comparison view.
- The built-in prompt-enhance tool tightens your description before you spend a credit.
Wan 2.5 specs & capabilities
Output
- Clip duration5, 10 seconds
- Resolutions720p · 1080p
- Aspect ratios16:9 · 9:16 · 1:1
- Generation modesText to Video · Image to Video
Controls
- Image to video (start frame)
- Reference images
- First & last frame keyframes
- Audio generation
- Negative prompt
- Video input / editing
- Extend video
Available versions on Flux 3 AI
Every Wan 2.5 endpoint below is selectable in the Create studio.
- Wan 2.5 (T2V)Wan 2.5 — text-to-video
- Wan 2.5 (I2V)Wan 2.5 — image-to-video
Wan 2.5 prompt examples
Copy one as a starting point, or send it straight to the Create studio.
“A lighthouse keeper climbing a spiral staircase at the storm’s peak, swinging lantern throwing shadows across stone walls, slow upward tilt, thunder rumbling outside”
Pairs dramatic lighting with the text-to-video audio toggle
“Bring this product photo to life: the perfume bottle rotates slowly on a marble counter as morning light sweeps across it, petals drifting past”
Photo animation with a reference image holding the product steady
“A minimalist yoga sequence on an empty beach at sunrise, static wide shot, calm pastel palette — negative prompt: on-screen text, extra limbs, camera shake”
Shows negative prompts stripping artifacts before they cost a retake
Create with Wan 2.5 in three steps
Describe your idea
Write a prompt or upload a starting image — the more specific the scene, motion, and style, the better Wan 2.5 performs.
Pick Wan 2.5 & settings
Choose duration, resolution, and aspect ratio. The studio shows the exact credit cost before you generate.
Generate & download
Wan 2.5 renders in the cloud — track progress in your library, then download watermark-free or share with a link.
Wan 2.5 questions, answered
What is Wan 2.5?
Wan 2.5 is a mid-generation release of Alibaba’s Wan video model, offered on Flux 3 AI in text-to-video and image-to-video modes. It renders 5- or 10-second clips up to 1080p and stands out in the family for supporting negative prompts in both modes.
Does Wan 2.5 support negative prompts?
Yes — both text-to-video and image-to-video accept a negative prompt. List what you don’t want (blur, on-screen text, extra objects) and the model steers away from it, which typically cuts down the number of takes you need.
Can Wan 2.5 animate a photo?
Yes. Image-to-video uses your uploaded picture as the starting frame and animates it according to the prompt. You can also attach a reference image so the subject stays consistent, useful when a product or character must look identical across clips.
Does Wan 2.5 generate audio?
Text-to-video includes a generate-audio toggle, so you can choose between a silent render and one delivered with sound. Image-to-video renders video only; pair it with your own track in an editor if the clip needs audio.
Wan 2.5 vs Wan 2.6 — what’s the difference?
Wan 2.6 goes longer (15-second text-to-video) and adds video-to-video plus Flash speed tiers, while Wan 2.5 counters with negative-prompt support in both modes and 1:1 framing on text-to-video. For maximum output control per credit, 2.5 is often the smarter buy.
What output formats does Wan 2.5 produce?
Clips are 5 or 10 seconds long at 720p or 1080p, in 16:9, 9:16, or 1:1. Downloads from your Flux 3 AI library are full quality with no platform watermark.