Wan 2.2 / 3.0
Alibaba's open Wan model family, run two ways on AutorunX: self-hosted Wan 2.2 for cost-efficient generation, and cloud Wan 3.0 for up to 30-second clips.
What Wan does
On its owned track, Wan 2.2 (AX-WAN) generates short video clips from text or an image, with fine-grained control options — first/last-frame conditioning, pose or depth-guided motion, inpainting/outpainting within a frame, and audio-driven generation via speech-to-video. It's the model to reach for when cost efficiency matters more than chasing the newest generation's fidelity ceiling.
On its cloud track, Wan 3.0 offers the current generation as a hosted service: up to 30-second clips, mixed image/video/audio/document references, and first/last-frame interpolation — useful when a shot needs more reach than the owned 10-second tier.
Capabilities
Key features
- Two tracks, one model family
- Wan 2.2 runs owned and self-hosted on AutorunX's own GPU fleet for cost efficiency, while Wan 3.0 runs as a proprietary cloud API via EvoLink for longer clips and mixed-reference inputs — pick the track that matches your budget and quality needs.
- VACE structural controls
- The owned Wan 2.2 tier supports VACE: depth-map guidance, pose conditioning, and inpainting/outpainting, giving direct structural control over motion and composition beyond what a text prompt alone provides.
- Speech-to-video (S2V)
- Wan 2.2 can generate motion driven directly by an audio track, similar in spirit to audio-driven avatar models but built into the same open Wan architecture used for general T2V/I2V work.
- First-and-last-frame interpolation
- Both tracks support supplying a start and end frame (FLF2V on the owned tier, first/last frame on Wan 3.0 image-to-video) so the model fills in the motion between two fixed compositions.
- Apache 2.0 open weights
- Because Wan 2.2's weights are released under the permissive Apache 2.0 license, AutorunX can self-host generation on its own infrastructure instead of routing every request through a third-party paid API.
How Wan works
Wan is Alibaba's Apache 2.0-licensed video model family, and AutorunX runs it on two different tracks. The owned track is Wan 2.2 — a 14B mixture-of-experts (MoE) model alongside a smaller 5B variant — self-hosted on AutorunX's own GPU fleet under the internal name AX-WAN. Because the weights are openly licensed, AutorunX can run inference on its own infrastructure rather than paying a third-party API per call, which is what makes it the cost-efficient option in the lineup.
Wan 2.2's MoE architecture activates a subset of its parameters per inference step rather than the full network, which is part of how it keeps generation cost down relative to a dense model of similar total size. It supports a wide input surface: standard text-to-video and image-to-video, first-frame conditioning, first-and-last-frame (FLF2V) interpolation, VACE controls (depth maps, pose guidance, inpainting, and outpainting for targeted edits within a frame), and speech-to-video (S2V), where an audio track drives the generated motion. Clips run 5-10 seconds at 480p-1080p.
The cloud track is Wan 3.0, Alibaba's current hosted generation, reached through EvoLink rather than AutorunX's own GPUs. One upstream model is exposed as three routes (text-to-video, image-to-video with first/last frame, and mixed-reference video). Clips run 2-30 seconds at 480p, 720p, or 1080p, with optional native audio at no extra rate. Earlier cloud Wan 2.6/2.7 lanes stay routable for old jobs but are hidden from the picker.
What people use Wan for
Cost-efficient bulk generation
Use owned Wan 2.2 for high-volume or exploratory generation where per-clip cost matters more than chasing the absolute newest model generation's fidelity.
Structurally controlled shots
Use VACE's depth, pose, or inpaint/outpaint controls on the owned tier when a shot needs precise structural guidance rather than pure prompt-driven generation.
Audio-driven motion on the open track
Use Wan 2.2's speech-to-video mode when you want audio-driven generation within the owned, self-hosted pipeline rather than a separate cloud avatar model.
Higher-end cloud renders
Switch to Wan 3.0 when a shot needs up to 30 seconds, mixed references, or the current cloud generation.
How to use Wan on AutorunX
Wan is available in Short Film, in the Video module.
- 01
Open Short Film or Sync Motion in Video
From the AutorunX dashboard, go to Video and open Short Film or Sync Motion — Wan is available as the owned, cost-efficient model option in both.
- 02
Set up your input
Write a prompt, attach a reference or first/last frame, or supply an audio track if you're using speech-to-video.
- 03
Pick Wan in the model picker
Open the model picker and choose owned Wan 2.2 for the cost-efficient option, or Wan 3.0 / Wan 3.0 Identity / Wan 3.0 (T2V) for the cloud tier.
- 04
Generate and review
Hit generate to render a 5-10 second clip (owned) or a 2-30 second clip (Wan 3.0). Use VACE or first/last-frame controls to refine composition as needed.
- Billed per second of generated video from your shared AutorunX credit wallet.
- 3 credits / sec (owned Wan 2.2) · 75 credits / sec (cloud Wan 3.0 @ 720p)
Tips for better results
- Default to owned Wan 2.2 for cost-sensitive runs
- Since it's self-hosted and priced at 3 credits per second (an 8-second clip is just 24 credits), use the owned tier as your default for drafts, tests, and high-volume generation before reaching for the cloud tier.
- Use VACE for precise composition needs
- When a prompt alone isn't landing the right pose or depth relationship, supply a pose or depth guide through VACE instead of iterating on prompt wording alone.
- Reserve first-and-last-frame for locked compositions
- Use FLF2V/first_last_frame specifically when both the opening and closing frame of a shot matter, not for general-purpose generation.
- Move to the cloud tier only when needed
- Reserve Wan 3.0 for shots where the owned tier's 10-second cap or quality ceiling isn't enough — the cloud tier costs more per clip (75 credits/sec at 720p).
Wan — frequently asked
Related models
Seedance 2.0
ByteDance / BytePlusByteDance's default AutorunX video engine for Short Film, Movie Maker, and Ad Remake — identity-preserving reference-to-video at up to 4K.
Seedream
ByteDance / BytePlusByteDance's Seedream generates and edits images from up to ten reference photos, holding identity and composition steady across a full shoot.
Kling 3.0
Kuaishou TechnologyKuaishou's flagship video model with native multi-lingual audio, in-video editing, and clips up to native 4K.
Ready to create with Wan?
Open Short Film and generate your first video in a couple of minutes.
Free plan, no card required.