Video GenerationAutorunX Owned Lane

Wan 2.2 / 2.6 / 2.7

Alibaba's open Wan model family, run two ways on AutorunX: self-hosted Wan 2.2 for cost-efficient generation, and cloud Wan 2.6/2.7 for higher-end output.

Selecting Wan in the AutorunX model picker

Picking Wan as the model for a generation.

DeveloperAlibaba
CategoryVideo generation (T2V, I2V, VACE, S2V)
Owned tierWan 2.2 (14B MoE/5B), self-hosted as "AX-WAN", Apache 2.0
Cloud tierWan 2.6/2.7, proprietary, via EvoLink
Clip length5-10s (owned) · 10s (cloud)
Resolution480p-1080p (owned) · 720p-1080p (cloud)
On AutorunXCost-efficient owned option for Short Film and Sync Motion

What Wan does

On its owned track, Wan 2.2 (AX-WAN) generates short video clips from text or an image, with fine-grained control options — first/last-frame conditioning, pose or depth-guided motion, inpainting/outpainting within a frame, and audio-driven generation via speech-to-video. It's the model to reach for when cost efficiency matters more than chasing the newest generation's fidelity ceiling.

On its cloud track, Wan 2.6 and 2.7 offer the newer generation of the same architecture family as a hosted service, trading self-hosted cost savings for access to Alibaba's latest closed-weight improvements — useful when a shot needs more polish than the owned tier delivers and the higher per-clip cost is acceptable.

Key features

Two tracks, one model family

Wan 2.2 runs owned and self-hosted on AutorunX's own GPU fleet for cost efficiency, while Wan 2.6/2.7 run as a proprietary cloud API via EvoLink for a newer, higher-end tier — pick the track that matches your budget and quality needs.

VACE structural controls

The owned Wan 2.2 tier supports VACE: depth-map guidance, pose conditioning, and inpainting/outpainting, giving direct structural control over motion and composition beyond what a text prompt alone provides.

Speech-to-video (S2V)

Wan 2.2 can generate motion driven directly by an audio track, similar in spirit to audio-driven avatar models but built into the same open Wan architecture used for general T2V/I2V work.

First-and-last-frame interpolation

Both tracks support supplying a start and end frame (FLF2V on the owned tier, first_last_frame on Wan 2.7) so the model fills in the motion between two fixed compositions.

Apache 2.0 open weights

Because Wan 2.2's weights are released under the permissive Apache 2.0 license, AutorunX can self-host generation on its own infrastructure instead of routing every request through a third-party paid API.

How Wan works

Wan is Alibaba's Apache 2.0-licensed video model family, and AutorunX runs it on two different tracks. The owned track is Wan 2.2 — a 14B mixture-of-experts (MoE) model alongside a smaller 5B variant — self-hosted on AutorunX's own Vast/RunPod GPU fleet under the internal name AX-WAN. Because the weights are openly licensed, AutorunX can run inference on its own infrastructure rather than paying a third-party API per call, which is what makes it the cost-efficient option in the lineup.

Wan 2.2's MoE architecture activates a subset of its parameters per inference step rather than the full network, which is part of how it keeps generation cost down relative to a dense model of similar total size. It supports a wide input surface: standard text-to-video and image-to-video, first-frame conditioning, first-and-last-frame (FLF2V) interpolation, VACE controls (depth maps, pose guidance, inpainting, and outpainting for targeted edits within a frame), and speech-to-video (S2V), where an audio track drives the generated motion. Clips run 5-10 seconds at 480p-1080p.

The cloud track — Wan 2.6 and 2.7 — is a separate, proprietary hosted tier reached through EvoLink rather than AutorunX's own hardware. Wan 2.6 supports image-to-video and text-to-video at 10-second clip lengths; Wan 2.7 adds a first_last_frame mode that takes two images (a start and end frame) to interpolate between. Both run at 720p-1080p. These are the same model family as the owned tier but a newer, closed-weight generation accessed as a paid API rather than self-hosted inference.

What people use Wan for

Cost-efficient bulk generation

Use owned Wan 2.2 for high-volume or exploratory generation where per-clip cost matters more than chasing the absolute newest model generation's fidelity.

Structurally controlled shots

Use VACE's depth, pose, or inpaint/outpaint controls on the owned tier when a shot needs precise structural guidance rather than pure prompt-driven generation.

Audio-driven motion on the open track

Use Wan 2.2's speech-to-video mode when you want audio-driven generation within the owned, self-hosted pipeline rather than a separate cloud avatar model.

Higher-end cloud renders

Switch to Wan 2.6/2.7 when a shot needs the newer generation's improvements and the higher cloud per-clip cost is justified.

Who built Wan

Alibaba is the Chinese technology and e-commerce conglomerate whose cloud and AI division develops the Wan (Wanxiang) video generation model family. Alibaba has released successive Wan generations under the permissive Apache 2.0 license, making the model weights freely usable, including commercially, which is why AutorunX can self-host it.

How to use Wan on AutorunX

Wan is available in Short Film in Video Lab.

1

Open Short Film or Sync Motion in Video Lab

From the AutorunX dashboard, go to Video Lab and open Short Film or Sync Motion — Wan is available as the owned, cost-efficient model option in both.

2

Set up your input

Write a prompt, attach a reference or first/last frame, or supply an audio track if you're using speech-to-video.

3

Pick Wan in the model picker

Open the model picker and choose owned Wan 2.2 for the cost-efficient option, or a cloud Wan 2.6/2.7 variant for the higher-end hosted tier.

4

Generate and review

Hit generate to render a 5-10 second clip (owned) or 10-second clip (cloud). Use VACE or first/last-frame controls to refine composition as needed.

An example generation rendered with Wan on AutorunX

An example generation rendered with Wan.

Credit usage

Billed per clip from your shared AutorunX credit wallet.

2 credits / clip (owned Wan 2.2) · 45-53 credits / clip (cloud Wan 2.6/2.7)

Tips for better results with Wan

Default to owned Wan 2.2 for cost-sensitive runs

Since it's self-hosted and priced at 2 credits per clip, use the owned tier as your default for drafts, tests, and high-volume generation before reaching for the cloud tier.

Use VACE for precise composition needs

When a prompt alone isn't landing the right pose or depth relationship, supply a pose or depth guide through VACE instead of iterating on prompt wording alone.

Reserve first-and-last-frame for locked compositions

Use FLF2V/first_last_frame specifically when both the opening and closing frame of a shot matter, not for general-purpose generation.

Move to the cloud tier only when needed

Reserve Wan 2.6/2.7 for shots where the owned tier's quality ceiling isn't enough, since the cloud tier costs meaningfully more per clip.

Wan — frequently asked questions

Ready to create with Wan?

Sign up for AutorunX to get 200 free credits across every lab, including Short Film.