Skip to main content
Vantaige
WAN (Wan-Video / Wan2GP) screenshot
WAN (Wan-Video / Wan2GP) logo

WAN (Wan-Video / Wan2GP)

Free

WAN is Alibaba's open-weight video generation series (2.1 through 2.7), offering text-to-video, image-to-video, and editing capabilities under Apache 2.0. It topped the VBench leaderboard at launch in February 2025 and became the default choice for self-hosted AI video work.

Features:APIOpen Source

WAN is a family of open-weight video generation models built by Alibaba's Tongyi Lab, part of DAMO Academy. Released publicly on February 25, 2025 as Wan 2.1, the series immediately topped the VBench leaderboard for open-source video generation, outscoring every prior open model and matching several commercial APIs at zero licensing cost. The weights are released under Apache 2.0, meaning anyone can download, run, fine-tune, and redistribute them for personal or commercial use without paying Alibaba anything. By April 2026, the series had advanced through versions 2.2, 2.5, 2.6, and 2.7, with each major release adding capabilities: cinematic camera control, native audio, and a reasoning-first "Thinking Mode" in the latest 2.7 API.

The core capabilities span text-to-video (T2V), image-to-video (I2V), first-last-frame-to-video (FLF2V), video creation and editing (VACE), speech-to-video (S2V), and character animation. The flagship 14B-parameter models generate 480p and 720p clips on consumer GPUs like an RTX 4090, while the compact 1.3B model runs on just 8.19 GB of VRAM, making it one of the only serious video models that works on a mid-range gaming card. A thriving ecosystem of LoRAs, ComfyUI workflows, and third-party wrappers (notably ComfyUI integration and the Wan2GP wrapper for low-VRAM setups) has grown up around the series since launch.

What WAN outputs in April 2026

Wan 2.1, the version most creators still run locally, generates clips up to 5 seconds at 480p or 720p, at 16 fps. The 14B T2V model handles cinematic motion, detailed environments, and accurate physics. The 1.3B T2V trades visual detail for accessibility. The I2V-14B variant animates a still image into a video, and the FLF2V-14B model interpolates between a start and end frame. VACE (Video All-in-one Creation and Editing), released May 2025, added inpainting, outpainting, and instruction-guided editing to the same weights.

Wan 2.2 (July 2025) introduced the first open-source Mixture-of-Experts architecture for video diffusion: 27 billion total parameters split across two expert networks (a high-noise denoising expert and a low-noise refinement expert), with only 14 billion active per inference step. Training data scaled up 65.6% in images and 83.2% in videos versus 2.1. New model variants added speech-to-video (S2V-14B) and full character animation (Animate-14B), which can animate a single photograph into a consistent video sequence.

Wan 2.5 (September 2025) added native audio: synchronized dialogue, sound effects, and ambient background noise generated alongside the video frames in a single pass. The model handles 1080p output and 10-second clips. Wan 2.7 (April 6, 2026) brought "Thinking Mode," where the model explicitly plans scene composition before generation, yielding better narrative coherence and sharper character identity. Wan 2.7, however, ships as closed weights available only through a paid API at $0.10 per second of 720p video. This is a notable departure from the open-weight ethos of every prior version.

For users who want to run WAN on limited hardware, the community-built Wan2GP wrapper by deepbeepmeep brings VRAM requirements down to as low as 6 GB on some configurations, adds a web UI, supports multiple models (WAN 2.1/2.2, HunyuanVideo, LTX Video, FLUX), and includes queue-based batch processing. It is the practical entry point for anyone without an RTX 4090.

Where WAN sits versus HunyuanVideo and Mochi 1

HunyuanVideo (Tencent, open-released December 2024) is built on a Causal 3D VAE with a dual-stream transformer and 13 billion parameters. Its architectural strength is multi-character scene coherence: scenes with 3-5 distinct characters maintain individual identities, correct spatial positioning, and coordinated movement better than WAN in head-to-head tests. The trade-off is severe: HunyuanVideo requires a minimum of 45-60 GB of GPU memory for 720p generation, putting it firmly in the data-center or multi-GPU workstation tier. WAN's 1.3B model, by contrast, runs on a single RTX 3060 with 8 GB VRAM. For creators on consumer hardware, HunyuanVideo is not a realistic alternative; for studios running A100 clusters, its character fidelity can justify the compute cost. You can compare both options at HunyuanVideo's listing.

Mochi 1 (Genmo, October 2024) is a 10 billion parameter model built on an Asymmetric Diffusion Transformer (AsymmDiT), with a visual stream that has nearly four times as many parameters as the text stream. It outputs at 480p at 30 fps, giving noticeably smoother motion than WAN 2.1's 16 fps. The catch is the hardware bar: single-GPU deployment requires approximately 60 GB of VRAM. Mochi 1 was the open-source benchmark before WAN 2.1 shipped four months later and took over the top spot on VBench. WAN 2.2's MoE architecture has since extended the gap in overall quality, and WAN's I2V, VACE, and editing pipelines give it a much broader feature surface. Mochi 1 remains a solid earlier entry worth checking at Mochi 1's listing, particularly for the motion smoothness if you have the hardware.

Against commercial services: Runway Gen-3 and Pika offer polished web interfaces and faster turnaround with no local setup, but charge per clip. Sora and Luma AI produce high-quality output but sit behind subscription or credit paywalls. WAN's core value proposition is specifically that it is free, runs offline, supports fine-tuning, and has no usage limits beyond what your GPU can handle.

"The videos generated by Wan 2.1 are so good, it's hard to believe they were made by a free AI app." - BGR editorial review, February 26, 2025
"That changes the calculus for anyone who built on or evaluated the Wan series based on its open-source accessibility." - Tellers.ai, on the Wan 2.7 closed-weights announcement, April 15, 2026

The real cost of generating video with WAN

For the open-weight versions (2.1 through 2.6), the cost to Alibaba is zero. You pay for your own electricity and hardware. An RTX 4090 generating 5-second 480p clips at FP16 precision uses roughly 4 minutes per clip unoptimized. With FP8 quantization or a speed LoRA running 8-10 inference steps, community benchmarks bring this to around 90 seconds per clip on the same GPU. On an RTX 4070 Ti (12 GB), the 1.3B model generates at comparable speed; the 14B model requires manual quantization or GGUF format.

Cloud rental paths exist for users without local GPUs. RunPod, ThinkDiffusion, and Spheron all offer WAN-configured GPU instances. A 4090 instance on RunPod costs roughly $0.74/hour; generating 10 clips per hour means about $0.07 per clip. Still cheaper than most commercial APIs for volume work, though Wan 2.7's API rate ($0.10/second of 720p) is competitive for users who need that version's Thinking Mode features without managing infrastructure.

The actual cost barrier is time: model files range from 5 GB (1.3B, GGUF quantized) to 28 GB (14B, full FP16). First-time setup in ComfyUI or Wan2GP requires downloading models to specific directory paths, configuring text encoders and VAE separately, and troubleshooting environment dependencies. Users without Python experience report spending 2-4 hours on initial setup. After that, generation is straightforward.

Where WAN consistently breaks

Character consistency across longer clips: Independent testing gave WAN 2.5 a 3.0/5 on character consistency, with "facial features and styles shifting noticeably between frames." This affects any prompt involving a specific character appearing throughout a clip: faces drift, clothing changes slightly, body proportions shift. The Animate-14B variant in WAN 2.2 addresses portrait animation specifically, but multi-character narrative scenes remain unreliable.

Complex environmental details: Detailed landscape prompts lose specific elements mid-clip. "Consistent mist over a reflective lake with soft morning backlight" may start correctly but the mist thins, the reflection changes, or the lighting shifts by the second half. Temporal consistency in dense environmental prompts is weaker than single-subject or simple background prompts.

AMD GPU performance: GitHub issue trackers for ComfyUI show WAN 2.2 on AMD/ROCm cards sometimes runs 4-5x slower on the second clip compared to the first, due to VRAM offloading behavior. A full restart restores speed. This is a known issue that third-party patches address inconsistently.

Slow motion artifacts in WAN 2.2: Some users report that WAN 2.2 defaults to generating slow-motion-looking footage unless frame rate settings are configured explicitly. The community workaround is forcing the step count and fps settings in ComfyUI rather than relying on defaults.

Setup friction for non-technical users: WAN is not a product you open and click. It requires a Python environment, model placement in specific subdirectories, and knowledge of inference parameters. Wan2GP and third-party cloud wrappers solve this, but they add latency and sometimes limit which model variants are available.

Best use cases versus skip-this scenarios

Use WAN when: You want the best open-source video generation without a per-clip cost. You have an RTX 3060 or better (or are willing to rent GPU time). You want to fine-tune on your own footage or style. You need offline generation without API calls, rate limits, or content policies. You're already in the ComfyUI ecosystem and want video added to existing image workflows. You want commercial rights without usage fees. You're building a pipeline that needs video editing, inpainting, or instruction-guided modification in addition to initial generation.

Skip WAN when: You need a no-setup web tool and aren't comfortable with command-line environments. You're working on multi-character narrative content where face consistency is critical. Commercial tools with character memory systems will outperform WAN 2.1-2.5 on this. You're on a machine with less than 6 GB of VRAM (even Wan2GP hits a floor). You want Wan 2.7's Thinking Mode reasoning quality without paying API rates. If your priority is the fastest possible iteration speed, Pika or Runway will generate clips in seconds versus minutes, which matters for rapid creative drafts.

The February 25, 2025 release of Wan 2.1 was a genuine turning point for the open-source video generation space. It didn't just close the gap to commercial models; according to VBench scores, it led the leaderboard at launch. The Squish Effect LoRA (March 2025), a community-trained model that makes any object deform with realistic physics, became one of the most-shared AI video demos of early 2025 and illustrated what the fine-tuning ecosystem could produce on top of open weights. That community creative energy, from LoRA packs to ComfyUI workflow templates to the Wan2GP VRAM optimizer, is the most reliable sign that a model has staying power in the open-source space.

User Reviews

No reviews yet. Be the first to share your experience!

Sign in to write a review.

Featured in collections

Curated lists that include WAN (Wan-Video / Wan2GP).

Related articles

Guides and articles related to WAN (Wan-Video / Wan2GP).