Video model · beta

How to run Wan 2.2 locally.

Wan 2.2 runs in ComfyUI with a model, the UMT5-XXL text encoder and a VAE. The 5B model fits an 8 GB card with ComfyUI’s offloading and has its own Wan 2.2 VAE. The 14B models come as a high-noise and a low-noise half, use the older Wan 2.1 VAE and want 16 to 24 GB, or GGUF files on 12 GB.

Updated 29 Sep 202612 min read

Maker
Wan-AI, Alibaba
Released
Jul 202528 July, 5B and both 14B models
Licence
Apache 2.0Commercial use allowed
Memory
8 GB and up5B with offloading. 14B from 12 GB as GGUF

5B or 14B.

Alibaba’s Wan team released Wan 2.2 on 28 July 2025, with ComfyUI support the same day. Three models are made for this kind of use, and they are built differently.

  • TI2V 5B is one dense 5B model that does text to video and image to video. It makes 720p at 24 fps, which for this model means 1280 × 704, and 121 frames, five seconds. It has its own VAE.
  • T2V A14B is text to video from two 14B experts. The high-noise model does the early steps, where layout and motion are set, and the low-noise model does the late steps, where detail goes in. You need both files, and ComfyUI runs them one after the other.
  • I2V A14B is the same pair for image to video. Unlike Wan 2.1’s image model, it needs no CLIP vision file.

Pick 5B for an 8 or 12 GB card, a Mac, or when you want quick tries at 720p. Pick 14B when quality comes first and you have 16 GB or more, or 12 GB with GGUF files and patience. One RTX 3060 owner who moved from 5B to 14B GGUF found the quality much better. Wan 2.2 also has speech-to-video and Animate models; they need their own workflows and aren’t covered here.

Files you need.

Every version takes the same text encoder, UMT5-XXL. The VAE is where people go wrong: only the 5B uses the Wan 2.2 VAE. Both 14B models use the Wan 2.1 VAE.

Wan 2.2 5B

  • Model wan2.2_ti2v_5B_fp16.safetensors ComfyUI/models/diffusion_models/
    10.0 GB Download
  • Text encoder umt5_xxl_fp8_e4m3fn_scaled.safetensors ComfyUI/models/text_encoders/
    6.7 GB Download
  • VAE wan2.2_vae.safetensors ComfyUI/models/vae/
    1.4 GB Download

Wan 2.2 14B, text to video

  • High-noise model wan2.2_t2v_high_noise_14B_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ or the fp16 file, 28.6 GB
    14.3 GB Download
  • Low-noise model wan2.2_t2v_low_noise_14B_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ or the fp16 file, 28.6 GB
    14.3 GB Download
  • Text encoder umt5_xxl_fp8_e4m3fn_scaled.safetensors ComfyUI/models/text_encoders/
    6.7 GB Download
  • VAE wan_2.1_vae.safetensors ComfyUI/models/vae/
    0.3 GB Download
  • 4-step LoRA, high wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors ComfyUI/models/loras/ optional, the Comfy template uses it
    1.2 GB Download
  • 4-step LoRA, low wan2.2_t2v_lightx2v_4steps_lora_v1.1_low_noise.safetensors ComfyUI/models/loras/ optional, the Comfy template uses it
    1.2 GB Download

Wan 2.2 14B, image to video

The same encoder and the same Wan 2.1 VAE, with the I2V pair and its own LoRAs. T2V LoRAs damage I2V output, so keep them apart.

  • High-noise model wan2.2_i2v_high_noise_14B_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ or the fp16 file, 28.6 GB
    14.3 GB Download
  • Low-noise model wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ or the fp16 file, 28.6 GB
    14.3 GB Download
  • 4-step LoRA, high wan2.2_i2v_lightx2v_4steps_lora_v1_high_noise.safetensors ComfyUI/models/loras/ optional
    1.2 GB Download
  • 4-step LoRA, low wan2.2_i2v_lightx2v_4steps_lora_v1_low_noise.safetensors ComfyUI/models/loras/ optional
    1.2 GB Download
ComfyUI/models, 14B text to video
models/
├── diffusion_models/
│   ├── wan2.2_t2v_high_noise_14B_fp8_scaled.safetensors
│   └── wan2.2_t2v_low_noise_14B_fp8_scaled.safetensors
├── loras/
│   ├── wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors
│   └── wan2.2_t2v_lightx2v_4steps_lora_v1.1_low_noise.safetensors
├── text_encoders/
│   └── umt5_xxl_fp8_e4m3fn_scaled.safetensors
└── vae/
    └── wan_2.1_vae.safetensors

fp16 files look a little better than fp8 and take twice the space. Among fp8 files, fp8_scaled beats plain fp8_e4m3fn. Comfy-Org has no fp8 build of the 5B.

Smaller files: GGUF

QuantStack’s GGUF builds load through the ComfyUI-GGUF nodes with Unet Loader (GGUF). For 14B the sizes below are per half: download the HighNoise and the LowNoise file of the same level.

ModelQ3_K_MQ4_K_MQ5_K_MQ6_KQ8_0
5B2.6 GB3.4 GB3.8 GB4.2 GB5.4 GB
14B T2V7.2 GB9.7 GB10.8 GB12.0 GB15.4 GB
14B I2V7.2 GB9.7 GB10.8 GB12.0 GB15.4 GB

14B sizes are per half, so double them. Q8_0 is close to fp16. Q4_K_M and Q5_K_M are what most 12 and 16 GB owners run. The 14B T2V also has a Q2_K at 5.3 GB.

What fits your computer.

Video needs far more working memory than an image, and it grows with size and frame count. ComfyUI moves what doesn’t fit into system RAM, so most setups below still run, only slower.

  • 6 GBTight

    5B as a small GGUF (Q4_K_M is 3.4 GB) at a smaller size with fewer frames. No reliable report at this tier; 8 GB is the practical floor.

  • 8 GB5B fits

    ComfyUI’s own docs say the 5B fits 8 GB with its offloading. For 14B, the advice given on Hugging Face is GGUF Q2 or Q3 per half.

  • 12 GBOffloads

    5B comfortably. 14B as Q4 GGUF: an RTX 3060 with 32 GB of RAM makes five seconds in 10 to 12 minutes with a 4-step LoRA.

  • 16 GBOffloads

    14B as GGUF Q4 to Q6, or fp8 with part of it in RAM. A 16 GB RTX 4090 laptop ran I2V Q3_K_M without spilling into RAM.

  • 24 GBFits

    14B fp8 pair at 640 × 640, 81 frames: Comfy measured about 84% of an RTX 4090D’s memory.

  • Mac 16 to 24 GBTight

    5B as fp16 or GGUF, with the euler sampler. fp8 files don’t run on Apple GPUs.

  • Mac 32 to 48 GB5B fits

    5B at 832 × 480, 121 frames ran on a 48 GB M4 Max at about 39 s per step. 14B only as GGUF.

  • Mac 64 GB+Tight

    The 14B fp16 pair is 57.2 GB of weights, so 96 to 128 GB is more realistic, or GGUF. A 128 GB M1 Ultra gave black video at 81 frames.

Set it up.

  1. Update ComfyUI

    An old ComfyUI is behind the most common Wan 2.2 error, the 36-channel one below. ComfyUI Desktop updates itself; the portable build has an update script, and if that doesn’t take, a fresh portable copy does. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the files

    From the lists above: model (or both halves), text encoder and the right VAE. Add the two 4-step LoRAs if you want the fast template.

  3. Put them in their folders

    Models in diffusion_models, encoder in text_encoders, VAE in vae, LoRAs in loras. Restart ComfyUI so the loaders list them.

  4. Open the template

    In the template browser, pick Wan 2.2 5B Video Generation, Wan 2.2 14B Text to Video or Wan 2.2 14B Image to Video. The 14B templates come with the 4-step LoRAs switched on.

  5. Check the loaders

    Load CLIP gets UMT5 with type wan. Load VAE gets wan2.2_vae for 5B and wan_2.1_vae for 14B. On 14B there are two Load Diffusion Model nodes: high-noise feeds the first KSamplerAdvanced, low-noise the second. For image to video, the 5B takes your picture in Wan22ImageToVideoLatent, the 14B in WanImageToVideo.

  6. Write a prompt and run

    Describe the scene and the motion. The templates carry Wan’s standard Chinese negative prompt; leave it in. The first run loads everything and takes longest.

Settings that work.

Wan’s own settings and Comfy’s templates differ. Wan’s scripts use more steps and a different shift; Comfy trades some quality for speed. Both are listed so you can choose.

5B, Wan’s settings

Steps
50Comfy template: 20
CFG
5
Sampler
uni_pceuler on a Mac
Scheduler
simple
Shift
5Comfy template: 8
Size
1280 × 704or 704 × 1280
Frames
1215 s at 24 fps
Text encoder
Load CLIP, wan

14B text to video, Wan’s settings

Steps
40Comfy examples: 20
CFG
4 / 3high / low. Comfy: 3.5
Sampler
euleras in Comfy’s examples
Shift
12Comfy examples: 8
Size
1280 × 720or 832 × 480
Frames
815 s at 16 fps
Switch
by noise levelComfy: half the steps each
Text encoder
Load CLIP, wan

For 14B image to video, Wan uses 40 steps, CFG 3.5 on both halves and shift 5. Comfy’s older non-LoRA example uses 20 steps, CFG 3.5 and shift 8. In KSamplerAdvanced, the high-noise sampler runs from step 0 to half, with leftover noise returned, and the low-noise sampler finishes.

14B with the 4-step LoRAs (Comfy’s template)

Steps
42 high, then 2 low
CFG
1
Sampler
euler
Scheduler
simple
Shift
5on both halves
Size
640 × 640
Frames
81at 16 fps
LoRA strength
1.0one LoRA per half

Four steps is the total, not per model. The LoRAs make motion slower, which lightx2v acknowledges. What helped people: a larger size (1280 × 720 moved at normal speed where 832 × 480 was slow), a stronger high-noise LoRA, newer LoRA releases, or a third sampler that runs one or two high-noise steps without the LoRA at CFG 3.5 first.

How fast.

GPUModelVideoTime
RTX 4090D 24 GB14B T2V fp8, no LoRA640 × 640, 81 frames536 s, 513 s after[1]
RTX 4090D 24 GB14B T2V fp8, 4-step LoRA640 × 640, 81 frames108 s, 71 s after[1]
RTX 4090D 24 GB14B I2V fp8, 4-step LoRA640 × 640, 81 frames97 s, 71 s after[1]
RTX 3060 12 GB14B T2V, Q4 GGUF448 × 415, 5 s20 min[2]
RTX 3060 12 GB14B, 4 to 8 step LoRA5 s10 to 15 min[3]
M4 Max 48 GB5B832 × 480, 121 frames39 s per step[4]

The first time includes loading the models. On the RTX 3060 with a LoRA, text to video took 10 to 12 minutes and image to video about 15. Wan’s own figure for the 5B is under nine minutes for five seconds of 720p on one consumer GPU.

When it goes wrong.

weight of size [5120, 36, 1, 2, 2], expected input[1, 32, 21, 96, 96] to have 36 channels, but got 32 channels instead
The classic 14B image-to-video error. Update ComfyUI (a fresh portable copy if the updater doesn’t work). Delete the WanImageToVideo (Flow2) node pack, which overwrites core code. And use wan_2.1_vae, not the Wan 2.2 VAE.
expected input[1, 16, …] to have 48 channels, but got 16 channels instead
The Wan 2.2 VAE on a 14B model. 14B uses wan_2.1_vae.safetensors; only the 5B takes wan2.2_vae.
size mismatch for encoder.conv1.weight
Wrong VAE for the model. Same fix: the 2.1 VAE for 14B, the 2.2 VAE for 5B.
mat1 and mat2 shapes cannot be multiplied (462x768 and 4096x5120)
The wrong text encoder. Use umt5_xxl_fp8_e4m3fn_scaled from Comfy-Org in Load CLIP with type wan.
lora key not loaded: blocks.0.cross_attn…
Some original lightx2v LoRA files don’t match ComfyUI’s key names. Kijai’s converted versions in Kijai/WanVideo_comfy load cleanly.
Everything moves in slow motion
A side effect of the 4-step LoRAs. Go larger, raise the high-noise LoRA strength, or add a first high-noise pass without the LoRA, as described above.
The 5B makes only brown fog
A broken install. A fresh ComfyUI with the models moved over fixed it.
Mac: the first frame is fine, then mush or bands
The uni_pc sampler on Apple GPUs. Switch to euler.
Trying to convert Float8_e4m3fn to the MPS backend
fp8 files on a Mac. Use fp16 files or GGUF.
ComfyUI gets killed or freezes on a 32 GB PC
Out of system RAM with the fp8 pair. Use GGUF, add swap, or try --cache-ram.
ComfyUI doesn’t see the GGUF files
Install ComfyUI-GGUF and swap Load Diffusion Model for Unet Loader (GGUF).

Low noise and high noise: what is each one for, and do I need both?

Hugging Face, Comfy-Org Wan 2.2 repackage

Why do the A14B models use the Wan 2.1 VAE and not the new one?

Hugging Face, Wan2.2-T2V-A14B

Questions.

Which VAE does Wan 2.2 use?

Only Wan 2.2 5B uses the new wan2.2_vae (1.4 GB). Both 14B models, text to video and image to video, use wan_2.1_vae (0.3 GB). Mixing them up gives a channel mismatch error.

Do I need both the high-noise and the low-noise file?

Yes, for 14B. The high-noise model does the first half of the steps and the low-noise model the second, in two KSamplerAdvanced nodes. With GGUF, download the same quant level for both halves.

Can Wan 2.2 run on 8 GB of VRAM?

The 5B can: ComfyUI’s docs say it fits 8 GB with its offloading, and GGUF files make it lighter. 14B on 8 GB means Q2 or Q3 GGUF halves, a lot of system RAM and a long wait.

Why is my Wan 2.2 video in slow motion?

The 4-step lightx2v LoRAs slow motion down, and their makers say so. A larger size, a stronger high-noise LoRA, newer LoRA releases or a first high-noise pass without the LoRA at CFG 3.5 all help.

How many frames can Wan 2.2 make?

The 14B models are made for 81 frames at 16 fps and the 5B for 121 frames at 24 fps, both about five seconds. Frame counts go in steps of four plus one, like 49, 81 or 121.

Does Wan 2.2 run on a Mac?

The 5B does, on Apple Silicon with 24 GB or more, using fp16 or GGUF files and the euler sampler. uni_pc gives broken video on Apple GPUs and fp8 files don’t load. 14B needs GGUF or a lot of memory.

Can I use Wan 2.2 commercially?

Yes. All three models are Apache 2.0, and Wan’s model card says they claim no rights over what you generate.

Sources: Wan2.2 repository, TI2V-5B model card, T2V-A14B config, ComfyUI Wan 2.2 tutorial, Comfy-Org workflow templates [1], ComfyUI examples, QuantStack GGUF thread [2], Civitai RTX 3060 workflow [3], ComfyUI issue #16644 [4], 36-channel thread, ComfyUI issue #9092, uni_pc on Mac, #16573, RAM thread, #10771, slow-motion thread.

HEISS UI

Both halves, paired for you.

HEISS UI runs Wan 2.2 text and image to video on the ComfyUI you already have, in the Video tab (beta). Pick a file and write a prompt.

  • The two halves, paired. Pick one half of the 14B pair and the other is found next to it, or fetched.
  • The right VAE, every time. 5B and 14B each get the one that belongs to them.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • Failures that explain themselves. When a run runs out of memory, it says so and offers the fix.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.