Video model · beta

How to run Wan 2.1 locally.

Wan 2.1 runs in ComfyUI with three files: the model, the UMT5-XXL text encoder and the Wan 2.1 VAE. The 1.3B text-to-video model needs about 8 GB and makes 480p clips; the 14B is 14.3 GB in fp8 and wants 16 to 24 GB, or GGUF files on 12 GB. Image to video uses a 14B model and one more file, a CLIP vision encoder.

Updated 29 Sep 20269 min read

Maker
Wan-AI, Alibaba
Released
Feb 202525 February, in ComfyUI two days later
Licence
Apache 2.0Commercial use allowed
Memory
8 GB and up1.3B at 480p. 14B from 12 GB as GGUF

1.3B or 14B.

Alibaba’s Wan team released Wan 2.1 on 25 February 2025. It has since been followed by Wan 2.2, but its 1.3B model is still one of the lightest ways to make video locally.

  • T2V 1.3B is text to video on small cards. Wan’s figure is 8.19 GB of VRAM, and five seconds of 480p took about four minutes on an RTX 4090 without any speed-ups. Keep it at 480p: Wan calls 720p on the 1.3B less stable.
  • T2V 14B is the full text-to-video model, at 480p or 720p, with about ten times the parameters.
  • I2V 14B turns a picture into video. It comes in a 480p and a 720p version and needs a CLIP vision file on top of the usual three.

There are more: first-and-last-frame (FLF2V), VACE for editing and control, and the Fun models. They use their own workflows. If you are starting fresh with 8 GB, Wan 2.2 5B is the other option to try.

Files you need.

All from Comfy-Org’s repackage. The text encoder and VAE are the same for every Wan 2.1 model, and the 14B models of Wan 2.2 use them too.

Text to video, 1.3B

  • Model wan2.1_t2v_1.3B_fp16.safetensors ComfyUI/models/diffusion_models/
    2.8 GB Download
  • Text encoder umt5_xxl_fp8_e4m3fn_scaled.safetensors ComfyUI/models/text_encoders/ or umt5_xxl_fp16.safetensors, 11.4 GB
    6.7 GB Download
  • VAE wan_2.1_vae.safetensors ComfyUI/models/vae/
    0.3 GB Download

Text to video, 14B

  • Model wan2.1_t2v_14B_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ or wan2.1_t2v_14B_fp16.safetensors, 28.6 GB
    14.3 GB Download

Image to video, 14B

  • Model wan2.1_i2v_480p_14B_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ or the fp16 file, 32.8 GB. A 720p version exists too
    16.4 GB Download
  • CLIP vision clip_vision_h.safetensors ComfyUI/models/clip_vision/
    1.3 GB Download
ComfyUI/models
models/
├── clip_vision/
│   └── clip_vision_h.safetensors        (image to video only)
├── diffusion_models/
│   └── wan2.1_t2v_1.3B_fp16.safetensors
├── text_encoders/
│   └── umt5_xxl_fp8_e4m3fn_scaled.safetensors
└── vae/
    └── wan_2.1_vae.safetensors

For the 14B there are two fp8 files. Pick fp8_scaled over plain fp8_e4m3fn: users rank the quality fp16, then bf16, then fp8_scaled, then fp8_e4m3fn.

Smaller files: GGUF

GGUF builds load through ComfyUI-GGUF with Unet Loader (GGUF). The encoder has one too, city96’s UMT5 GGUF, at 6.0 GB for Q8_0.

ModelQ4_K_MQ5_K_MQ8_0
T2V 1.3B1.0 GBnot listed1.5 GB
T2V 14B10.1 GB11.3 GB15.9 GB

Q8_0 is close to fp16. On the 1.3B, GGUF saves little next to the 2.8 GB fp16 file; it matters for the 14B.

What fits your computer.

480p, 33 to 81 frames. ComfyUI loads the encoder first and moves it aside before sampling, so the model file sets the limit.

  • 6 GBTight

    1.3B with ComfyUI’s offloading. Wan’s own figure is 8.19 GB, so expect it to lean on system RAM. Keep frames low.

  • 8 GB1.3B fits

    1.3B at 832 × 480. The 14B doesn’t fit without GGUF and a lot of system RAM.

  • 12 GBOffloads

    1.3B comfortably. 14B text to video as GGUF Q4_K_M or Q5_K_M. 14B image to video was too much for one 12 GB owner.

  • 16 GBOffloads

    14B fp8 with part of it in system RAM, or GGUF Q5_K_M.

  • 24 GBFits

    14B fp8 text to video and image to video.

  • Mac 24 GB+1.3B fits

    1.3B fp16 with the euler sampler. M3 Max and M4 Max owners ran the example workflow; its uni_pc sampler gave them distorted video.

  • Mac 64 GB+Tight

    14B as fp16 (28.6 GB) or GGUF, with the fp16 encoder. fp8 files don’t run on Apple GPUs.

Set it up.

  1. Update ComfyUI

    Wan 2.1 has been in ComfyUI since 27 February 2025, so any current version runs it. ComfyUI Desktop updates itself. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the files

    Model, text encoder and VAE from the lists above, plus clip_vision_h for image to video. Use Comfy-Org’s files: a similarly named model from elsewhere gave one user only noise.

  3. Put them in their folders

    Model in diffusion_models, encoder in text_encoders, VAE in vae, CLIP vision in clip_vision. Restart ComfyUI.

  4. Open the template

    In the template browser, pick Wan 2.1 Text to Video or Wan 2.1 Image to Video. The text template is set up for the 1.3B.

  5. Check the loaders

    Load CLIP gets UMT5 with type wan. Load VAE gets wan_2.1_vae. For image to video, Load CLIP Vision gets clip_vision_h, and your picture goes into WanImageToVideo.

  6. Write a prompt and run

    Describe the scene and the motion. Keep the template’s Chinese negative prompt: it is Wan’s standard one.

Settings that work.

Comfy’s template follows Wan’s advice for the 1.3B. For the 14B, Wan’s own script uses a lower shift and CFG, more steps and more frames.

1.3B

Steps
30Wan’s script: 50
CFG
6Wan’s advice for 1.3B
Sampler
uni_pceuler on a Mac
Scheduler
simple
Shift
8Wan: 8 to 12
Size
832 × 480
Frames
3381 for five seconds
FPS
16

14B text to video, Wan’s settings

Steps
50Comfy template: 30
CFG
5Comfy template: 6
Sampler
uni_pceuler on a Mac
Scheduler
simple
Shift
5Comfy template: 8
Size
1280 × 720or 832 × 480
Frames
815 s at 16 fps
FPS
16

For image to video, Wan’s script uses 40 steps, with shift 3 at 832 × 480. Comfy’s template runs 20 steps at CFG 6, shift 8, 512 × 512 and 33 frames. Frame counts go in steps of four plus one, like 33, 49 or 81.

When it goes wrong.

ComfyUI shows “reconnecting” or a token error at the text encoder
The wrong text encoder. comfyanonymous’s answer: use the UMT5 file from Comfy-Org’s repackage.
mat1 and mat2 shapes cannot be multiplied (512x768 and 4096x5120)
A CLIP-L style encoder went into a model that expects UMT5. Load umt5_xxl_fp8_e4m3fn_scaled with type wan.
Only noise comes out
A similarly named model from another site. The Comfy-Org files fixed it.
Mac: distorted or banded video
The uni_pc sampler on Apple GPUs. Switch to euler.
Trying to convert Float8_e4m3fn to the MPS backend
fp8 files on a Mac. Use fp16 files or GGUF.
Image to video runs out of memory on 12 GB
The 14B I2V is heavy. People pointed to Wan 2.1 Fun 1.3B InP, a small image-to-video model, or GGUF.
ComfyUI doesn’t see the GGUF files
Install ComfyUI-GGUF and swap Load Diffusion Model for Unet Loader (GGUF).

Is there an image-to-video model at a lower size? My 12 GB card can’t run the 14B.

Hugging Face, Comfy-Org Wan 2.1 repackage

What’s the difference between fp8 e4m3fn and fp8 scaled, and which should I use?

Hugging Face, Comfy-Org Wan 2.1 repackage

Questions.

How much VRAM does Wan 2.1 need?

The 1.3B text-to-video model needs about 8 GB by Wan’s figure and runs on 8 GB cards at 480p. The 14B is 14.3 GB in fp8, so it wants 16 to 24 GB, or GGUF files and system RAM on 12 GB.

Which text encoder does Wan 2.1 use?

UMT5-XXL, loaded with Load CLIP set to type wan. Comfy-Org’s umt5_xxl_fp8_e4m3fn_scaled (6.7 GB) is the usual one; the fp16 version is 11.4 GB. Wan 2.2 uses the same encoder.

What extra file does Wan 2.1 image to video need?

A CLIP vision encoder, clip_vision_h.safetensors (1.3 GB), in the models/clip_vision folder. Wan 2.2’s image-to-video model doesn’t need one.

Can Wan 2.1 make videos longer than five seconds?

Not in one go: clips stay around five seconds, 81 frames at 16 fps. Longer videos are made by chaining clips. Sound needs a separate model such as MMAudio.

Should I use Wan 2.1 or Wan 2.2?

Wan 2.2 is newer, and its 14B models use the same text encoder and VAE as Wan 2.1 14B. For 8 GB cards, the choice is between Wan 2.1 1.3B and Wan 2.2 5B, which makes 720p but needs its own VAE.

Does Wan 2.1 run on a Mac?

Yes, the 1.3B runs on Apple Silicon. Use fp16 files and switch the sampler from uni_pc to euler, because uni_pc gives distorted video on Apple GPUs. fp8 files don’t load on a Mac.

Sources: Wan2.1 repository, Wan2.1 generate.py, ComfyUI Wan 2.1 tutorial, Comfy-Org Wan 2.1 repackage, encoder thread, Mac output thread, ComfyUI issue #7027, WanVideoWrapper issue #461, length and audio thread.

HEISS UI

Type, and get a video.

HEISS UI runs Wan 2.1 text to video on the ComfyUI you already have, in the Video tab (beta). No graph, just a prompt.

  • Video from the prompt box. The same box as your images. Set length and frame rate, then generate.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • Failures that explain themselves. When a run runs out of memory, it says so and offers the fix.

Image to video with Wan 2.1 runs as your own workflow. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.