1.3B or 14B.
Alibaba’s Wan team released Wan 2.1 on 25 February 2025. It has since been followed by Wan 2.2, but its 1.3B model is still one of the lightest ways to make video locally.
- T2V 1.3B is text to video on small cards. Wan’s figure is 8.19 GB of VRAM, and five seconds of 480p took about four minutes on an RTX 4090 without any speed-ups. Keep it at 480p: Wan calls 720p on the 1.3B less stable.
- T2V 14B is the full text-to-video model, at 480p or 720p, with about ten times the parameters.
- I2V 14B turns a picture into video. It comes in a 480p and a 720p version and needs a CLIP vision file on top of the usual three.
There are more: first-and-last-frame (FLF2V), VACE for editing and control, and the Fun models. They use their own workflows. If you are starting fresh with 8 GB, Wan 2.2 5B is the other option to try.
Files you need.
All from Comfy-Org’s repackage. The text encoder and VAE are the same for every Wan 2.1 model, and the 14B models of Wan 2.2 use them too.
Text to video, 1.3B
-
Model2.8 GB Download
wan2.1_t2v_1.3B_fp16.safetensorsComfyUI/models/diffusion_models/ -
Text encoder6.7 GB Download
umt5_xxl_fp8_e4m3fn_scaled.safetensorsComfyUI/models/text_encoders/ or umt5_xxl_fp16.safetensors, 11.4 GB -
VAE0.3 GB Download
wan_2.1_vae.safetensorsComfyUI/models/vae/
Text to video, 14B
-
Model14.3 GB Download
wan2.1_t2v_14B_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ or wan2.1_t2v_14B_fp16.safetensors, 28.6 GB
Image to video, 14B
-
Model16.4 GB Download
wan2.1_i2v_480p_14B_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ or the fp16 file, 32.8 GB. A 720p version exists too -
CLIP vision1.3 GB Download
clip_vision_h.safetensorsComfyUI/models/clip_vision/
models/
├── clip_vision/
│ └── clip_vision_h.safetensors (image to video only)
├── diffusion_models/
│ └── wan2.1_t2v_1.3B_fp16.safetensors
├── text_encoders/
│ └── umt5_xxl_fp8_e4m3fn_scaled.safetensors
└── vae/
└── wan_2.1_vae.safetensors
For the 14B there are two fp8 files. Pick fp8_scaled over plain fp8_e4m3fn: users rank the quality fp16, then bf16, then fp8_scaled, then fp8_e4m3fn.
Smaller files: GGUF
GGUF builds load through ComfyUI-GGUF with Unet Loader (GGUF). The encoder has one too, city96’s UMT5 GGUF, at 6.0 GB for Q8_0.
Q8_0 is close to fp16. On the 1.3B, GGUF saves little next to the 2.8 GB fp16 file; it matters for the 14B.
What fits your computer.
480p, 33 to 81 frames. ComfyUI loads the encoder first and moves it aside before sampling, so the model file sets the limit.
- 6 GBTight
1.3B with ComfyUI’s offloading. Wan’s own figure is 8.19 GB, so expect it to lean on system RAM. Keep frames low.
- 8 GB1.3B fits
1.3B at 832 × 480. The 14B doesn’t fit without GGUF and a lot of system RAM.
- 12 GBOffloads
1.3B comfortably. 14B text to video as GGUF Q4_K_M or Q5_K_M. 14B image to video was too much for one 12 GB owner.
- 16 GBOffloads
14B fp8 with part of it in system RAM, or GGUF Q5_K_M.
- 24 GBFits
14B fp8 text to video and image to video.
- Mac 24 GB+1.3B fits
1.3B fp16 with the euler sampler. M3 Max and M4 Max owners ran the example workflow; its uni_pc sampler gave them distorted video.
- Mac 64 GB+Tight
14B as fp16 (28.6 GB) or GGUF, with the fp16 encoder. fp8 files don’t run on Apple GPUs.
Set it up.
-
Update ComfyUI
Wan 2.1 has been in ComfyUI since 27 February 2025, so any current version runs it. ComfyUI Desktop updates itself. In a manual install:
Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the files
Model, text encoder and VAE from the lists above, plus
clip_vision_hfor image to video. Use Comfy-Org’s files: a similarly named model from elsewhere gave one user only noise. -
Put them in their folders
Model in
diffusion_models, encoder intext_encoders, VAE invae, CLIP vision inclip_vision. Restart ComfyUI. -
Open the template
In the template browser, pick Wan 2.1 Text to Video or Wan 2.1 Image to Video. The text template is set up for the 1.3B.
-
Check the loaders
Load CLIP gets UMT5 with type
wan. Load VAE getswan_2.1_vae. For image to video, Load CLIP Vision getsclip_vision_h, and your picture goes intoWanImageToVideo. -
Write a prompt and run
Describe the scene and the motion. Keep the template’s Chinese negative prompt: it is Wan’s standard one.
Settings that work.
Comfy’s template follows Wan’s advice for the 1.3B. For the 14B, Wan’s own script uses a lower shift and CFG, more steps and more frames.
1.3B
- Steps
- 30Wan’s script: 50
- CFG
- 6Wan’s advice for 1.3B
- Sampler
- uni_pceuler on a Mac
- Scheduler
- simple
- Shift
- 8Wan: 8 to 12
- Size
- 832 × 480
- Frames
- 3381 for five seconds
- FPS
- 16
14B text to video, Wan’s settings
- Steps
- 50Comfy template: 30
- CFG
- 5Comfy template: 6
- Sampler
- uni_pceuler on a Mac
- Scheduler
- simple
- Shift
- 5Comfy template: 8
- Size
- 1280 × 720or 832 × 480
- Frames
- 815 s at 16 fps
- FPS
- 16
For image to video, Wan’s script uses 40 steps, with shift 3 at 832 × 480. Comfy’s template runs 20 steps at CFG 6, shift 8, 512 × 512 and 33 frames. Frame counts go in steps of four plus one, like 33, 49 or 81.
When it goes wrong.
- ComfyUI shows “reconnecting” or a token error at the text encoder
- The wrong text encoder. comfyanonymous’s answer: use the UMT5 file from Comfy-Org’s repackage.
mat1 and mat2 shapes cannot be multiplied (512x768 and 4096x5120)- A CLIP-L style encoder went into a model that expects UMT5. Load
umt5_xxl_fp8_e4m3fn_scaledwith typewan. - Only noise comes out
- A similarly named model from another site. The Comfy-Org files fixed it.
- Mac: distorted or banded video
- The uni_pc sampler on Apple GPUs. Switch to euler.
Trying to convert Float8_e4m3fn to the MPS backend- fp8 files on a Mac. Use fp16 files or GGUF.
- Image to video runs out of memory on 12 GB
- The 14B I2V is heavy. People pointed to Wan 2.1 Fun 1.3B InP, a small image-to-video model, or GGUF.
- ComfyUI doesn’t see the GGUF files
- Install ComfyUI-GGUF and swap Load Diffusion Model for Unet Loader (GGUF).
Is there an image-to-video model at a lower size? My 12 GB card can’t run the 14B.
What’s the difference between fp8 e4m3fn and fp8 scaled, and which should I use?
Questions.
How much VRAM does Wan 2.1 need?
The 1.3B text-to-video model needs about 8 GB by Wan’s figure and runs on 8 GB cards at 480p. The 14B is 14.3 GB in fp8, so it wants 16 to 24 GB, or GGUF files and system RAM on 12 GB.
Which text encoder does Wan 2.1 use?
UMT5-XXL, loaded with Load CLIP set to type wan. Comfy-Org’s umt5_xxl_fp8_e4m3fn_scaled (6.7 GB) is the usual one; the fp16 version is 11.4 GB. Wan 2.2 uses the same encoder.
What extra file does Wan 2.1 image to video need?
A CLIP vision encoder, clip_vision_h.safetensors (1.3 GB), in the models/clip_vision folder. Wan 2.2’s image-to-video model doesn’t need one.
Can Wan 2.1 make videos longer than five seconds?
Not in one go: clips stay around five seconds, 81 frames at 16 fps. Longer videos are made by chaining clips. Sound needs a separate model such as MMAudio.
Should I use Wan 2.1 or Wan 2.2?
Wan 2.2 is newer, and its 14B models use the same text encoder and VAE as Wan 2.1 14B. For 8 GB cards, the choice is between Wan 2.1 1.3B and Wan 2.2 5B, which makes 720p but needs its own VAE.
Does Wan 2.1 run on a Mac?
Yes, the 1.3B runs on Apple Silicon. Use fp16 files and switch the sampler from uni_pc to euler, because uni_pc gives distorted video on Apple GPUs. fp8 files don’t load on a Mac.
Sources: Wan2.1 repository, Wan2.1 generate.py, ComfyUI Wan 2.1 tutorial, Comfy-Org Wan 2.1 repackage, encoder thread, Mac output thread, ComfyUI issue #7027, WanVideoWrapper issue #461, length and audio thread.