480p or 720p.
Tencent released HunyuanVideo 1.5 on 20 November 2025: an 8.3B model that makes five-second clips at 24 fps, from text or from a picture. Every model is its own file, trained for one resolution and one task.
- 480p text to video and image to video. The ones to start with on 12 to 16 GB.
- 720p text to video and image to video. What Comfy’s templates use. Sharper, and much slower.
- CFG-distilled versions of both run at CFG 1, one pass per step instead of two. Tencent still asks for 50 steps.
- 480p I2V step-distilled, added on 5 December, runs in 8 or 12 steps. Tencent says an RTX 4090 makes a clip with it in about 75 seconds.
- Super-resolution models take a finished clip up to 720p or 1080p. More on that under settings.
Image to video needs one extra file, a SigLIP vision encoder. Text in the video, like a sign or a title, comes from the second text encoder, ByT5.
Files you need.
All from Comfy-Org’s repackage. The encoders and the VAE are shared by every model.
Shared parts
-
Text encoder9.4 GB Download
qwen_2.5_vl_7b_fp8_scaled.safetensorsComfyUI/models/text_encoders/ or qwen_2.5_vl_7b.safetensors, bf16, 16.6 GB -
Glyph encoder0.4 GB Download
byt5_small_glyphxl_fp16.safetensorsComfyUI/models/text_encoders/ -
VAE2.5 GB Download
hunyuanvideo15_vae_fp16.safetensorsComfyUI/models/vae/
Text to video, pick one
-
480p, CFG-distilled8.3 GB Download
hunyuanvideo1.5_480p_t2v_cfg_distilled_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ fp16: 16.7 GB -
480p16.7 GB Download
hunyuanvideo1.5_480p_t2v_fp16.safetensorsComfyUI/models/diffusion_models/ -
720p16.7 GB Download
hunyuanvideo1.5_720p_t2v_fp16.safetensorsComfyUI/models/diffusion_models/ the Comfy template’s model
Image to video, pick one, plus the vision file
-
480p, step-distilled8.3 GB Download
hunyuanvideo1.5_480p_i2v_step_distilled_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ fp16: 16.7 GB -
720p16.7 GB Download
hunyuanvideo1.5_720p_i2v_fp16.safetensorsComfyUI/models/diffusion_models/ the Comfy template’s model. CFG-distilled fp8: 8.3 GB -
Vision encoder0.9 GB Download
sigclip_vision_patch14_384.safetensorsComfyUI/models/clip_vision/
Optional
-
1080p upscaler8.3 GB Download
hunyuanvideo1.5_1080p_sr_distilled_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ fp16: 16.7 GB. Needs the latent upsampler below -
Latent upsampler0.2 GB Download
hunyuanvideo15_latent_upsampler_1080p.safetensorsComfyUI/models/latent_upscale_models/ -
4-step LoRA0.3 GB Download
hunyuanvideo1.5_t2v_480p_lightx2v_4step_lora_rank_32_bf16.safetensorsComfyUI/models/loras/ for 480p text to video
models/
├── clip_vision/
│ └── sigclip_vision_patch14_384.safetensors (image to video)
├── diffusion_models/
│ └── hunyuanvideo1.5_480p_t2v_cfg_distilled_fp8_scaled.safetensors
├── text_encoders/
│ ├── qwen_2.5_vl_7b_fp8_scaled.safetensors
│ └── byt5_small_glyphxl_fp16.safetensors
└── vae/
└── hunyuanvideo15_vae_fp16.safetensors
The plain 720p models have no fp8 file. If they run out of memory, Comfy’s template note says to set weight_dtype in Load Diffusion Model to fp8_e4m3fn.
Smaller files: GGUF
jayn7’s GGUF builds load through ComfyUI-GGUF with Unet Loader (GGUF). The same uploader has 720p text-to-video and image-to-video GGUFs at similar sizes.
| Model | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|
| 480p T2V | 5.1 GB | 6.1 GB | 7.0 GB | 9.0 GB |
The repo has the plain and the CFG-distilled 480p model. Q8_0 is close to fp16.
What fits your computer.
The model file is only part of it. At 121 frames the VAE decode at the end is where 12 GB cards run out, and VAE Decode (Tiled) is the fix.
- 6 to 8 GBUntested
No ComfyUI report found at this size. The 480p Q4_K_M GGUF is 5.1 GB. WanGP lists HunyuanVideo 1.5 from 6 GB.
- 12 GBOffloads
480p in fp8 or GGUF. An RTX 3060 ran 480p at 121 frames. An RTX 4070 was fine at 81 frames and spilled over at 121, in the VAE decode.
- 16 GB480p fits
480p fp8. An RTX 4060 Ti 16 GB ran the CFG-distilled 480p model at 640 × 480, 81 frames.
- 24 GBFits
720p fp16. The default 720p text-to-video template took about 20 minutes on an RTX 4090.
- 32 GBFits
RTX 5090: 848 × 480 at 121 frames in 284 s with the 720p model.
- Mac 48 GB+Untested
No verified Mac report. fp8 files don’t load on Apple GPUs; the fp16 model, bf16 encoder and VAE add up to 35.8 GB.
- 12 GB graphics cards
- RTX 3060 12 GB
- 16 GB cards
- RTX 5060 Ti 16 GB
- 24 GB cards
- RTX 5090
- Mac with Apple Silicon
Set it up.
-
Update ComfyUI
HunyuanVideo 1.5 needs core nodes from late November 2025, such as
HunyuanVideo15ImageToVideo. ComfyUI Desktop was behind at first, so check that it has updated. In a manual install:Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the files
The two text encoders, the VAE and one model from the lists above. For image to video, add the SigLIP vision file.
-
Put them in their folders
Model in
diffusion_models, both encoders intext_encoders, VAE invae, SigLIP inclip_vision. Restart ComfyUI. -
Open the template
In the template browser, pick Hunyuan Video 1.5 Text to Video or Hunyuan Video 1.5 Image to Video. Both are set up for the 720p models, with a 1080p upscale stage that starts switched off.
-
Check the loaders
DualCLIPLoader gets Qwen2.5-VL and ByT5 with type
hunyuan_video_15. For image to video, Load CLIP Vision gets the SigLIP file. If you swapped in a 480p model, set the size to 848 × 480 and the shift to 5. -
Write a prompt and run
Be specific about subject, motion and camera. Tencent’s own code rewrites prompts with a large language model before generating; in ComfyUI you write the full prompt yourself.
Settings that work.
Tencent sets the shift per model, and it matters. Comfy’s templates use 20 steps instead of 50, because 50 “just take too long”, and shift 7 on everything.
480p, Tencent’s settings
- Steps
- 50Comfy template: 20
- CFG
- 61 for CFG-distilled
- Sampler
- euler
- Scheduler
- simple
- Shift
- 5Comfy template: 7
- Size
- 848 × 480
- Frames
- 1215 s at 24 fps
- Negative
- emptyTencent’s default
| Model | CFG | Shift | Steps |
|---|---|---|---|
| 480p T2V and I2V | 6 | 5 | 50 |
| 720p T2V | 6 | 9 | 50 |
| 720p I2V | 6 | 7 | 50 |
| 480p CFG-distilled | 1 | 5 | 50 |
| 720p I2V CFG-distilled | 1 | 7 | 50 |
| 480p I2V step-distilled | 1 | 7 | 8 or 12 |
| Upscale to 720p | 1 | 2 | 6 |
| Upscale to 1080p | 1 | 2 | 8 |
From Tencent’s model card, repeated in Comfy’s template notes. Comfy’s 720p templates use shift 7 for text to video too.
The 4-step lightx2v LoRA is the fast route for 480p text to video, at CFG 1. It also seems to work at 720p, but people found the results faded and low in contrast, especially in low light.
How fast.
| GPU | Model | Video | Time |
|---|---|---|---|
| RTX 5090 | 720p model | 848 × 480, 121 frames | 284 s[1] |
| RTX 5090 | 720p model | 848 × 480, 245 frames | 498 s[1] |
| RTX 4090 | 720p model, undervolted | 848 × 480, 121 frames | 297 s[1] |
| RTX 4090 | 720p T2V template | 1280 × 720, 121 frames | 20 min[2] |
| RTX 4090 | 480p I2V step-distilled | 480p | 75 s[3] |
| RTX 2080 Ti | 4 steps | 512 × 512, 49 frames | 80 s[4] |
Tencent’s 75 s figure is “within 75 seconds” for the step-distilled model. On the 2080 Ti, the VAE alone took about 20 s, and Wan 2.2 in a rapid merge took 31 s for the same clip.
When it goes wrong.
- The video is black
- SageAttention on an RTX 30-series card. Start ComfyUI without
--use-sage-attentionfor this model. - Out of memory, or a crawl, at 121 frames on 12 GB
- The VAE decode. Use VAE Decode (Tiled), use fewer frames, and on Windows set “Prefer No Sysmem Fallback” in the NVIDIA settings.
Expected tensor to have size 98 at dimension 1, but got size 64- A super-resolution file sits in the main model loader. Load a t2v or i2v model there; the SR model belongs in the upscale stage.
- Node
HunyuanVideo15ImageToVideois missing - ComfyUI is too old. Update it; the Desktop app got the node later than the portable build.
- Out of memory while loading the 720p model
- Set
weight_dtypein Load Diffusion Model tofp8_e4m3fn, or use a 480p fp8 file. - A CFG-distilled model looks unfinished
- Tencent says the CFG-distilled models must use 50 steps. 20 is too few for them.
- Faded, low-contrast video with the 4-step LoRA
- A known trade-off of that LoRA. Use more steps without it, or the 480p step-distilled model for image to video.
Is there a 720p version of the text-to-video lightx2v LoRA? HunyuanVideo needs a lot more prompt micromanaging than Wan 2.2.
Where is the HunyuanVideo 1.5 workflow for ComfyUI, and why does Desktop not have the node yet?
Questions.
Can I use HunyuanVideo 1.5 in the EU or the UK?
Not under its licence. The Tencent Hunyuan Community License states that it does not apply in the European Union, the United Kingdom or South Korea, so it grants no rights to use the model there. Elsewhere, commercial use is allowed below 100 million monthly users.
How much VRAM does HunyuanVideo 1.5 need?
Tencent’s minimum is 14 GB with offloading. In ComfyUI, the 480p fp8 or GGUF models run on 12 GB cards like the RTX 3060, with tiled VAE decoding at 121 frames. The 720p fp16 models are happiest on 24 GB.
Should I use the 480p or the 720p model?
480p on 12 to 16 GB and for quick tries: it has fp8 and GGUF files at 5 to 9 GB. 720p on 24 GB or more when you want the sharper picture and can wait. Each model is trained for its own resolution.
What shift should I use for HunyuanVideo 1.5?
Tencent uses shift 5 for the 480p models, 9 for 720p text to video, 7 for 720p image to video and the 480p step-distilled model, and 2 for the upscalers. Comfy’s templates use 7 throughout.
How do I get 1080p video from HunyuanVideo 1.5?
With the 1080p super-resolution model and its latent upsampler, as a second pass of 8 steps at CFG 1 and shift 2. Comfy’s templates include that stage, switched off. Use tiled VAE decoding, since 1080p decoding can run out of memory even on 32 GB.
Why is my HunyuanVideo 1.5 video black?
Usually SageAttention on an RTX 30-series card. Start ComfyUI without the --use-sage-attention flag for this model. One user also got black output from a lightx2v CFG-distilled fp8 file.
Does HunyuanVideo 1.5 run on a Mac?
There is no verified report yet. fp8 files don’t load on Apple GPUs, so a Mac needs the fp16 model and the bf16 encoder, about 36 GB together with the VAE, which points to 48 GB of memory or more.
Sources: HunyuanVideo 1.5 model card, licence, Comfy-Org repackage, Comfy-Org workflow templates, generation times thread [1], repackage test thread [2], step-distilled thread [3], ComfyUI issue #10888 [4], 12 GB thread, #10855, ComfyUI issue #10823, SR model thread.