Video model · beta

How to run HunyuanVideo 1.5 locally.

HunyuanVideo 1.5 runs in ComfyUI with an 8.3B model, two text encoders (Qwen2.5-VL 7B and ByT5), its own VAE and, for image to video, a SigLIP vision file. Tencent’s minimum is 14 GB of VRAM with offloading; the 480p models in fp8 (8.3 GB) or GGUF run on 12 GB cards. Its licence doesn’t cover the EU, the UK or South Korea.

Updated 29 Sep 202611 min read

Maker
Tencent Hunyuan
Released
Nov 202520 November; 480p I2V step-distilled 5 December
Licence
Tencent Hunyuan CommunityNot licensed in the EU, UK or South Korea
Memory
12 GB and up480p in fp8 or GGUF. Tencent: 14 GB

480p or 720p.

Tencent released HunyuanVideo 1.5 on 20 November 2025: an 8.3B model that makes five-second clips at 24 fps, from text or from a picture. Every model is its own file, trained for one resolution and one task.

  • 480p text to video and image to video. The ones to start with on 12 to 16 GB.
  • 720p text to video and image to video. What Comfy’s templates use. Sharper, and much slower.
  • CFG-distilled versions of both run at CFG 1, one pass per step instead of two. Tencent still asks for 50 steps.
  • 480p I2V step-distilled, added on 5 December, runs in 8 or 12 steps. Tencent says an RTX 4090 makes a clip with it in about 75 seconds.
  • Super-resolution models take a finished clip up to 720p or 1080p. More on that under settings.

Image to video needs one extra file, a SigLIP vision encoder. Text in the video, like a sign or a title, comes from the second text encoder, ByT5.

Files you need.

All from Comfy-Org’s repackage. The encoders and the VAE are shared by every model.

Shared parts

  • Text encoder qwen_2.5_vl_7b_fp8_scaled.safetensors ComfyUI/models/text_encoders/ or qwen_2.5_vl_7b.safetensors, bf16, 16.6 GB
    9.4 GB Download
  • Glyph encoder byt5_small_glyphxl_fp16.safetensors ComfyUI/models/text_encoders/
    0.4 GB Download
  • VAE hunyuanvideo15_vae_fp16.safetensors ComfyUI/models/vae/
    2.5 GB Download

Text to video, pick one

  • 480p, CFG-distilled hunyuanvideo1.5_480p_t2v_cfg_distilled_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ fp16: 16.7 GB
    8.3 GB Download
  • 480p hunyuanvideo1.5_480p_t2v_fp16.safetensors ComfyUI/models/diffusion_models/
    16.7 GB Download
  • 720p hunyuanvideo1.5_720p_t2v_fp16.safetensors ComfyUI/models/diffusion_models/ the Comfy template’s model
    16.7 GB Download

Image to video, pick one, plus the vision file

  • 480p, step-distilled hunyuanvideo1.5_480p_i2v_step_distilled_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ fp16: 16.7 GB
    8.3 GB Download
  • 720p hunyuanvideo1.5_720p_i2v_fp16.safetensors ComfyUI/models/diffusion_models/ the Comfy template’s model. CFG-distilled fp8: 8.3 GB
    16.7 GB Download
  • Vision encoder sigclip_vision_patch14_384.safetensors ComfyUI/models/clip_vision/
    0.9 GB Download

Optional

  • 1080p upscaler hunyuanvideo1.5_1080p_sr_distilled_fp8_scaled.safetensors ComfyUI/models/diffusion_models/ fp16: 16.7 GB. Needs the latent upsampler below
    8.3 GB Download
  • Latent upsampler hunyuanvideo15_latent_upsampler_1080p.safetensors ComfyUI/models/latent_upscale_models/
    0.2 GB Download
  • 4-step LoRA hunyuanvideo1.5_t2v_480p_lightx2v_4step_lora_rank_32_bf16.safetensors ComfyUI/models/loras/ for 480p text to video
    0.3 GB Download
ComfyUI/models
models/
├── clip_vision/
│   └── sigclip_vision_patch14_384.safetensors   (image to video)
├── diffusion_models/
│   └── hunyuanvideo1.5_480p_t2v_cfg_distilled_fp8_scaled.safetensors
├── text_encoders/
│   ├── qwen_2.5_vl_7b_fp8_scaled.safetensors
│   └── byt5_small_glyphxl_fp16.safetensors
└── vae/
    └── hunyuanvideo15_vae_fp16.safetensors

The plain 720p models have no fp8 file. If they run out of memory, Comfy’s template note says to set weight_dtype in Load Diffusion Model to fp8_e4m3fn.

Smaller files: GGUF

jayn7’s GGUF builds load through ComfyUI-GGUF with Unet Loader (GGUF). The same uploader has 720p text-to-video and image-to-video GGUFs at similar sizes.

ModelQ4_K_MQ5_K_MQ6_KQ8_0
480p T2V5.1 GB6.1 GB7.0 GB9.0 GB

The repo has the plain and the CFG-distilled 480p model. Q8_0 is close to fp16.

What fits your computer.

The model file is only part of it. At 121 frames the VAE decode at the end is where 12 GB cards run out, and VAE Decode (Tiled) is the fix.

  • 6 to 8 GBUntested

    No ComfyUI report found at this size. The 480p Q4_K_M GGUF is 5.1 GB. WanGP lists HunyuanVideo 1.5 from 6 GB.

  • 12 GBOffloads

    480p in fp8 or GGUF. An RTX 3060 ran 480p at 121 frames. An RTX 4070 was fine at 81 frames and spilled over at 121, in the VAE decode.

  • 16 GB480p fits

    480p fp8. An RTX 4060 Ti 16 GB ran the CFG-distilled 480p model at 640 × 480, 81 frames.

  • 24 GBFits

    720p fp16. The default 720p text-to-video template took about 20 minutes on an RTX 4090.

  • 32 GBFits

    RTX 5090: 848 × 480 at 121 frames in 284 s with the 720p model.

  • Mac 48 GB+Untested

    No verified Mac report. fp8 files don’t load on Apple GPUs; the fp16 model, bf16 encoder and VAE add up to 35.8 GB.

Set it up.

  1. Update ComfyUI

    HunyuanVideo 1.5 needs core nodes from late November 2025, such as HunyuanVideo15ImageToVideo. ComfyUI Desktop was behind at first, so check that it has updated. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the files

    The two text encoders, the VAE and one model from the lists above. For image to video, add the SigLIP vision file.

  3. Put them in their folders

    Model in diffusion_models, both encoders in text_encoders, VAE in vae, SigLIP in clip_vision. Restart ComfyUI.

  4. Open the template

    In the template browser, pick Hunyuan Video 1.5 Text to Video or Hunyuan Video 1.5 Image to Video. Both are set up for the 720p models, with a 1080p upscale stage that starts switched off.

  5. Check the loaders

    DualCLIPLoader gets Qwen2.5-VL and ByT5 with type hunyuan_video_15. For image to video, Load CLIP Vision gets the SigLIP file. If you swapped in a 480p model, set the size to 848 × 480 and the shift to 5.

  6. Write a prompt and run

    Be specific about subject, motion and camera. Tencent’s own code rewrites prompts with a large language model before generating; in ComfyUI you write the full prompt yourself.

Settings that work.

Tencent sets the shift per model, and it matters. Comfy’s templates use 20 steps instead of 50, because 50 “just take too long”, and shift 7 on everything.

480p, Tencent’s settings

Steps
50Comfy template: 20
CFG
61 for CFG-distilled
Sampler
euler
Scheduler
simple
Shift
5Comfy template: 7
Size
848 × 480
Frames
1215 s at 24 fps
Negative
emptyTencent’s default
ModelCFGShiftSteps
480p T2V and I2V6550
720p T2V6950
720p I2V6750
480p CFG-distilled1550
720p I2V CFG-distilled1750
480p I2V step-distilled178 or 12
Upscale to 720p126
Upscale to 1080p128

From Tencent’s model card, repeated in Comfy’s template notes. Comfy’s 720p templates use shift 7 for text to video too.

The 4-step lightx2v LoRA is the fast route for 480p text to video, at CFG 1. It also seems to work at 720p, but people found the results faded and low in contrast, especially in low light.

How fast.

GPUModelVideoTime
RTX 5090720p model848 × 480, 121 frames284 s[1]
RTX 5090720p model848 × 480, 245 frames498 s[1]
RTX 4090720p model, undervolted848 × 480, 121 frames297 s[1]
RTX 4090720p T2V template1280 × 720, 121 frames20 min[2]
RTX 4090480p I2V step-distilled480p75 s[3]
RTX 2080 Ti4 steps512 × 512, 49 frames80 s[4]

Tencent’s 75 s figure is “within 75 seconds” for the step-distilled model. On the 2080 Ti, the VAE alone took about 20 s, and Wan 2.2 in a rapid merge took 31 s for the same clip.

When it goes wrong.

The video is black
SageAttention on an RTX 30-series card. Start ComfyUI without --use-sage-attention for this model.
Out of memory, or a crawl, at 121 frames on 12 GB
The VAE decode. Use VAE Decode (Tiled), use fewer frames, and on Windows set “Prefer No Sysmem Fallback” in the NVIDIA settings.
Expected tensor to have size 98 at dimension 1, but got size 64
A super-resolution file sits in the main model loader. Load a t2v or i2v model there; the SR model belongs in the upscale stage.
Node HunyuanVideo15ImageToVideo is missing
ComfyUI is too old. Update it; the Desktop app got the node later than the portable build.
Out of memory while loading the 720p model
Set weight_dtype in Load Diffusion Model to fp8_e4m3fn, or use a 480p fp8 file.
A CFG-distilled model looks unfinished
Tencent says the CFG-distilled models must use 50 steps. 20 is too few for them.
Faded, low-contrast video with the 4-step LoRA
A known trade-off of that LoRA. Use more steps without it, or the 480p step-distilled model for image to video.

Is there a 720p version of the text-to-video lightx2v LoRA? HunyuanVideo needs a lot more prompt micromanaging than Wan 2.2.

Hugging Face, Comfy-Org HunyuanVideo 1.5 repackage

Where is the HunyuanVideo 1.5 workflow for ComfyUI, and why does Desktop not have the node yet?

GitHub, ComfyUI issue #10823

Questions.

Can I use HunyuanVideo 1.5 in the EU or the UK?

Not under its licence. The Tencent Hunyuan Community License states that it does not apply in the European Union, the United Kingdom or South Korea, so it grants no rights to use the model there. Elsewhere, commercial use is allowed below 100 million monthly users.

How much VRAM does HunyuanVideo 1.5 need?

Tencent’s minimum is 14 GB with offloading. In ComfyUI, the 480p fp8 or GGUF models run on 12 GB cards like the RTX 3060, with tiled VAE decoding at 121 frames. The 720p fp16 models are happiest on 24 GB.

Should I use the 480p or the 720p model?

480p on 12 to 16 GB and for quick tries: it has fp8 and GGUF files at 5 to 9 GB. 720p on 24 GB or more when you want the sharper picture and can wait. Each model is trained for its own resolution.

What shift should I use for HunyuanVideo 1.5?

Tencent uses shift 5 for the 480p models, 9 for 720p text to video, 7 for 720p image to video and the 480p step-distilled model, and 2 for the upscalers. Comfy’s templates use 7 throughout.

How do I get 1080p video from HunyuanVideo 1.5?

With the 1080p super-resolution model and its latent upsampler, as a second pass of 8 steps at CFG 1 and shift 2. Comfy’s templates include that stage, switched off. Use tiled VAE decoding, since 1080p decoding can run out of memory even on 32 GB.

Why is my HunyuanVideo 1.5 video black?

Usually SageAttention on an RTX 30-series card. Start ComfyUI without the --use-sage-attention flag for this model. One user also got black output from a lightx2v CFG-distilled fp8 file.

Does HunyuanVideo 1.5 run on a Mac?

There is no verified report yet. fp8 files don’t load on Apple GPUs, so a Mac needs the fp16 model and the bf16 encoder, about 36 GB together with the VAE, which points to 48 GB of memory or more.

Sources: HunyuanVideo 1.5 model card, licence, Comfy-Org repackage, Comfy-Org workflow templates, generation times thread [1], repackage test thread [2], step-distilled thread [3], ComfyUI issue #10888 [4], 12 GB thread, #10855, ComfyUI issue #10823, SR model thread.

HEISS UI

A picture in, video out.

HEISS UI runs HunyuanVideo 1.5 text and image to video on the ComfyUI you already have, in the Video tab (beta).

  • Image to video, wired for you. Pick the image-to-video file and the composer asks for a start picture. The rest is set up.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • Failures that explain themselves. When a run runs out of memory, it says so and offers the fix.

The 1080p upscale stage isn’t built in yet. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.