Video model · beta

How to run MiniMax H3 locally.

MiniMax H3 runs in ComfyUI 0.30 or newer with four files: a 21.0 GB model, a Qwen3-VL 32B text encoder (15.7 GB), and a video and an audio VAE, and it makes video with sound in one pass. It is built for 24 GB cards with plenty of system RAM; 16 GB works at small sizes. Its licence excludes the EU, the UK, South Korea and the US unless you apply.

Updated 29 Sep 202611 min read

Maker
MiniMax
Released
Jul 2026Open weights since late July
Licence
MiniMax H3 CommunityNot licensed in the EU, UK, South Korea or US without an application
Memory
24 GB and up16 GB at small sizes, with lots of RAM

What it makes.

H3 is MiniMax’s open video model, on Hugging Face since late July 2026. It generates the picture and a 32 kHz stereo soundtrack together: voices, sound effects and music, all from one prompt.

  • Length: 5 to 15 seconds at 24 fps. The model card says 4 to 15. Frame counts snap to steps of 17, plus 5: 124 frames is five seconds.
  • Size: 768p, which is 1344 × 768 at 16:9. The 2K upscale MiniMax shows off runs only on their API.
  • Two checkpoints: FL2VA does text to video, or animates from a first and last frame. Ref2VA takes references: up to nine images, three videos and three audio clips.
  • No CFG, no negative prompt. The released weights are guidance-distilled, so one pass per step.

Under the hood it is a 33B transformer, but about 13B of that sits in branches whose outputs can be computed ahead of time. The pruned files leave those out, which is why they are 21 GB and not 34.

Files you need.

All four from Comfy-Org’s repackage. These are the ones Comfy’s template uses.

  • Model minimax_h3_fl2va_pruned_int8_convrot.safetensors ComfyUI/models/diffusion_models/
    21.0 GB Download
  • Text encoder qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors ComfyUI/models/text_encoders/ doesn’t need a Blackwell GPU
    15.7 GB Download
  • Video VAE minimax_h3_video_vae_int8_convrot.safetensors ComfyUI/models/vae/ fp16: 5.2 GB
    2.8 GB Download
  • Audio VAE minimax_h3_audio_vae_fp32.safetensors ComfyUI/models/vae/
    0.6 GB Download
  • Turbo LoRA minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors ComfyUI/models/loras/ optional, 8 steps. A 4-step 768p LoRA is in the Comfy-Org repo
    2.0 GB Download
ComfyUI/models
models/
├── diffusion_models/
│   └── minimax_h3_fl2va_pruned_int8_convrot.safetensors
├── loras/
│   └── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
├── text_encoders/
│   └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
└── vae/
    ├── minimax_h3_video_vae_int8_convrot.safetensors
    └── minimax_h3_audio_vae_fp32.safetensors

Other builds

Comfy-Org’s readme says to prefer int8_convrot if your PyTorch is built for CUDA 13.0 (cu130), and fp8 only if int8 won’t run. Ref2VA comes in the same sizes.

BuildModelText encoder
int8_convrot21.0 GB27.1 GB
nvfp4_awqnone15.7 GB
fp8_scaled21.0 GBnone
w6a816.0 GBnone
bf1640.2 GB51.5 GB
Unpruned int834.0 GBnone
Unpruned bf1666.3 GBnone

Model sizes are the pruned FL2VA files unless the row says unpruned.

What fits your computer.

H3 streams its weights between system RAM and the GPU, so RAM matters as much as VRAM. The diffusers docs expect around 75 GB of host RAM at int8 on a 12 to 16 GB card.

  • 8 GBBarely

    No 8 GB card report. One 16 GB owner capped ComfyUI at 8 GB and still ran 640 × 384, streaming about 270 GB from disk per run.

  • 12 GBTight

    Small canvases like 960 × 544, with the weights in system RAM. An RTX 4070 hung at “Model Initializing”, a known intermittent issue.

  • 16 GBOffloads

    RTX 5070 Ti with 32 GB of RAM: 640 × 384, 226 frames, with --disable-pinned-memory and --reserve-vram 1.5.

  • 24 GBFits

    RTX 4090 with 62 GB of RAM: 1344 × 768, five seconds, with SageAttention. VRAM peaked at 15.7 GB.

  • 32 GBFits

    RTX 5090. One owner measured Turbo at 8 steps 2.2 times faster than 20 steps; clips of 345 frames and more stalled.

  • Mac 48 GB+Use h3.c

    No verified ComfyUI run on a Mac. antirez’s h3.c runs H3 on Metal, peaking at 27.8 GB (int8) to 39.1 GB (bf16) at 512 × 512.

Set it up.

  1. Update ComfyUI

    H3 needs ComfyUI 0.30.0 or later. Desktop follows stable releases, so it can lag behind. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the files

    Model, text encoder and both VAEs from the list above, about 40 GB. Add a Turbo LoRA if you want 8-step clips.

  3. Put them in their folders

    Model in diffusion_models, encoder in text_encoders, both VAEs in vae, LoRA in loras. Restart ComfyUI.

  4. Open the template

    In the template browser, pick MiniMax H3: Text to Video, MiniMax H3: Image to Video or MiniMax H3: Reference to Video. The image template does text to video too when no picture is connected.

  5. Check the settings

    Load CLIP uses type minimax. There are two VAE loaders, one for video and one for audio. Set the duration in seconds; a math node turns it into a valid frame count. Keep the Resolution Selector’s multiple at 32.

  6. Write a structured prompt

    Describe the shots, camera moves and sound in one block, in the labelled sections from MiniMax’s prompting guide. Freeform prompts follow poorly. MiniMax’s own prompt rewriter, which the card calls critical to quality, is only on their API.

Settings that work.

Steps
20
CFG
noneBasicGuider, no negative
Sampler
res_multistep
Scheduler
simple
Shift
12ComfyUI’s default for H3
Size
1344 × 768768 px short edge
Length
5 to 15 smodel card: 4 to 15 s
Frames
124 for 5 ssteps of 17, plus 5

With a Turbo LoRA

8-step LoRA
8 stepsstrength 1.0
4-step LoRA
4 stepsmade for 768p
CFG
none
Sampler
res_multistep

Match the steps to the LoRA. In one RTX 4090 test, 4 steps looked mushy, 8 were coherent and usable, and 12 was the tester’s pick for everyday use.

Megapixels16:9 sizeUse
0.4864 × 480the image template’s default
0.5960 × 54412 to 16 GB cards
0.71152 × 640a middle step
0.981344 × 768native 768p

Sizes from Comfy’s template note, at a multiple of 32. Avoid 1.0 MP: 1376 × 768 goes over H3’s pixel cap.

How fast.

GPUSetupVideoTime
RTX 4090Turbo LoRA, 4 steps1344 × 768, 5 s275 s[1]
RTX 4090Turbo LoRA, 8 steps1344 × 768, 5 s420 s[1]
RTX 4090Turbo LoRA, 12 steps1344 × 768, 5 s605 s[1]
RTX 4090Base, 20 steps1344 × 768, 5 s990 s[1]
RTX 5070 Ti 16 GBRef2VA, 20 steps, EasyCache640 × 384, 226 frames190 s[2]
M5 Maxh3.c, 20 steps512 × 512, 22 frames16.7 s[3]

The RTX 4090 ran ComfyUI 0.30.0 with SageAttention and 62 GB of RAM. The M5 Max figure is from h3.c, a separate Metal program, not ComfyUI.

When it goes wrong.

It hangs at “Model Initializing...”
A known, intermittent issue with dynamic VRAM. Remove MultiGPU and other custom nodes that touch model loading, and try again.
CUDA out of memory on 24 GB
Keep dynamic VRAM on, which is the default, and use a smaller canvas or fewer frames.
Windows crashes with an access violation in the video VAE decode
Start ComfyUI with --disable-async-offload --disable-pinned-memory.
The PC swaps heavily with 32 GB of RAM
Add --disable-pinned-memory.
to() received an invalid combination of arguments - got (NestedTensor…)
VAE Decode (Tiled) doesn’t work with H3. Use the plain VAE Decode.
shape '[25600, 2560]' is invalid for input of size 3930620
Load CLIP failing on the encoder, most likely an older ComfyUI without nvfp4 support. Update, and check the type is minimax.
Static or popping in the audio
Update ComfyUI. For one user, a fresh portable 0.30.0 fixed it.
Garbled speech in the first moment
It depends on the seed. One user fixed it by writing dialogue in plain quotes instead of the <d> tags.
Mushy video with the Turbo LoRA
4 steps is too few for most clips. Use the 8-step LoRA at 8 steps.
The GGUF file won’t load
ComfyUI-GGUF doesn’t support H3 yet. Use the int8 safetensors files.

Can this run on my 16 GB PC? The answer in the thread: ComfyUI 0.30 can do it.

Hugging Face, MiniMax-H3

Prompt adherence really changes with resolution: small sizes follow the shot list, 768p much less.

Hugging Face, MiniMax-H3

Questions.

Can I use MiniMax H3 in the EU, the UK or the US?

Not under the standard licence. The MiniMax H3 Community License excludes the European Union, the United Kingdom, South Korea and the United States. People there can apply for a licence at platform.minimax.io/h3-license.

How much VRAM does MiniMax H3 need?

It is built for 24 GB cards with plenty of system RAM; an RTX 4090 with 62 GB of RAM runs it at 1344 × 768. 16 GB cards work at small sizes like 640 × 384, with the weights streaming from RAM and disk.

Do MiniMax H3 GGUF files work in ComfyUI?

Not yet. The H3 GGUFs on Hugging Face are made for stable-diffusion.cpp, and ComfyUI-GGUF can’t load them until pending changes land. Use the int8_convrot safetensors files.

How long can MiniMax H3 videos be?

5 to 15 seconds at 24 fps, per the diffusers docs; the model card says 4 to 15. Frame counts snap to steps of 17 plus 5, so five seconds is 124 frames.

Why is there no CFG or negative prompt for MiniMax H3?

The released weights are guidance-distilled. ComfyUI’s template uses a BasicGuider with no CFG, and there is no negative prompt to write.

Which MiniMax H3 Turbo LoRA should I use?

The 8-step FL2V LoRA at strength 1.0 with 8 steps is the safe choice. The 4-step LoRA is made for 768p and looked mushy in one RTX 4090 test. Ref2VA has its own 4-step LoRA.

Does MiniMax H3 run on a Mac?

There is no verified ComfyUI run on a Mac. antirez’s h3.c runs H3 natively on Apple Silicon with Metal, using about 28 to 39 GB at 512 × 512.

Sources: MiniMax H3 model card, licence, licence Q&A thread, ComfyUI MiniMax H3 tutorial, diffusers H3 docs, Comfy-Org repackage, RTX 4090 test [1], 16 GB notes [2], RTX 5090 notes, h3.c [3], ComfyUI issue #15628, #15337, #15274, #15614.

HEISS UI

Video with sound.

HEISS UI runs MiniMax H3 on the ComfyUI you already have, in the Video tab (beta). Write a prompt and get picture and sound in one file.

  • Picture and sound, one file. The finished clip lands in your gallery with its audio.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • Failures that explain themselves. When a run runs out of memory, it says so and offers the fix.

Image to video runs as your own workflow. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.