What it makes.
H3 is MiniMax’s open video model, on Hugging Face since late July 2026. It generates the picture and a 32 kHz stereo soundtrack together: voices, sound effects and music, all from one prompt.
- Length: 5 to 15 seconds at 24 fps. The model card says 4 to 15. Frame counts snap to steps of 17, plus 5: 124 frames is five seconds.
- Size: 768p, which is 1344 × 768 at 16:9. The 2K upscale MiniMax shows off runs only on their API.
- Two checkpoints: FL2VA does text to video, or animates from a first and last frame. Ref2VA takes references: up to nine images, three videos and three audio clips.
- No CFG, no negative prompt. The released weights are guidance-distilled, so one pass per step.
Under the hood it is a 33B transformer, but about 13B of that sits in branches whose outputs can be computed ahead of time. The pruned files leave those out, which is why they are 21 GB and not 34.
Files you need.
All four from Comfy-Org’s repackage. These are the ones Comfy’s template uses.
-
Model21.0 GB Download
minimax_h3_fl2va_pruned_int8_convrot.safetensorsComfyUI/models/diffusion_models/ -
Text encoder15.7 GB Download
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensorsComfyUI/models/text_encoders/ doesn’t need a Blackwell GPU -
Video VAE2.8 GB Download
minimax_h3_video_vae_int8_convrot.safetensorsComfyUI/models/vae/ fp16: 5.2 GB -
Audio VAE0.6 GB Download
minimax_h3_audio_vae_fp32.safetensorsComfyUI/models/vae/ -
Turbo LoRA2.0 GB Download
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensorsComfyUI/models/loras/ optional, 8 steps. A 4-step 768p LoRA is in the Comfy-Org repo
models/
├── diffusion_models/
│ └── minimax_h3_fl2va_pruned_int8_convrot.safetensors
├── loras/
│ └── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
├── text_encoders/
│ └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
└── vae/
├── minimax_h3_video_vae_int8_convrot.safetensors
└── minimax_h3_audio_vae_fp32.safetensors
Other builds
Comfy-Org’s readme says to prefer int8_convrot if your PyTorch is built for CUDA 13.0 (cu130), and fp8 only if int8 won’t run. Ref2VA comes in the same sizes.
| Build | Model | Text encoder |
|---|---|---|
| int8_convrot | 21.0 GB | 27.1 GB |
| nvfp4_awq | none | 15.7 GB |
| fp8_scaled | 21.0 GB | none |
| w6a8 | 16.0 GB | none |
| bf16 | 40.2 GB | 51.5 GB |
| Unpruned int8 | 34.0 GB | none |
| Unpruned bf16 | 66.3 GB | none |
Model sizes are the pruned FL2VA files unless the row says unpruned.
What fits your computer.
H3 streams its weights between system RAM and the GPU, so RAM matters as much as VRAM. The diffusers docs expect around 75 GB of host RAM at int8 on a 12 to 16 GB card.
- 8 GBBarely
No 8 GB card report. One 16 GB owner capped ComfyUI at 8 GB and still ran 640 × 384, streaming about 270 GB from disk per run.
- 12 GBTight
Small canvases like 960 × 544, with the weights in system RAM. An RTX 4070 hung at “Model Initializing”, a known intermittent issue.
- 16 GBOffloads
RTX 5070 Ti with 32 GB of RAM: 640 × 384, 226 frames, with
--disable-pinned-memoryand--reserve-vram 1.5. - 24 GBFits
RTX 4090 with 62 GB of RAM: 1344 × 768, five seconds, with SageAttention. VRAM peaked at 15.7 GB.
- 32 GBFits
RTX 5090. One owner measured Turbo at 8 steps 2.2 times faster than 20 steps; clips of 345 frames and more stalled.
- Mac 48 GB+Use h3.c
No verified ComfyUI run on a Mac. antirez’s h3.c runs H3 on Metal, peaking at 27.8 GB (int8) to 39.1 GB (bf16) at 512 × 512.
Set it up.
-
Update ComfyUI
H3 needs ComfyUI 0.30.0 or later. Desktop follows stable releases, so it can lag behind. In a manual install:
Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the files
Model, text encoder and both VAEs from the list above, about 40 GB. Add a Turbo LoRA if you want 8-step clips.
-
Put them in their folders
Model in
diffusion_models, encoder intext_encoders, both VAEs invae, LoRA inloras. Restart ComfyUI. -
Open the template
In the template browser, pick MiniMax H3: Text to Video, MiniMax H3: Image to Video or MiniMax H3: Reference to Video. The image template does text to video too when no picture is connected.
-
Check the settings
Load CLIP uses type
minimax. There are two VAE loaders, one for video and one for audio. Set the duration in seconds; a math node turns it into a valid frame count. Keep the Resolution Selector’s multiple at 32. -
Write a structured prompt
Describe the shots, camera moves and sound in one block, in the labelled sections from MiniMax’s prompting guide. Freeform prompts follow poorly. MiniMax’s own prompt rewriter, which the card calls critical to quality, is only on their API.
Settings that work.
- Steps
- 20
- CFG
- noneBasicGuider, no negative
- Sampler
- res_multistep
- Scheduler
- simple
- Shift
- 12ComfyUI’s default for H3
- Size
- 1344 × 768768 px short edge
- Length
- 5 to 15 smodel card: 4 to 15 s
- Frames
- 124 for 5 ssteps of 17, plus 5
With a Turbo LoRA
- 8-step LoRA
- 8 stepsstrength 1.0
- 4-step LoRA
- 4 stepsmade for 768p
- CFG
- none
- Sampler
- res_multistep
Match the steps to the LoRA. In one RTX 4090 test, 4 steps looked mushy, 8 were coherent and usable, and 12 was the tester’s pick for everyday use.
| Megapixels | 16:9 size | Use |
|---|---|---|
| 0.4 | 864 × 480 | the image template’s default |
| 0.5 | 960 × 544 | 12 to 16 GB cards |
| 0.7 | 1152 × 640 | a middle step |
| 0.98 | 1344 × 768 | native 768p |
Sizes from Comfy’s template note, at a multiple of 32. Avoid 1.0 MP: 1376 × 768 goes over H3’s pixel cap.
How fast.
| GPU | Setup | Video | Time |
|---|---|---|---|
| RTX 4090 | Turbo LoRA, 4 steps | 1344 × 768, 5 s | 275 s[1] |
| RTX 4090 | Turbo LoRA, 8 steps | 1344 × 768, 5 s | 420 s[1] |
| RTX 4090 | Turbo LoRA, 12 steps | 1344 × 768, 5 s | 605 s[1] |
| RTX 4090 | Base, 20 steps | 1344 × 768, 5 s | 990 s[1] |
| RTX 5070 Ti 16 GB | Ref2VA, 20 steps, EasyCache | 640 × 384, 226 frames | 190 s[2] |
| M5 Max | h3.c, 20 steps | 512 × 512, 22 frames | 16.7 s[3] |
The RTX 4090 ran ComfyUI 0.30.0 with SageAttention and 62 GB of RAM. The M5 Max figure is from h3.c, a separate Metal program, not ComfyUI.
When it goes wrong.
- It hangs at “Model Initializing...”
- A known, intermittent issue with dynamic VRAM. Remove MultiGPU and other custom nodes that touch model loading, and try again.
- CUDA out of memory on 24 GB
- Keep dynamic VRAM on, which is the default, and use a smaller canvas or fewer frames.
- Windows crashes with an access violation in the video VAE decode
- Start ComfyUI with
--disable-async-offload --disable-pinned-memory. - The PC swaps heavily with 32 GB of RAM
- Add
--disable-pinned-memory. to() received an invalid combination of arguments - got (NestedTensor…)- VAE Decode (Tiled) doesn’t work with H3. Use the plain VAE Decode.
shape '[25600, 2560]' is invalid for input of size 3930620- Load CLIP failing on the encoder, most likely an older ComfyUI without nvfp4 support. Update, and check the type is
minimax. - Static or popping in the audio
- Update ComfyUI. For one user, a fresh portable 0.30.0 fixed it.
- Garbled speech in the first moment
- It depends on the seed. One user fixed it by writing dialogue in plain quotes instead of the
<d>tags. - Mushy video with the Turbo LoRA
- 4 steps is too few for most clips. Use the 8-step LoRA at 8 steps.
- The GGUF file won’t load
- ComfyUI-GGUF doesn’t support H3 yet. Use the int8 safetensors files.
Can this run on my 16 GB PC? The answer in the thread: ComfyUI 0.30 can do it.
Prompt adherence really changes with resolution: small sizes follow the shot list, 768p much less.
Questions.
Can I use MiniMax H3 in the EU, the UK or the US?
Not under the standard licence. The MiniMax H3 Community License excludes the European Union, the United Kingdom, South Korea and the United States. People there can apply for a licence at platform.minimax.io/h3-license.
How much VRAM does MiniMax H3 need?
It is built for 24 GB cards with plenty of system RAM; an RTX 4090 with 62 GB of RAM runs it at 1344 × 768. 16 GB cards work at small sizes like 640 × 384, with the weights streaming from RAM and disk.
Do MiniMax H3 GGUF files work in ComfyUI?
Not yet. The H3 GGUFs on Hugging Face are made for stable-diffusion.cpp, and ComfyUI-GGUF can’t load them until pending changes land. Use the int8_convrot safetensors files.
How long can MiniMax H3 videos be?
5 to 15 seconds at 24 fps, per the diffusers docs; the model card says 4 to 15. Frame counts snap to steps of 17 plus 5, so five seconds is 124 frames.
Why is there no CFG or negative prompt for MiniMax H3?
The released weights are guidance-distilled. ComfyUI’s template uses a BasicGuider with no CFG, and there is no negative prompt to write.
Which MiniMax H3 Turbo LoRA should I use?
The 8-step FL2V LoRA at strength 1.0 with 8 steps is the safe choice. The 4-step LoRA is made for 768p and looked mushy in one RTX 4090 test. Ref2VA has its own 4-step LoRA.
Does MiniMax H3 run on a Mac?
There is no verified ComfyUI run on a Mac. antirez’s h3.c runs H3 natively on Apple Silicon with Metal, using about 28 to 39 GB at 512 × 512.
Sources: MiniMax H3 model card, licence, licence Q&A thread, ComfyUI MiniMax H3 tutorial, diffusers H3 docs, Comfy-Org repackage, RTX 4090 test [1], 16 GB notes [2], RTX 5090 notes, h3.c [3], ComfyUI issue #15628, #15337, #15274, #15614.