Short answer
24 GB runs nearly every model HEISS UI supports from the card: Qwen-Image in fp8, HiDream, Krea 2, SD 3.5 Large, MiniMax H3 in its pruned form, and Wan 2.2 14B one half at a time. Flux.2 Dev and the full-precision Krea 2 still reach into system RAM.
Between the two, the 4090 does the same work in about half the time: 11 against 26 seconds for Flux.1 Dev.
RTX 3090 and 4090.
| RTX 3090 / 3090 Ti | RTX 4090 | |
|---|---|---|
| Bandwidth | 936 / 1008 GB/s | 1008 GB/s |
| Computes fp8 | No: fp8 files save memory, not time | Yes |
| SDXL, 1024 px, 20 steps | 6.2 s[1] | 2.3 to 3.6 s[1] |
| Flux.1 Dev fp8, 20 steps | 26 s[2] | 11.3 s[2]with SageAttention |
| Power | 350 W | 450 W |
5.4% of Steam users have 24 GB, August 2026. The RTX 5090 Laptop GPU also has 24 GB, at much lower power: see laptops.
What fits in 24 GB.
Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.
- Qwen-Imagefp8 · 20.4 GBFits
Comfy measured 86% of a 24 GB card.
- HiDream I1fp8 · 17.1 GBFits
Its four text encoders run first.
- Krea 2 Turbo and Rawfp8 · 13.1 GBFits
bf16 (26.3 GB) offloads a little.
- Flux.1 Devfp8 12 GB · bf16 23.8 GBFits
bf16 is tight.
- SD 3.5 Large, ERNIE-Image, Ideogram 415 to 19 GBFits
In fp8. Ideogram 4 runs its two models side by side.
- Flux.2 Klein 9B, Z-Image, Qwen-Image 2.1, Chromafull precisionFits
bf16 files, no compromises.
- MiniMax H3pruned int8 · 21 GBTight
The text encoder leaves the card after encoding.
- Wan 2.2 14Bfp8 · 2 × 14.3 GBFits
One half at a time. fp16 halves (28.6 GB) offload.
- Wan 2.2 5B, HunyuanVideo 1.5fp16 · 10 and 16.7 GBFits
720p.
- Flux.2 Devfp8 35.5 GB · GGUF Q4_K_M 20.1 GBOffloads
The GGUF fits; fp8 needs 64 GB of RAM.
Set up ComfyUI.
Install or update ComfyUI
ComfyUI Desktop or the portable NVIDIA build with a current driver.
Add nothing, mostly
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.On 24 GB there’s no memory flag to add.
--use-sage-attentionmade the fastest 4090 result in the Flux thread above, but it needs the SageAttention package installed in ComfyUI’s Python first.64 GB of RAM for the biggest models
Flux.2 Dev in fp8 (35.5 GB, plus an 18 GB text encoder) and Krea 2 at full precision go past 24 GB. With 64 GB of system RAM they offload smoothly; with 32 GB they lean on the pagefile.
How long it takes.
| Card | Model and settings | Time |
|---|---|---|
| RTX 4090 | SDXL, 1024 px, 20 steps | 2.3 to 3.6 s[1] |
| RTX 3090 | SDXL, 1024 px, 20 steps | 6.2 s[1] |
| RTX 4090 | Flux.1 Dev fp8, 20 steps, SageAttention | 11.3 s[2] |
| RTX 3090 | Flux.1 Dev fp8, 20 steps | 26 s[2] |
| RTX 4090D | Qwen-Image fp8, 1328 px, 20 steps34 s with the 8-step lightx2v LoRA. Warm runs. | 71 s[3] |
| RTX 4090 | Wan 2.2 14B GGUF Q4_K_M, 832 × 480, 81 frames, 20 stepsComfyUI, 14.9 GB peak. | 516 s[4] |
| RTX 4090 | Wan 2.2 5B, 1280 × 704, 121 frames, 50 stepsWan’s own script, not ComfyUI. | 667 s[4] |
The Wan 2.2 times use 20 and 50 steps. With the 4-step lightx2v LoRAs the 14B takes a fraction of that; there’s no clean published 24 GB number for it yet.
Flux.2 Dev on 24 GB.
Flux.2 Dev is a 32B model. NVIDIA’s fp8 version cut its memory by about 40% and ComfyUI streams the rest from system RAM [5], but that’s still 35.5 GB of model plus an 18 GB Mistral text encoder (12.3 GB in fp4). Two ways onto 24 GB:
- The fp8 file with 64 GB of RAM. ComfyUI offloads what doesn’t fit; each step is slower than on a 32 GB card.
- A GGUF Q4_K_M (20.1 GB) fits the card through ComfyUI-GGUF, at some quality cost.
Video at 720p.
24 GB is where 720p video gets comfortable. Wan 2.2 5B and HunyuanVideo 1.5 run at 720p from the card. Wan 2.2 14B in fp8 runs one 14.3 GB half at a time; at 1280 × 720 and 81 frames, expect minutes per clip even with the 4-step LoRAs.
When it goes wrong.
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
Questions.
Used RTX 3090 or new RTX 5060 Ti 16 GB?
The 3090 fits more (24 GB) and moves memory twice as fast; the 5060 Ti computes fp8 and nvfp4, is new and draws half the power. They ran Flux.1 Dev fp8 at about the same speed. For Qwen-Image in fp8, HiDream and 720p video, the 3090.
RTX 3090 or RTX 4090 for ComfyUI?
Same 24 GB, so the same models fit. The 4090 computes fp8 and is about twice as fast: 11 against 26 seconds for Flux.1 Dev fp8 in the same thread.
4090 or 5090?
The 5090 adds 8 GB and 78% more bandwidth, and computes nvfp4. It runs Krea 2 at full precision and Flux.2 Dev with less offloading. If your models fit in 24 GB, the 4090 is plenty.
Is 64 GB of RAM worth it with a 24 GB card?
If you want Flux.2 Dev in fp8, Krea 2 at full precision or long Wan 2.2 14B clips, yes. For everything else, 32 GB is enough.
Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [3] Qwen-Image in ComfyUI, Comfy docs, [4] Wan 2.2 locally on an RTX 4090, ComputingForGeeks, [5] FLUX.2 on RTX GPUs, NVIDIA blog, [6] MiniMax H3 in ComfyUI, Comfy blog, [7] Dynamic VRAM in ComfyUI, Comfy blog, [8] Steam Hardware Survey, August 2026.