Hardware guide

What runs on an RTX 3090 or 4090.

24 GB runs nearly every model in ComfyUI from the card: Qwen-Image in fp8, HiDream, Krea 2, Wan 2.2 14B one half at a time, and MiniMax H3 in its pruned int8 form. Flux.2 Dev and full-precision Krea 2 still need system RAM. The 4090 is about twice as fast as the 3090 for the same work.

Updated 29 Sep 20265 min read

Memory
24 GBGDDR6X
Bandwidth
936 to 1008 GB/s3090 to 4090
fp8 compute
4090 onlyAda; the 3090 is Ampere
System RAM
64 GBfor Flux.2 Dev

Short answer

24 GB runs nearly every model HEISS UI supports from the card: Qwen-Image in fp8, HiDream, Krea 2, SD 3.5 Large, MiniMax H3 in its pruned form, and Wan 2.2 14B one half at a time. Flux.2 Dev and the full-precision Krea 2 still reach into system RAM.

Between the two, the 4090 does the same work in about half the time: 11 against 26 seconds for Flux.1 Dev.

RTX 3090 and 4090.

RTX 3090 / 3090 TiRTX 4090
Bandwidth936 / 1008 GB/s1008 GB/s
Computes fp8No: fp8 files save memory, not timeYes
SDXL, 1024 px, 20 steps6.2 s[1]2.3 to 3.6 s[1]
Flux.1 Dev fp8, 20 steps26 s[2]11.3 s[2]with SageAttention
Power350 W450 W

5.4% of Steam users have 24 GB, August 2026. The RTX 5090 Laptop GPU also has 24 GB, at much lower power: see laptops.

What fits in 24 GB.

Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.

Set up ComfyUI.

  1. Install or update ComfyUI

    ComfyUI Desktop or the portable NVIDIA build with a current driver.

  2. Add nothing, mostly

    Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).

    That changes the old advice. --lowvram now does nothing while Dynamic VRAM is on, and --novram, --highvram and --gpu-only switch it off. --normalvram is gone, and an old shortcut that still passes it stops ComfyUI with unrecognized arguments.

    On 24 GB there’s no memory flag to add. --use-sage-attention made the fastest 4090 result in the Flux thread above, but it needs the SageAttention package installed in ComfyUI’s Python first.

  3. 64 GB of RAM for the biggest models

    Flux.2 Dev in fp8 (35.5 GB, plus an 18 GB text encoder) and Krea 2 at full precision go past 24 GB. With 64 GB of system RAM they offload smoothly; with 32 GB they lean on the pagefile.

How long it takes.

CardModel and settingsTime
RTX 4090SDXL, 1024 px, 20 steps2.3 to 3.6 s[1]
RTX 3090SDXL, 1024 px, 20 steps6.2 s[1]
RTX 4090Flux.1 Dev fp8, 20 steps, SageAttention11.3 s[2]
RTX 3090Flux.1 Dev fp8, 20 steps26 s[2]
RTX 4090DQwen-Image fp8, 1328 px, 20 steps34 s with the 8-step lightx2v LoRA. Warm runs.71 s[3]
RTX 4090Wan 2.2 14B GGUF Q4_K_M, 832 × 480, 81 frames, 20 stepsComfyUI, 14.9 GB peak.516 s[4]
RTX 4090Wan 2.2 5B, 1280 × 704, 121 frames, 50 stepsWan’s own script, not ComfyUI.667 s[4]

The Wan 2.2 times use 20 and 50 steps. With the 4-step lightx2v LoRAs the 14B takes a fraction of that; there’s no clean published 24 GB number for it yet.

Flux.2 Dev on 24 GB.

Flux.2 Dev is a 32B model. NVIDIA’s fp8 version cut its memory by about 40% and ComfyUI streams the rest from system RAM [5], but that’s still 35.5 GB of model plus an 18 GB Mistral text encoder (12.3 GB in fp4). Two ways onto 24 GB:

  • The fp8 file with 64 GB of RAM. ComfyUI offloads what doesn’t fit; each step is slower than on a 32 GB card.
  • A GGUF Q4_K_M (20.1 GB) fits the card through ComfyUI-GGUF, at some quality cost.

Video at 720p.

24 GB is where 720p video gets comfortable. Wan 2.2 5B and HunyuanVideo 1.5 run at 720p from the card. Wan 2.2 14B in fp8 runs one 14.3 GB half at a time; at 1280 × 720 and 81 frames, expect minutes per clip even with the 4-step LoRAs.

When it goes wrong.

torch.OutOfMemoryError: CUDA out of memory
The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)
Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
It suddenly got much slower
The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram
An old launch flag. Remove it from the shortcut or the .bat file. Current ComfyUI doesn’t need memory flags on NVIDIA.
The first image takes minutes, the next ones don’t
The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.

Questions.

Used RTX 3090 or new RTX 5060 Ti 16 GB?

The 3090 fits more (24 GB) and moves memory twice as fast; the 5060 Ti computes fp8 and nvfp4, is new and draws half the power. They ran Flux.1 Dev fp8 at about the same speed. For Qwen-Image in fp8, HiDream and 720p video, the 3090.

RTX 3090 or RTX 4090 for ComfyUI?

Same 24 GB, so the same models fit. The 4090 computes fp8 and is about twice as fast: 11 against 26 seconds for Flux.1 Dev fp8 in the same thread.

4090 or 5090?

The 5090 adds 8 GB and 78% more bandwidth, and computes nvfp4. It runs Krea 2 at full precision and Flux.2 Dev with less offloading. If your models fit in 24 GB, the 4090 is plenty.

Is 64 GB of RAM worth it with a 24 GB card?

If you want Flux.2 Dev in fp8, Krea 2 at full precision or long Wan 2.2 14B clips, yes. For everything else, 32 GB is enough.

Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [3] Qwen-Image in ComfyUI, Comfy docs, [4] Wan 2.2 locally on an RTX 4090, ComputingForGeeks, [5] FLUX.2 on RTX GPUs, NVIDIA blog, [6] MiniMax H3 in ComfyUI, Comfy blog, [7] Dynamic VRAM in ComfyUI, Comfy blog, [8] Steam Hardware Survey, August 2026.

HEISS UI

Made for a 24 GB card.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. It reads your card and does the fitting for you.

  • The model that fits your 24 GB is marked. Pick a first model and the size that suits this card is already chosen. One tap downloads it with everything it needs.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.