Short answer
16 GB runs every image model HEISS UI supports up to about 15 GB from the card: Krea 2, Flux.1 Dev, Flux.2 Klein 9B, Z-Image, Qwen-Image 2.1, SD 3.5 Large, plus Wan 2.2 5B and HunyuanVideo 1.5 at 720p. Qwen-Image in fp8, HiDream, Wan 2.2 14B and Flux.2 Dev still use system RAM.
On a 5070 Ti or 5080, the nvfp4 files leave room to spare.
The 16 GB cards.
| Card | Bandwidth | Computes natively | Flux.1 Dev fp8, 20 steps |
|---|---|---|---|
| RTX 5080 | 960 GB/s | fp8, nvfp4 | 16.7 s[1]6.7 s tuned |
| RTX 5070 Ti | 896 GB/s | fp8, nvfp4 | No published number |
| RTX 4080 / Super | 717 / 736 GB/s | fp8 | SDXL: 6.5 s[2] |
| RTX 4070 Ti Super | 672 GB/s | fp8 | No published number |
| RTX 5060 Ti / 4060 Ti 16 GB | 448 / 288 GB/s | fp8 (and nvfp4 on the 5060 Ti) | Their own page |
AMD’s 16 GB cards (RX 9070 XT, 9070, 7800 XT, 7600 XT) fit the same models: see the AMD guide.
What fits in 16 GB.
Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.
- Krea 2 Turbo and Rawfp8 · 13.1 GBFits
Or nvfp4 (7.7 GB) on RTX 50. The bf16 files (26.3 GB) offload.
- Flux.1 Devfp8 · about 12 GBFits
16.7 s at 20 steps on a 5080, 6.7 s tuned.
- Flux.2 Klein 9Bfp8Fits
Comfy measures 19.6 GB at bf16; fp8 fits.
- Z-Image, Qwen-Image 2.1, Chroma, MageFlow, Ideogram 44 to 14 GBFits
Ideogram 4 in nvfp4 on RTX 50; its fp8 pair offloads a little.
- SD 3.5 Large, ERNIE-Image14.9 and 16.1 GBTight
One image at a time at 1024 px.
- Qwen-ImageGGUF Q5_K_M · 14.9 GBTight
fp8 (20.4 GB) offloads.
- HiDream I1fp8 · 17.1 GBOffloads
A little over, plus four text encoders.
- Wan 2.2 5Bfp16 · 10 GBFits
1280 × 704, 121 frames.
- HunyuanVideo 1.5fp8 · 8.3 GBFits
720p.
- Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads
Each half fits alone; the switch goes through RAM. 832 × 480 is easy, 1280 × 720 is slow.
- MiniMax H3w6a8 · 16 GBOffloads
With its nvfp4 text encoder (15.7 GB). 64 GB of RAM.
- Flux.2 Devfp8 · 35.5 GBOffloads
64 GB of RAM.
Set up ComfyUI.
Install or update ComfyUI
ComfyUI Desktop or the portable NVIDIA build. RTX 50 cards need a PyTorch built for CUDA 12.8 or newer; the current portable ships CUDA 13.0.
Keep the default memory handling
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.--reserve-vram 1keeps 1 GB free for Windows and the browser. Add it if the desktop stutters while ComfyUI renders.--vram-headroom 1asks Dynamic VRAM to keep extra room free, counting what other apps use.--fast-diskprefers streaming weights from disk over RAM. ComfyUI says it can be faster with a fast NVMe SSD.--disable-dynamic-vrambrings back the old loader, for a custom node that breaks with the new one. ComfyUI says this flag will be removed soon.
Pagefile and RAM
Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with
The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.With 16 GB of VRAM, 32 GB of RAM covers the image models. Wan 2.2 14B, HiDream and Flux.2 Dev are happier with 64 GB.
Where 16 GB still runs out.
- Qwen-Image in fp8 is 20.4 GB. Use a GGUF Q4_K_M (13.1 GB) or Q5_K_M (14.9 GB), or let it offload.
- Krea 2 at full precision is 26.3 GB. The fp8 file is what fits.
- Wan 2.2 14B swaps two 14.3 GB halves through RAM, which is where 16 GB cards with 32 GB or less crash [4].
- Flux.2 Dev needs a 32 GB card or 64 GB of RAM.
Some 16 GB owners with 32 GB of RAM saw memory handling get worse after ComfyUI updates in 2026 [5]. If a workflow that used to work now stalls, update again first; if it persists, --disable-dynamic-vram brings back the old loader [6].
5070 Ti or 5080?
Same 16 GB, same formats. The 5080 has about 20% more cores and slightly more bandwidth, so it’s a bit faster at everything; it doesn’t fit anything the 5070 Ti can’t. For ComfyUI, the difference in price buys speed, not models.
When it goes wrong.
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
Questions.
Is 16 GB of VRAM enough for ComfyUI in 2026?
For images, nearly always: Krea 2, Flux.1 Dev, Flux.2 Klein, Z-Image and Qwen-Image 2.1 all fit. For Qwen-Image in fp8, HiDream, Flux.2 Dev and Wan 2.2 14B it offloads to system RAM, which works with 32 to 64 GB of it.
Can I run Wan 2.2 14B on 16 GB?
Yes, as two fp8 halves with offloading. Give it 32 GB of RAM at least, 64 GB for 720p or long clips, and a large pagefile on Windows.
Should I use nvfp4 on a 5070 Ti or 5080?
For speed and headroom, yes: Krea 2 Turbo is 7.7 GB in nvfp4 against 13.1 GB in fp8. For the makers’ reference quality, compare against the fp8 file on your own prompts.
Why did ComfyUI get slower after an update?
ComfyUI changed how it manages memory in 2026, and some updates regressed on 16 GB cards. Update to the latest version. If it’s still slower, start ComfyUI with --disable-dynamic-vram to compare.
Sources: [1] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [2] SDXL GPU benchmark thread, ComfyUI discussion #2970, [3] Krea 2 files, Comfy-Org on Hugging Face, [4] Wan 2.2 crashes after a GPU upgrade, ComfyUI issue #12451, [5] Memory regression on 16 GB cards, ComfyUI issue #12541, [6] Dynamic VRAM on by default, ComfyUI discussion #12699, [7] Dynamic VRAM in ComfyUI, Comfy blog, [8] Steam Hardware Survey, August 2026.