Hardware guide

What runs on 16 GB of VRAM.

16 GB runs every image model up to about 15 GB from the card, Krea 2 and Flux.1 Dev included, and 720p video with Wan 2.2 5B and HunyuanVideo 1.5. Qwen-Image in fp8, Wan 2.2 14B and Flux.2 Dev still need system RAM. This covers the RTX 5070 Ti, 5080, 4080 and 4070 Ti Super.

Updated 29 Sep 20265 min read

Cards
RTX 5070 Ti, 50804080, 4070 Ti Super
Steam users
26.9%have 16 GB, Aug 2026
Best start
Krea 2 Turbofp8, or nvfp4 on RTX 50
System RAM
32 to 64 GB64 GB for Wan 2.2 14B

Short answer

16 GB runs every image model HEISS UI supports up to about 15 GB from the card: Krea 2, Flux.1 Dev, Flux.2 Klein 9B, Z-Image, Qwen-Image 2.1, SD 3.5 Large, plus Wan 2.2 5B and HunyuanVideo 1.5 at 720p. Qwen-Image in fp8, HiDream, Wan 2.2 14B and Flux.2 Dev still use system RAM.

On a 5070 Ti or 5080, the nvfp4 files leave room to spare.

The 16 GB cards.

CardBandwidthComputes nativelyFlux.1 Dev fp8, 20 steps
RTX 5080960 GB/sfp8, nvfp416.7 s[1]6.7 s tuned
RTX 5070 Ti896 GB/sfp8, nvfp4No published number
RTX 4080 / Super717 / 736 GB/sfp8SDXL: 6.5 s[2]
RTX 4070 Ti Super672 GB/sfp8No published number
RTX 5060 Ti / 4060 Ti 16 GB448 / 288 GB/sfp8 (and nvfp4 on the 5060 Ti)Their own page

AMD’s 16 GB cards (RX 9070 XT, 9070, 7800 XT, 7600 XT) fit the same models: see the AMD guide.

What fits in 16 GB.

Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.

Set up ComfyUI.

  1. Install or update ComfyUI

    ComfyUI Desktop or the portable NVIDIA build. RTX 50 cards need a PyTorch built for CUDA 12.8 or newer; the current portable ships CUDA 13.0.

  2. Keep the default memory handling

    Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).

    That changes the old advice. --lowvram now does nothing while Dynamic VRAM is on, and --novram, --highvram and --gpu-only switch it off. --normalvram is gone, and an old shortcut that still passes it stops ComfyUI with unrecognized arguments.

    • --reserve-vram 1 keeps 1 GB free for Windows and the browser. Add it if the desktop stutters while ComfyUI renders.
    • --vram-headroom 1 asks Dynamic VRAM to keep extra room free, counting what other apps use.
    • --fast-disk prefers streaming weights from disk over RAM. ComfyUI says it can be faster with a fast NVMe SSD.
    • --disable-dynamic-vram brings back the old loader, for a custom node that breaks with the new one. ComfyUI says this flag will be removed soon.
  3. Pagefile and RAM

    Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.

    With 16 GB of VRAM, 32 GB of RAM covers the image models. Wan 2.2 14B, HiDream and Flux.2 Dev are happier with 64 GB.

Where 16 GB still runs out.

  • Qwen-Image in fp8 is 20.4 GB. Use a GGUF Q4_K_M (13.1 GB) or Q5_K_M (14.9 GB), or let it offload.
  • Krea 2 at full precision is 26.3 GB. The fp8 file is what fits.
  • Wan 2.2 14B swaps two 14.3 GB halves through RAM, which is where 16 GB cards with 32 GB or less crash [4].
  • Flux.2 Dev needs a 32 GB card or 64 GB of RAM.

Some 16 GB owners with 32 GB of RAM saw memory handling get worse after ComfyUI updates in 2026 [5]. If a workflow that used to work now stalls, update again first; if it persists, --disable-dynamic-vram brings back the old loader [6].

5070 Ti or 5080?

Same 16 GB, same formats. The 5080 has about 20% more cores and slightly more bandwidth, so it’s a bit faster at everything; it doesn’t fit anything the 5070 Ti can’t. For ComfyUI, the difference in price buys speed, not models.

When it goes wrong.

torch.OutOfMemoryError: CUDA out of memory
The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)
Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
It suddenly got much slower
The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram
An old launch flag. Remove it from the shortcut or the .bat file. Current ComfyUI doesn’t need memory flags on NVIDIA.
The first image takes minutes, the next ones don’t
The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.

Questions.

Is 16 GB of VRAM enough for ComfyUI in 2026?

For images, nearly always: Krea 2, Flux.1 Dev, Flux.2 Klein, Z-Image and Qwen-Image 2.1 all fit. For Qwen-Image in fp8, HiDream, Flux.2 Dev and Wan 2.2 14B it offloads to system RAM, which works with 32 to 64 GB of it.

Can I run Wan 2.2 14B on 16 GB?

Yes, as two fp8 halves with offloading. Give it 32 GB of RAM at least, 64 GB for 720p or long clips, and a large pagefile on Windows.

Should I use nvfp4 on a 5070 Ti or 5080?

For speed and headroom, yes: Krea 2 Turbo is 7.7 GB in nvfp4 against 13.1 GB in fp8. For the makers’ reference quality, compare against the fp8 file on your own prompts.

Why did ComfyUI get slower after an update?

ComfyUI changed how it manages memory in 2026, and some updates regressed on 16 GB cards. Update to the latest version. If it’s still slower, start ComfyUI with --disable-dynamic-vram to compare.

Sources: [1] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [2] SDXL GPU benchmark thread, ComfyUI discussion #2970, [3] Krea 2 files, Comfy-Org on Hugging Face, [4] Wan 2.2 crashes after a GPU upgrade, ComfyUI issue #12451, [5] Memory regression on 16 GB cards, ComfyUI issue #12541, [6] Dynamic VRAM on by default, ComfyUI discussion #12699, [7] Dynamic VRAM in ComfyUI, Comfy blog, [8] Steam Hardware Survey, August 2026.

HEISS UI

Sized for 16 GB.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. It reads your card and does the fitting for you.

  • The model that fits your 16 GB is marked. Pick a first model and the size that suits this card is already chosen. One tap downloads it with everything it needs.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.