Short answer
The 16 GB RTX 5060 Ti runs nearly every image model HEISS UI supports from its own memory, Krea 2 and Flux.1 Dev included, and it computes fp8 and nvfp4 natively. Wan 2.2 14B, Qwen-Image in fp8 and Flux.2 Dev need system RAM.
It’s not a fast card. It’s a card that rarely runs out, and that matters more here.
What fits in 16 GB.
Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.
- Krea 2 Turbofp8 13.1 GB · nvfp4 7.7 GBFits
fp8 for the makers’ reference quality, nvfp4 for speed.
- Flux.2 Klein 4B and 9Bbf16 7.8 GB · 9B fp8Fits
Both at 4 steps.
- Flux.1 Devfp8 · about 12 GBFits
About 26 s at 20 steps with SageAttention.
- Z-Image, Qwen-Image 2.1, Chroma, MageFlow4 to 12 GBFits
Any precision, bf16 included.
- Ideogram 4nvfp4 · 2 × 5.5 GBFits
The fp8 pair (2 × 9.3 GB) offloads a little.
- SDXLfp16 · 6.9 GBFits
With room for LoRAs, ControlNets and batches.
- Qwen-ImageGGUF Q4_K_M · 13.1 GBTight
The fp8 (20.4 GB) and nvfp4 (19.8 GB) files offload.
- SD 3.5 Large, ERNIE-Image14.9 and 16.1 GBTight
One image at a time at 1024 px.
- HiDream I1fp8 · 17.1 GBOffloads
Plus four text encoders.
- Wan 2.2 5Bfp16 · 10 GBFits
Native 1280 × 704, 121 frames.
- HunyuanVideo 1.5fp8 · 8.3 GBFits
720p is within reach.
- Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads
Each half fits alone; switching between them goes through RAM. 32 GB minimum.
- MiniMax H3w6a8 16 GB + nvfp4 encoder 15.7 GBOffloads
64 GB of RAM.
- Flux.2 Devfp8 · 35.5 GBOffloads
Plus its 12 to 18 GB text encoder. 64 GB of RAM, and slow.
Use the nvfp4 files.
The 5060 Ti is a Blackwell card, so it computes nvfp4, a 4-bit format, directly. Comfy-Org ships it for Krea 2 Turbo (7.7 GB instead of 13.1), Z-Image Turbo (4.5 GB), Ideogram 4 (2 × 5.5 GB) and Qwen-Image [5]. The file names end in _nvfp4.safetensors; the text encoders and VAE stay the same.
- Model7.7 GBDownload
krea2_turbo_nvfp4.safetensorsComfyUI/models/diffusion_models/ - Text encoder5.2 GBDownload
qwen3vl_4b_fp8_scaled.safetensorsComfyUI/models/text_encoders/ - VAE0.3 GBDownload
qwen_image_vae.safetensorsComfyUI/models/vae/
On the RTX 4060 Ti 16 GB (Ada), fp8 is native and nvfp4 is emulated: use the fp8 files.
Set up ComfyUI.
Install or update ComfyUI
RTX 50 cards need a PyTorch built for CUDA 12.8 or newer. Current ComfyUI Desktop and the portable build (CUDA 13.0) have it; an install from before 2025 doesn’t, and fails with
no kernel image is available for execution on the device.Leave the memory flags alone
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.--reserve-vram 1keeps 1 GB free for Windows and the browser. Add it if the desktop stutters while ComfyUI renders.--vram-headroom 1asks Dynamic VRAM to keep extra room free, counting what other apps use.--fast-diskprefers streaming weights from disk over RAM. ComfyUI says it can be faster with a fast NVMe SSD.--disable-dynamic-vrambrings back the old loader, for a custom node that breaks with the new one. ComfyUI says this flag will be removed soon.
Make it fail fast instead of crawling
On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the
python.exeComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.Pagefile and RAM
Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with
The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.
How long it takes.
| Card | Model and settings | Time |
|---|---|---|
| RTX 5060 Ti 16 GB | Flux.1 Dev fp8, 1024 px, 20 steps--use-sage-attention --fast | 25.7 s[1] |
| RTX 5060 Ti 16 GB | Qwen-Image 2.1 int8, 1760 × 2368, 20 steps32 GB of RAM, models on a hard disk. | 234 to 330 s[2] |
| RTX 5060 Ti 16 GB | Qwen-Image 2.1 GGUF Q8_0, same settings | 158 s[2] |
| RTX 4060 Ti 16 GB | SDXL, 1024 px, 20 stepsStock; 3.96 it/s tuned and overclocked. | 2.6 it/s[3] |
| RTX 3090 | Flux.1 Dev fp8, 1024 px, 20 stepsFor comparison. | 26 s[1] |
The Qwen-Image 2.1 test is worth a look: on this card the Q8 GGUF beat Comfy’s own int8 file, and its owner found the int8 version distorted bodies more often. ComfyUI recommends its native formats; test both on your card [7].
Video, RAM and the pagefile.
Wan 2.2 14B is where 16 GB cards trip. Each half is 14.3 GB in fp8 and fits on its own, but switching from the high-noise half to the low-noise half goes through system RAM. One owner who upgraded from an RTX 2060 to a 5060 Ti 16 GB, with 28 GB of RAM, found Wan 2.2 14B crashing on the second model; a 100 GB pagefile fixed it [4].
- 32 GB of RAM is the minimum for Wan 2.2 14B; 64 GB removes the pagefile from the picture.
- Keep the pagefile on System managed size, or set 64 GB or more if it still crashes.
- Wan 2.2 5B and HunyuanVideo 1.5 fit without any of this.
5060 Ti 16 GB or a used RTX 3090?
| RTX 5060 Ti 16 GB | RTX 3090 | |
|---|---|---|
| Graphics memory | 16 GB | 24 GB |
| Bandwidth | 448 GB/s | 936 GB/s |
| fp8 and nvfp4 compute | Yes | No |
| Flux.1 Dev fp8, 20 steps | 25.7 s[1] | 26 s[1] |
| Qwen-Image fp8, HiDream | Offloads | Fits |
| Power | 180 W | 350 W |
The 3090 fits more and moves memory twice as fast; the 5060 Ti is new, draws half the power and computes the new formats. For models up to 16 GB they end up close. For Qwen-Image, HiDream and 720p Wan 2.2, the 3090’s 24 GB wins. See RTX 3090 and 4090.
When it goes wrong.
- Wan 2.2 14B crashes on the second model
- The high-noise half is still in RAM when the low-noise half loads. Raise the pagefile to 64 GB or more, add RAM, or use the Q8 GGUF halves (15.4 GB each).
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
Questions.
RTX 5060 Ti 8 GB or 16 GB for AI images?
16 GB. For image and video models the 8 GB version runs a much shorter list. Check the box, or Task Manager once it’s in: both are sold as RTX 5060 Ti.
Is it worth upgrading from an RTX 3060 12 GB?
Yes, if you want Krea 2, Flux.1 Dev or 720p video without offloading. It’s also faster and computes fp8 and nvfp4. If your models already fit in 12 GB, 32 GB of system RAM is the cheaper upgrade.
5060 Ti 16 GB or 5070?
The 5070 is faster on anything that fits 12 GB. The 5060 Ti fits more. For ComfyUI, where running out means offloading, most people are better served by 16 GB.
Why does Wan 2.2 crash on the second model?
Switching from the high-noise half to the low-noise half goes through system RAM. With 32 GB or less, raise the pagefile to 64 GB or more, or add RAM.
int8 or GGUF for Qwen-Image 2.1?
Start with ComfyUI’s int8 file, which it recommends. On a 5060 Ti with 32 GB of RAM a Q8 GGUF ran faster in one test, but the main ComfyUI-GGUF release doesn’t load Qwen-Image 2.1 GGUFs yet, so that route needs a build of the pack that does.
Sources: [1] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [2] Qwen-Image 2.1 int8 and GGUF on an RTX 5060 Ti, ComfyUI issue #16470, [3] SDXL GPU benchmark thread, ComfyUI discussion #2970, [4] Wan 2.2 crashes after a GPU upgrade, ComfyUI issue #12451, [5] Krea 2 files, Comfy-Org on Hugging Face, [6] Dynamic VRAM in ComfyUI, Comfy blog, [7] Is GGUF still worth it, ComfyUI-GGUF issue #463, [8] Steam Hardware Survey, August 2026, [9] System memory fallback, NVIDIA.