Hardware guide

What runs on an RTX 5060 Ti 16 GB.

The 16 GB RTX 5060 Ti runs nearly every image model in ComfyUI from its own memory, Krea 2 and Flux.1 Dev included, and it computes the new fp8 and nvfp4 formats natively. Wan 2.2 14B and the 20 GB models need system RAM. This also covers the RTX 4060 Ti 16 GB.

Updated 29 Sep 20266 min read

Memory
16 GB GDDR7an 8 GB version exists
Bandwidth
448 GB/s128-bit, PCIe 5.0 x8
Architecture
Blackwellfp8 and nvfp4 compute
On Steam
2.55%Aug 2026

Short answer

The 16 GB RTX 5060 Ti runs nearly every image model HEISS UI supports from its own memory, Krea 2 and Flux.1 Dev included, and it computes fp8 and nvfp4 natively. Wan 2.2 14B, Qwen-Image in fp8 and Flux.2 Dev need system RAM.

It’s not a fast card. It’s a card that rarely runs out, and that matters more here.

What fits in 16 GB.

Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.

  • Krea 2 Turbofp8 13.1 GB · nvfp4 7.7 GBFits

    fp8 for the makers’ reference quality, nvfp4 for speed.

  • Flux.2 Klein 4B and 9Bbf16 7.8 GB · 9B fp8Fits

    Both at 4 steps.

  • Flux.1 Devfp8 · about 12 GBFits

    About 26 s at 20 steps with SageAttention.

  • Z-Image, Qwen-Image 2.1, Chroma, MageFlow4 to 12 GBFits

    Any precision, bf16 included.

  • Ideogram 4nvfp4 · 2 × 5.5 GBFits

    The fp8 pair (2 × 9.3 GB) offloads a little.

  • SDXLfp16 · 6.9 GBFits

    With room for LoRAs, ControlNets and batches.

  • Qwen-ImageGGUF Q4_K_M · 13.1 GBTight

    The fp8 (20.4 GB) and nvfp4 (19.8 GB) files offload.

  • SD 3.5 Large, ERNIE-Image14.9 and 16.1 GBTight

    One image at a time at 1024 px.

  • HiDream I1fp8 · 17.1 GBOffloads

    Plus four text encoders.

  • Wan 2.2 5Bfp16 · 10 GBFits

    Native 1280 × 704, 121 frames.

  • HunyuanVideo 1.5fp8 · 8.3 GBFits

    720p is within reach.

  • Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads

    Each half fits alone; switching between them goes through RAM. 32 GB minimum.

  • MiniMax H3w6a8 16 GB + nvfp4 encoder 15.7 GBOffloads

    64 GB of RAM.

  • Flux.2 Devfp8 · 35.5 GBOffloads

    Plus its 12 to 18 GB text encoder. 64 GB of RAM, and slow.

Use the nvfp4 files.

The 5060 Ti is a Blackwell card, so it computes nvfp4, a 4-bit format, directly. Comfy-Org ships it for Krea 2 Turbo (7.7 GB instead of 13.1), Z-Image Turbo (4.5 GB), Ideogram 4 (2 × 5.5 GB) and Qwen-Image [5]. The file names end in _nvfp4.safetensors; the text encoders and VAE stay the same.

  • Modelkrea2_turbo_nvfp4.safetensorsComfyUI/models/diffusion_models/
    7.7 GBDownload
  • Text encoderqwen3vl_4b_fp8_scaled.safetensorsComfyUI/models/text_encoders/
    5.2 GBDownload
  • VAEqwen_image_vae.safetensorsComfyUI/models/vae/
    0.3 GBDownload

On the RTX 4060 Ti 16 GB (Ada), fp8 is native and nvfp4 is emulated: use the fp8 files.

Set up ComfyUI.

  1. Install or update ComfyUI

    RTX 50 cards need a PyTorch built for CUDA 12.8 or newer. Current ComfyUI Desktop and the portable build (CUDA 13.0) have it; an install from before 2025 doesn’t, and fails with no kernel image is available for execution on the device.

  2. Leave the memory flags alone

    Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).

    That changes the old advice. --lowvram now does nothing while Dynamic VRAM is on, and --novram, --highvram and --gpu-only switch it off. --normalvram is gone, and an old shortcut that still passes it stops ComfyUI with unrecognized arguments.

    • --reserve-vram 1 keeps 1 GB free for Windows and the browser. Add it if the desktop stutters while ComfyUI renders.
    • --vram-headroom 1 asks Dynamic VRAM to keep extra room free, counting what other apps use.
    • --fast-disk prefers streaming weights from disk over RAM. ComfyUI says it can be faster with a fast NVMe SSD.
    • --disable-dynamic-vram brings back the old loader, for a custom node that breaks with the new one. ComfyUI says this flag will be removed soon.
  3. Make it fail fast instead of crawling

    On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the python.exe ComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.

  4. Pagefile and RAM

    Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.

How long it takes.

CardModel and settingsTime
RTX 5060 Ti 16 GBFlux.1 Dev fp8, 1024 px, 20 steps--use-sage-attention --fast25.7 s[1]
RTX 5060 Ti 16 GBQwen-Image 2.1 int8, 1760 × 2368, 20 steps32 GB of RAM, models on a hard disk.234 to 330 s[2]
RTX 5060 Ti 16 GBQwen-Image 2.1 GGUF Q8_0, same settings158 s[2]
RTX 4060 Ti 16 GBSDXL, 1024 px, 20 stepsStock; 3.96 it/s tuned and overclocked.2.6 it/s[3]
RTX 3090Flux.1 Dev fp8, 1024 px, 20 stepsFor comparison.26 s[1]

The Qwen-Image 2.1 test is worth a look: on this card the Q8 GGUF beat Comfy’s own int8 file, and its owner found the int8 version distorted bodies more often. ComfyUI recommends its native formats; test both on your card [7].

Video, RAM and the pagefile.

Wan 2.2 14B is where 16 GB cards trip. Each half is 14.3 GB in fp8 and fits on its own, but switching from the high-noise half to the low-noise half goes through system RAM. One owner who upgraded from an RTX 2060 to a 5060 Ti 16 GB, with 28 GB of RAM, found Wan 2.2 14B crashing on the second model; a 100 GB pagefile fixed it [4].

  • 32 GB of RAM is the minimum for Wan 2.2 14B; 64 GB removes the pagefile from the picture.
  • Keep the pagefile on System managed size, or set 64 GB or more if it still crashes.
  • Wan 2.2 5B and HunyuanVideo 1.5 fit without any of this.

5060 Ti 16 GB or a used RTX 3090?

RTX 5060 Ti 16 GBRTX 3090
Graphics memory16 GB24 GB
Bandwidth448 GB/s936 GB/s
fp8 and nvfp4 computeYesNo
Flux.1 Dev fp8, 20 steps25.7 s[1]26 s[1]
Qwen-Image fp8, HiDreamOffloadsFits
Power180 W350 W

The 3090 fits more and moves memory twice as fast; the 5060 Ti is new, draws half the power and computes the new formats. For models up to 16 GB they end up close. For Qwen-Image, HiDream and 720p Wan 2.2, the 3090’s 24 GB wins. See RTX 3090 and 4090.

When it goes wrong.

Wan 2.2 14B crashes on the second model
The high-noise half is still in RAM when the low-noise half loads. Raise the pagefile to 64 GB or more, add RAM, or use the Q8 GGUF halves (15.4 GB each).
torch.OutOfMemoryError: CUDA out of memory
The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)
Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
It suddenly got much slower
The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram
An old launch flag. Remove it from the shortcut or the .bat file. Current ComfyUI doesn’t need memory flags on NVIDIA.
The first image takes minutes, the next ones don’t
The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.

Questions.

RTX 5060 Ti 8 GB or 16 GB for AI images?

16 GB. For image and video models the 8 GB version runs a much shorter list. Check the box, or Task Manager once it’s in: both are sold as RTX 5060 Ti.

Is it worth upgrading from an RTX 3060 12 GB?

Yes, if you want Krea 2, Flux.1 Dev or 720p video without offloading. It’s also faster and computes fp8 and nvfp4. If your models already fit in 12 GB, 32 GB of system RAM is the cheaper upgrade.

5060 Ti 16 GB or 5070?

The 5070 is faster on anything that fits 12 GB. The 5060 Ti fits more. For ComfyUI, where running out means offloading, most people are better served by 16 GB.

Why does Wan 2.2 crash on the second model?

Switching from the high-noise half to the low-noise half goes through system RAM. With 32 GB or less, raise the pagefile to 64 GB or more, or add RAM.

int8 or GGUF for Qwen-Image 2.1?

Start with ComfyUI’s int8 file, which it recommends. On a 5060 Ti with 32 GB of RAM a Q8 GGUF ran faster in one test, but the main ComfyUI-GGUF release doesn’t load Qwen-Image 2.1 GGUFs yet, so that route needs a build of the pack that does.

Sources: [1] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [2] Qwen-Image 2.1 int8 and GGUF on an RTX 5060 Ti, ComfyUI issue #16470, [3] SDXL GPU benchmark thread, ComfyUI discussion #2970, [4] Wan 2.2 crashes after a GPU upgrade, ComfyUI issue #12451, [5] Krea 2 files, Comfy-Org on Hugging Face, [6] Dynamic VRAM in ComfyUI, Comfy blog, [7] Is GGUF still worth it, ComfyUI-GGUF issue #463, [8] Steam Hardware Survey, August 2026, [9] System memory fallback, NVIDIA.

HEISS UI

Sized for 16 GB.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. It reads your card and does the fitting for you.

  • The model that fits your 16 GB is marked. Pick a first model and the size that suits this card is already chosen. One tap downloads it with everything it needs.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.