Hardware guide

What runs on 8 GB of VRAM.

On an 8 GB card, the RTX 4060, 5060, 3070 or the 8 GB 4060 Ti and 5060 Ti, ComfyUI runs SDXL, Flux.2 Klein 4B, Z-Image Turbo and Wan 2.2 5B well. Bigger models run from system RAM, slower. RTX 50 cards also run the new nvfp4 files, which fit more.

Updated 29 Sep 20266 min read

Cards
RTX 4060, 5060, 3070and the 8 GB 4060 Ti, 5060 Ti
Steam users
25.7%have 8 GB, Aug 2026
Best start
Flux.2 Klein 4Bfp8, 4 steps
System RAM
32 GBfor models that offload

Short answer

On 8 GB, SDXL, Flux.2 Klein 4B, Z-Image Turbo and MageFlow run entirely on the card, and Wan 2.2 5B makes short clips. Krea 2, Flux.1 Dev and Wan 2.2 14B run too, with part of the model in system RAM, if the computer has 32 GB of it.

An RTX 50 card goes further: the nvfp4 files fit where fp8 didn’t, Krea 2 Turbo included.

Which 8 GB card you have.

A quarter of Steam users have 8 GB of graphics memory [9]. The cards differ in two ways that matter here: how fast they read memory, and which small number formats they compute natively.

CardBandwidthComputes nativelyNote
RTX 5060448 GB/sfp8, nvfp4PCIe 5.0 x8
RTX 5060 Ti 8 GB448 GB/sfp8, nvfp4A 16 GB version exists
RTX 4060272 GB/sfp8PCIe 4.0 x8
RTX 4060 Ti 8 GB288 GB/sfp8A 16 GB version exists
RTX 3070 / 3070 Ti448 / 608 GB/snoneLike the RTX 3060 Ti
RTX 2070 / 2080448 GB/snoneOldest the current portable build supports

“None” still runs fp8 and nvfp4 files: ComfyUI unpacks them as it goes. They save memory there, not time.

What fits in 8 GB.

Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM and streams it in: it works, slower.

  • SDXLfp16 · 6.9 GBFits

    1024 px, one image at a time. Peaked at 5.6 GB on an RTX 4060 Laptop.

  • Flux.2 Klein 4Bfp8 · 4.1 GBFits

    4 steps. The best image model that sits fully on 8 GB.

  • Z-Image Turboint8 6.2 GB · nvfp4 4.5 GBFits

    nvfp4 on RTX 50, int8 on everything else.

  • MageFlow Turboint8 · 4.2 GBFits

    4 steps.

  • SD 1.5, Anima, Sana, Lumina 22 to 5 GBFits

    Small models.

  • Krea 2 Turbonvfp4 · 7.7 GB · RTX 50Tight

    Made for RTX 50. Other cards use the fp8 file (13.1 GB), which offloads.

  • Flux.1 DevGGUF Q4_K_S · 6.8 GBTight

    Or fp8 with offloading and 32 GB of RAM.

  • ChromaGGUF Q4_K_M · 5.6 GBTight

    The fp8 file (9.2 GB) offloads.

  • Wan 2.2 5Bfp16 10 GB · GGUF 3.4 to 5.4 GBTight

    ComfyUI’s docs say it fits 8 GB with its offloading. A GGUF Q5 to Q8 leaves more room.

  • Ideogram 4nvfp4 · 2 × 5.5 GBOffloads

    Two models side by side. On RTX 50 the nvfp4 pair is the lightest way.

  • Qwen-Image 2.1int8 · 7.3 GBOffloads

    With its 9.4 GB text encoder, the set needs system RAM.

  • HunyuanVideo 1.5fp8 · 8.3 GBOffloads

    480p.

  • Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads

    With the 4-step LoRAs and 32 GB of RAM, a short 480p clip takes about two minutes.

  • Qwen-Image, HiDream, SD 3.5 Large15 to 20 GBOffloads

    Only with 64 GB of RAM.

  • Flux.2 Dev, MiniMax H335 GB and 16 to 21 GBNo

    Plus text encoders of 12 to 18 GB. A 24 GB card is the sensible floor.

RTX 50: use the nvfp4 files.

Comfy-Org now ships several models in nvfp4, a 4-bit format that RTX 50 cards compute directly. For 8 GB that’s the difference between offloading and fitting [7]:

Modelfp8 or int8nvfp4
Krea 2 Turbo13.1 GB7.7 GB
Z-Image Turbo6.2 GB (int8)4.5 GB
Ideogram 42 × 9.3 GB2 × 5.5 GB
Qwen-Image20.4 GB19.8 GB

File names end in _nvfp4.safetensors. On RTX 30 and 40 cards ComfyUI emulates the format, so the speed is an RTX 50 thing; fp8, int8 and GGUF are the usual choice there.

Set up ComfyUI for an 8 GB card.

  1. Install or update ComfyUI

    ComfyUI Desktop or the portable NVIDIA build, with a current driver. The portable build ships CUDA 13.0 and supports the 20 series and newer.

  2. Leave the memory flags alone

    Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).

    That changes the old advice. --lowvram now does nothing while Dynamic VRAM is on, and --novram, --highvram and --gpu-only switch it off. --normalvram is gone, and an old shortcut that still passes it stops ComfyUI with unrecognized arguments.

    If the desktop stutters during a render, add --reserve-vram 1:

    run_nvidia_gpu.bat
    .\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --reserve-vram 1
    pause
  3. Make it fail fast instead of crawling

    On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the python.exe ComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.

  4. Give Windows pagefile room

    Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.

How long it takes.

CardModel and settingsTime
RTX 4060SDXL, 1024 px, 20 steps15.9 s[1]
RTX 4060 LaptopSDXL (WAI-Illustrious), 1024 px, 20 stepsNo --lowvram, 5.6 GB peak.15.8 s[2]
RTX 3070SDXL, 1024 px, 20 steps10.9 s[1]
RTX 3060 TiSDXL, 1024 px, 20 steps13.3 s[1]
RTX 4060Wan 2.2 14B, 4 steps, 480 × 480, 33 frames32 GB of RAM. The 5B took 95 to 179 s here and looked worse.111 s[3]

There are no published RTX 5060 or 5060 Ti 8 GB numbers for these workloads yet.

Video on 8 GB.

Trade size for length. An RTX 5060 Ti 8 GB owner who pushed Wan 2.2 5B to 768 × 768 got two seconds of video; the advice was 672 × 384 and 33 frames at 24 fps, 12 to 16 steps, and to go up from there [4]. The distilled 4-step Wan 2.2 14B at 480 × 480 turned out both faster and better than the 5B on an RTX 4060 [3].

6 GB cards.

An RTX 3060 Laptop, RTX 4050 Laptop or RTX 2060 has 6 GB. SDXL still works at 1024 px, best with a Lightning or DMD2 checkpoint that needs 4 to 8 steps. Flux.2 Klein 4B in fp8 fits barely; one user with a 4 GB card watched it spill into shared memory while Z-Image and Wan 2.2 stayed fine [8]. For video, Wan 2.2 5B as a GGUF Q4 at 480p and short clips.

RAM is the real limit.

System RAMWhat it changes
16 GBEverything marked Fits. Offloaded models load slowly and lean on the pagefile.
32 GBKrea 2, Flux.1 Dev, Qwen-Image 2.1 and Wan 2.2 14B become practical.
64 GBThe 20 GB models, and Wan 2.2 14B without pagefile trouble.

When it goes wrong.

torch.OutOfMemoryError: CUDA out of memory
The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)
Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
It suddenly got much slower
The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram
An old launch flag. Remove it from the shortcut or the .bat file. Current ComfyUI doesn’t need memory flags on NVIDIA.
The first image takes minutes, the next ones don’t
The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.

Questions.

Can I run Flux on 8 GB of VRAM?

Yes. Flux.2 Klein 4B in fp8 runs fully on the card. Flux.1 Dev runs as a GGUF Q4 or in fp8 with offloading and 32 GB of RAM. Flux.2 Dev doesn’t fit a sensible 8 GB setup.

Can I run Wan 2.2 on 8 GB with 16 GB of RAM?

Wan 2.2 5B, yes, as a GGUF at around 480p and 33 to 49 frames. The 14B pair leans hard on the pagefile with 16 GB; 32 GB of RAM is the practical floor for it.

RTX 4060 or RTX 3060 12 GB for AI images?

The 4060 computes fp8 natively and draws less power; the 3060 12 GB fits more. For models that fit in 8 GB they’re close. For Flux.2 Klein at full precision, Krea 2 or Wan 2.2 5B in fp16, the 12 GB card is the better fit.

What is nvfp4, and should I use it?

A 4-bit format that RTX 50 cards compute directly. On a 5060 or 5060 Ti, yes: Krea 2 Turbo drops from 13.1 GB in fp8 to 7.7 GB. On RTX 30 and 40 cards ComfyUI emulates it, so fp8, int8 or GGUF is the usual choice there.

GGUF or fp8?

ComfyUI recommends its own fp8 and int8 files, and they’re usually faster. A GGUF Q4_K_M to Q8_0 is the way in when RAM is short. The ComfyUI-GGUF pack doesn’t load Krea 2, Ideogram 4, MiniMax H3 or Qwen-Image 2.1 GGUFs yet.

Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] SDXL on an RTX 4060 Laptop, lilting channel, [3] Wan 2.2 on an RTX 4060 8 GB, lilting channel, [4] Photo to video on 8 GB, Hugging Face forum, [5] Wan 2.2 in ComfyUI, Comfy docs, [6] Dynamic VRAM in ComfyUI, Comfy blog, [7] Krea 2 files, Comfy-Org on Hugging Face, [8] Flux.2 Klein spilling into shared memory on 4 GB, ComfyUI issue #11913, [9] Steam Hardware Survey, August 2026, [10] System memory fallback, NVIDIA.

HEISS UI

Made to fit 8 GB.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. It reads your card and does the fitting for you.

  • The model that fits your 8 GB is marked. Pick a first model and the size that suits this card is already chosen. One tap downloads it with everything it needs.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.