Short answer
On 8 GB, SDXL, Flux.2 Klein 4B, Z-Image Turbo and MageFlow run entirely on the card, and Wan 2.2 5B makes short clips. Krea 2, Flux.1 Dev and Wan 2.2 14B run too, with part of the model in system RAM, if the computer has 32 GB of it.
An RTX 50 card goes further: the nvfp4 files fit where fp8 didn’t, Krea 2 Turbo included.
Which 8 GB card you have.
A quarter of Steam users have 8 GB of graphics memory [9]. The cards differ in two ways that matter here: how fast they read memory, and which small number formats they compute natively.
| Card | Bandwidth | Computes natively | Note |
|---|---|---|---|
| RTX 5060 | 448 GB/s | fp8, nvfp4 | PCIe 5.0 x8 |
| RTX 5060 Ti 8 GB | 448 GB/s | fp8, nvfp4 | A 16 GB version exists |
| RTX 4060 | 272 GB/s | fp8 | PCIe 4.0 x8 |
| RTX 4060 Ti 8 GB | 288 GB/s | fp8 | A 16 GB version exists |
| RTX 3070 / 3070 Ti | 448 / 608 GB/s | none | Like the RTX 3060 Ti |
| RTX 2070 / 2080 | 448 GB/s | none | Oldest the current portable build supports |
“None” still runs fp8 and nvfp4 files: ComfyUI unpacks them as it goes. They save memory there, not time.
What fits in 8 GB.
Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM and streams it in: it works, slower.
- SDXLfp16 · 6.9 GBFits
1024 px, one image at a time. Peaked at 5.6 GB on an RTX 4060 Laptop.
- Flux.2 Klein 4Bfp8 · 4.1 GBFits
4 steps. The best image model that sits fully on 8 GB.
- Z-Image Turboint8 6.2 GB · nvfp4 4.5 GBFits
nvfp4 on RTX 50, int8 on everything else.
- MageFlow Turboint8 · 4.2 GBFits
4 steps.
- SD 1.5, Anima, Sana, Lumina 22 to 5 GBFits
Small models.
- Krea 2 Turbonvfp4 · 7.7 GB · RTX 50Tight
Made for RTX 50. Other cards use the fp8 file (13.1 GB), which offloads.
- Flux.1 DevGGUF Q4_K_S · 6.8 GBTight
Or fp8 with offloading and 32 GB of RAM.
- ChromaGGUF Q4_K_M · 5.6 GBTight
The fp8 file (9.2 GB) offloads.
- Wan 2.2 5Bfp16 10 GB · GGUF 3.4 to 5.4 GBTight
ComfyUI’s docs say it fits 8 GB with its offloading. A GGUF Q5 to Q8 leaves more room.
- Ideogram 4nvfp4 · 2 × 5.5 GBOffloads
Two models side by side. On RTX 50 the nvfp4 pair is the lightest way.
- Qwen-Image 2.1int8 · 7.3 GBOffloads
With its 9.4 GB text encoder, the set needs system RAM.
- HunyuanVideo 1.5fp8 · 8.3 GBOffloads
480p.
- Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads
With the 4-step LoRAs and 32 GB of RAM, a short 480p clip takes about two minutes.
- Qwen-Image, HiDream, SD 3.5 Large15 to 20 GBOffloads
Only with 64 GB of RAM.
- Flux.2 Dev, MiniMax H335 GB and 16 to 21 GBNo
Plus text encoders of 12 to 18 GB. A 24 GB card is the sensible floor.
RTX 50: use the nvfp4 files.
Comfy-Org now ships several models in nvfp4, a 4-bit format that RTX 50 cards compute directly. For 8 GB that’s the difference between offloading and fitting [7]:
| Model | fp8 or int8 | nvfp4 |
|---|---|---|
| Krea 2 Turbo | 13.1 GB | 7.7 GB |
| Z-Image Turbo | 6.2 GB (int8) | 4.5 GB |
| Ideogram 4 | 2 × 9.3 GB | 2 × 5.5 GB |
| Qwen-Image | 20.4 GB | 19.8 GB |
File names end in _nvfp4.safetensors. On RTX 30 and 40 cards ComfyUI emulates the format, so the speed is an RTX 50 thing; fp8, int8 and GGUF are the usual choice there.
Set up ComfyUI for an 8 GB card.
Install or update ComfyUI
ComfyUI Desktop or the portable NVIDIA build, with a current driver. The portable build ships CUDA 13.0 and supports the 20 series and newer.
Leave the memory flags alone
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.If the desktop stutters during a render, add
--reserve-vram 1:run_nvidia_gpu.bat.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --reserve-vram 1 pause
Make it fail fast instead of crawling
On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the
python.exeComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.Give Windows pagefile room
Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with
The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.
How long it takes.
| Card | Model and settings | Time |
|---|---|---|
| RTX 4060 | SDXL, 1024 px, 20 steps | 15.9 s[1] |
| RTX 4060 Laptop | SDXL (WAI-Illustrious), 1024 px, 20 stepsNo --lowvram, 5.6 GB peak. | 15.8 s[2] |
| RTX 3070 | SDXL, 1024 px, 20 steps | 10.9 s[1] |
| RTX 3060 Ti | SDXL, 1024 px, 20 steps | 13.3 s[1] |
| RTX 4060 | Wan 2.2 14B, 4 steps, 480 × 480, 33 frames32 GB of RAM. The 5B took 95 to 179 s here and looked worse. | 111 s[3] |
There are no published RTX 5060 or 5060 Ti 8 GB numbers for these workloads yet.
Video on 8 GB.
Trade size for length. An RTX 5060 Ti 8 GB owner who pushed Wan 2.2 5B to 768 × 768 got two seconds of video; the advice was 672 × 384 and 33 frames at 24 fps, 12 to 16 steps, and to go up from there [4]. The distilled 4-step Wan 2.2 14B at 480 × 480 turned out both faster and better than the 5B on an RTX 4060 [3].
6 GB cards.
An RTX 3060 Laptop, RTX 4050 Laptop or RTX 2060 has 6 GB. SDXL still works at 1024 px, best with a Lightning or DMD2 checkpoint that needs 4 to 8 steps. Flux.2 Klein 4B in fp8 fits barely; one user with a 4 GB card watched it spill into shared memory while Z-Image and Wan 2.2 stayed fine [8]. For video, Wan 2.2 5B as a GGUF Q4 at 480p and short clips.
RAM is the real limit.
| System RAM | What it changes |
|---|---|
| 16 GB | Everything marked Fits. Offloaded models load slowly and lean on the pagefile. |
| 32 GB | Krea 2, Flux.1 Dev, Qwen-Image 2.1 and Wan 2.2 14B become practical. |
| 64 GB | The 20 GB models, and Wan 2.2 14B without pagefile trouble. |
When it goes wrong.
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
Questions.
Can I run Flux on 8 GB of VRAM?
Yes. Flux.2 Klein 4B in fp8 runs fully on the card. Flux.1 Dev runs as a GGUF Q4 or in fp8 with offloading and 32 GB of RAM. Flux.2 Dev doesn’t fit a sensible 8 GB setup.
Can I run Wan 2.2 on 8 GB with 16 GB of RAM?
Wan 2.2 5B, yes, as a GGUF at around 480p and 33 to 49 frames. The 14B pair leans hard on the pagefile with 16 GB; 32 GB of RAM is the practical floor for it.
RTX 4060 or RTX 3060 12 GB for AI images?
The 4060 computes fp8 natively and draws less power; the 3060 12 GB fits more. For models that fit in 8 GB they’re close. For Flux.2 Klein at full precision, Krea 2 or Wan 2.2 5B in fp16, the 12 GB card is the better fit.
What is nvfp4, and should I use it?
A 4-bit format that RTX 50 cards compute directly. On a 5060 or 5060 Ti, yes: Krea 2 Turbo drops from 13.1 GB in fp8 to 7.7 GB. On RTX 30 and 40 cards ComfyUI emulates it, so fp8, int8 or GGUF is the usual choice there.
GGUF or fp8?
ComfyUI recommends its own fp8 and int8 files, and they’re usually faster. A GGUF Q4_K_M to Q8_0 is the way in when RAM is short. The ComfyUI-GGUF pack doesn’t load Krea 2, Ideogram 4, MiniMax H3 or Qwen-Image 2.1 GGUFs yet.
Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] SDXL on an RTX 4060 Laptop, lilting channel, [3] Wan 2.2 on an RTX 4060 8 GB, lilting channel, [4] Photo to video on 8 GB, Hugging Face forum, [5] Wan 2.2 in ComfyUI, Comfy docs, [6] Dynamic VRAM in ComfyUI, Comfy blog, [7] Krea 2 files, Comfy-Org on Hugging Face, [8] Flux.2 Klein spilling into shared memory on 4 GB, ComfyUI issue #11913, [9] Steam Hardware Survey, August 2026, [10] System memory fallback, NVIDIA.