Short answer
A 12 GB card runs SDXL, Flux.2 Klein 4B at full precision, Z-Image, Qwen-Image 2.1, Chroma and Wan 2.2 5B from its own memory, and Krea 2 Turbo in fp8 just about. Wan 2.2 14B, Qwen-Image and Flux.2 Dev need system RAM.
On an RTX 5070, the nvfp4 files put Krea 2 Turbo at 7.7 GB, with room to spare.
The 12 GB cards.
This page covers the RTX 4070 family, the 5070 and the RTX 3080 12 GB and 3080 Ti. The RTX 3060 12 GB has its own page.
| Card | Bandwidth | Computes natively | SDXL, 1024 px, 20 steps |
|---|---|---|---|
| RTX 5070 | 672 GB/s | fp8, nvfp4 | No published number |
| RTX 4070 / Super / Ti | 504 GB/s | fp8 | 4070: 7.1 s[1] |
| RTX 3080 12 GB / 3080 Ti | 912 GB/s | none | 3080 Ti: 5.6 s[1] |
| RTX 3060 12 GB | 360 GB/s | none | Its own page |
The RTX 5070 is the third most common card on Steam in August 2026, at 3.77%.
What fits in 12 GB.
Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.
- SDXLfp16 · 6.9 GBFits
With LoRAs and a batch of two.
- Flux.2 Klein 4Bbf16 · 7.8 GBFits
Full precision, 4 steps.
- Flux.2 Klein 9BGGUF Q8_0 · 10 GBTight
Or Q5_K_M (7 GB). Comfy measures 19.6 GB at bf16.
- Z-Image Turboint8 · 6.2 GBFits
The bf16 file (12.3 GB) is tight; nvfp4 (4.5 GB) on a 5070.
- Qwen-Image 2.1int8 · 7.3 GBFits
The text encoder runs first and then makes room.
- Chromafp8 · 9.2 GBFits
Or a GGUF Q8_0 (9.7 GB).
- MageFlowint8 · 4.2 GBFits
4 steps with the Turbo version.
- Krea 2 Turbofp8 13.1 GB · nvfp4 7.7 GBTight
fp8 peaked at 11.8 GB at 1080p on a 3060 12 GB. On a 5070, the nvfp4 file fits easily.
- Flux.1 Devfp8 · about 12 GBTight
Or GGUF Q8_0 (12.7 GB) / Q6_K (9.9 GB). 32 GB of RAM.
- Ideogram 4fp8 2 × 9.3 GB · nvfp4 2 × 5.5 GBOffloads
Two models side by side; on a 5070 the nvfp4 pair nearly fits.
- Wan 2.2 5Bfp16 · 10 GBFits
Native 1280 × 704. Start with shorter clips.
- HunyuanVideo 1.5fp8 · 8.3 GBFits
480p; 720p offloads.
- Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads
832 × 480 with the 4-step LoRAs and 32 GB of RAM.
- Qwen-Image, HiDream, SD 3.5 Large, ERNIE15 to 20 GBOffloads
Qwen-Image as a GGUF Q4_K_M (13.1 GB) is the easiest of these.
- MiniMax H316 to 21 GBOffloads
Comfy runs it on an RTX 3060 with heavy offloading. 64 GB of RAM.
- Flux.2 Devfp8 · 35.5 GBNo
Only with 64 GB of RAM, and slowly.
Set up ComfyUI.
Install or update ComfyUI
ComfyUI Desktop or the portable NVIDIA build with a current driver. The RTX 5070 needs a PyTorch built for CUDA 12.8 or newer; the current portable build ships CUDA 13.0.
Leave the memory flags alone
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.--reserve-vram 1keeps 1 GB free for Windows and the browser. Add it if the desktop stutters while ComfyUI renders.--vram-headroom 1asks Dynamic VRAM to keep extra room free, counting what other apps use.--fast-diskprefers streaming weights from disk over RAM. ComfyUI says it can be faster with a fast NVMe SSD.--disable-dynamic-vrambrings back the old loader, for a custom node that breaks with the new one. ComfyUI says this flag will be removed soon.
Make it fail fast instead of crawling
On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the
python.exeComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.Pagefile and RAM
Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with
The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.With 16 GB of RAM, the models marked Fits run well; the rest want 32 GB.
RTX 5070: nvfp4 changes the list.
The 5070 computes nvfp4, a 4-bit format Comfy-Org now ships for several models [4]. On 12 GB that moves models from tight to comfortable:
- Model7.7 GBDownload
krea2_turbo_nvfp4.safetensorsComfyUI/models/diffusion_models/ - Model4.5 GBDownload
z_image_turbo_nvfp4.safetensorsComfyUI/models/diffusion_models/
The text encoders and VAEs are the same as for the fp8 versions. On an RTX 4070 or 3080, ComfyUI emulates nvfp4, so stay with fp8 or int8 there.
How long it takes.
| Card | Model and settings | Time |
|---|---|---|
| RTX 3080 Ti | SDXL, 1024 px, 20 steps | 5.6 s[1] |
| RTX 4070 | SDXL, 1024 px, 20 steps | 7.1 s[1] |
| RTX 3060 12 GB | Flux.2 Klein 4B fp8, 1024 px, 4 stepsThe slowest 12 GB card, as a floor. | 9.2 s[3] |
| RTX 3060 12 GB | Krea 2 Turbo fp8, 1280 × 720, 6 steps | 28.4 s[2] |
There are no clean published numbers for Flux.1 Dev or Wan 2.2 on a 4070 or 5070 yet. They sit between the 3060 12 GB and the 16 GB cards; how far depends on how much offloads.
Video on 12 GB.
- Wan 2.2 5B fits in fp16 at its native 1280 × 704.
- Wan 2.2 14B runs as two fp8 halves with the 4-step LoRAs at 832 × 480, with 32 GB of RAM. At 1280 × 720, Wan’s own script ran out of memory on a 12 GB card before the first frame [5].
- HunyuanVideo 1.5 fits in fp8 at 480p.
The RTX 3080 10 GB.
The original 3080 has 10 GB. It’s fast (760 GB/s) and sits between the tiers: everything on the 8 GB list fits with room, and most of the 12 GB models marked Fits here fit too. Krea 2 and Flux.1 Dev in fp8 offload.
12 or 16 GB?
If you’re buying, 16 GB fits Krea 2, Flux.1 Dev and HunyuanVideo at 720p without offloading. An RTX 5060 Ti 16 GB is slower than a 5070 whenever a model fits in 12 GB, and faster when it doesn’t.
When it goes wrong.
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
Questions.
Can a 12 GB card run Flux Dev?
Flux.1 Dev, yes: in fp8 or as a GGUF Q6 to Q8, with 32 GB of system RAM. Flux.2 Dev needs 64 GB of RAM and patience on 12 GB.
Can I run Wan 2.2 14B on 12 GB?
Yes, at 832 × 480 with the two fp8 halves and the 4-step lightx2v LoRAs, and 32 GB of RAM. Wan 2.2 5B fits outright.
RTX 5070 or RTX 5060 Ti 16 GB for ComfyUI?
The 5070 is faster on anything that fits in 12 GB. The 5060 Ti 16 GB fits more: Krea 2 in fp8, Flux.1 Dev and 720p video without offloading. With nvfp4 files the 5070 closes much of that gap.
Is the RTX 3080 10 GB still good for this?
Yes for SDXL, Flux.2 Klein 4B in fp8, Z-Image and Wan 2.2 5B, and it’s quick. Krea 2 and Flux.1 Dev in fp8 offload on 10 GB.
Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] Krea 2 Turbo on an RTX 3060 12 GB, MediaPixel, [3] Flux.2 Klein and MageFlow on an RTX 3060 12 GB, MediaPixel, [4] Krea 2 files, Comfy-Org on Hugging Face, [5] Wan 2.2 A14B on an RTX 3060 12 GB, Wan2.2 issue #144, [6] Dynamic VRAM in ComfyUI, Comfy blog, [7] Steam Hardware Survey, August 2026, [8] System memory fallback, NVIDIA.