Hardware guide

What runs on 12 GB of VRAM.

A 12 GB card runs SDXL, Flux.2 Klein 4B at full precision, Z-Image, Qwen-Image 2.1 and Wan 2.2 5B from its own memory, and Krea 2 Turbo just about. On an RTX 5070, the nvfp4 files add room. This covers the RTX 4070 family, 5070, 3080 12 GB and 3080 Ti.

Updated 29 Sep 20266 min read

Cards
RTX 4070, 50703080 12 GB, 3080 Ti
Steam users
13.0%have 12 GB, Aug 2026
Best start
Flux.2 Klein 4Bfull precision
System RAM
32 GBfor models that offload

Short answer

A 12 GB card runs SDXL, Flux.2 Klein 4B at full precision, Z-Image, Qwen-Image 2.1, Chroma and Wan 2.2 5B from its own memory, and Krea 2 Turbo in fp8 just about. Wan 2.2 14B, Qwen-Image and Flux.2 Dev need system RAM.

On an RTX 5070, the nvfp4 files put Krea 2 Turbo at 7.7 GB, with room to spare.

The 12 GB cards.

This page covers the RTX 4070 family, the 5070 and the RTX 3080 12 GB and 3080 Ti. The RTX 3060 12 GB has its own page.

CardBandwidthComputes nativelySDXL, 1024 px, 20 steps
RTX 5070672 GB/sfp8, nvfp4No published number
RTX 4070 / Super / Ti504 GB/sfp84070: 7.1 s[1]
RTX 3080 12 GB / 3080 Ti912 GB/snone3080 Ti: 5.6 s[1]
RTX 3060 12 GB360 GB/snoneIts own page

The RTX 5070 is the third most common card on Steam in August 2026, at 3.77%.

What fits in 12 GB.

Fits runs from the card. Tight fits at the default size. Offloads keeps part of the model in system RAM: it works, slower.

  • SDXLfp16 · 6.9 GBFits

    With LoRAs and a batch of two.

  • Flux.2 Klein 4Bbf16 · 7.8 GBFits

    Full precision, 4 steps.

  • Flux.2 Klein 9BGGUF Q8_0 · 10 GBTight

    Or Q5_K_M (7 GB). Comfy measures 19.6 GB at bf16.

  • Z-Image Turboint8 · 6.2 GBFits

    The bf16 file (12.3 GB) is tight; nvfp4 (4.5 GB) on a 5070.

  • Qwen-Image 2.1int8 · 7.3 GBFits

    The text encoder runs first and then makes room.

  • Chromafp8 · 9.2 GBFits

    Or a GGUF Q8_0 (9.7 GB).

  • MageFlowint8 · 4.2 GBFits

    4 steps with the Turbo version.

  • Krea 2 Turbofp8 13.1 GB · nvfp4 7.7 GBTight

    fp8 peaked at 11.8 GB at 1080p on a 3060 12 GB. On a 5070, the nvfp4 file fits easily.

  • Flux.1 Devfp8 · about 12 GBTight

    Or GGUF Q8_0 (12.7 GB) / Q6_K (9.9 GB). 32 GB of RAM.

  • Ideogram 4fp8 2 × 9.3 GB · nvfp4 2 × 5.5 GBOffloads

    Two models side by side; on a 5070 the nvfp4 pair nearly fits.

  • Wan 2.2 5Bfp16 · 10 GBFits

    Native 1280 × 704. Start with shorter clips.

  • HunyuanVideo 1.5fp8 · 8.3 GBFits

    480p; 720p offloads.

  • Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads

    832 × 480 with the 4-step LoRAs and 32 GB of RAM.

  • Qwen-Image, HiDream, SD 3.5 Large, ERNIE15 to 20 GBOffloads

    Qwen-Image as a GGUF Q4_K_M (13.1 GB) is the easiest of these.

  • MiniMax H316 to 21 GBOffloads

    Comfy runs it on an RTX 3060 with heavy offloading. 64 GB of RAM.

  • Flux.2 Devfp8 · 35.5 GBNo

    Only with 64 GB of RAM, and slowly.

Set up ComfyUI.

  1. Install or update ComfyUI

    ComfyUI Desktop or the portable NVIDIA build with a current driver. The RTX 5070 needs a PyTorch built for CUDA 12.8 or newer; the current portable build ships CUDA 13.0.

  2. Leave the memory flags alone

    Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).

    That changes the old advice. --lowvram now does nothing while Dynamic VRAM is on, and --novram, --highvram and --gpu-only switch it off. --normalvram is gone, and an old shortcut that still passes it stops ComfyUI with unrecognized arguments.

    • --reserve-vram 1 keeps 1 GB free for Windows and the browser. Add it if the desktop stutters while ComfyUI renders.
    • --vram-headroom 1 asks Dynamic VRAM to keep extra room free, counting what other apps use.
    • --fast-disk prefers streaming weights from disk over RAM. ComfyUI says it can be faster with a fast NVMe SSD.
    • --disable-dynamic-vram brings back the old loader, for a custom node that breaks with the new one. ComfyUI says this flag will be removed soon.
  3. Make it fail fast instead of crawling

    On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the python.exe ComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.

  4. Pagefile and RAM

    Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.

    With 16 GB of RAM, the models marked Fits run well; the rest want 32 GB.

RTX 5070: nvfp4 changes the list.

The 5070 computes nvfp4, a 4-bit format Comfy-Org now ships for several models [4]. On 12 GB that moves models from tight to comfortable:

  • Modelkrea2_turbo_nvfp4.safetensorsComfyUI/models/diffusion_models/
    7.7 GBDownload
  • Modelz_image_turbo_nvfp4.safetensorsComfyUI/models/diffusion_models/
    4.5 GBDownload

The text encoders and VAEs are the same as for the fp8 versions. On an RTX 4070 or 3080, ComfyUI emulates nvfp4, so stay with fp8 or int8 there.

How long it takes.

CardModel and settingsTime
RTX 3080 TiSDXL, 1024 px, 20 steps5.6 s[1]
RTX 4070SDXL, 1024 px, 20 steps7.1 s[1]
RTX 3060 12 GBFlux.2 Klein 4B fp8, 1024 px, 4 stepsThe slowest 12 GB card, as a floor.9.2 s[3]
RTX 3060 12 GBKrea 2 Turbo fp8, 1280 × 720, 6 steps28.4 s[2]

There are no clean published numbers for Flux.1 Dev or Wan 2.2 on a 4070 or 5070 yet. They sit between the 3060 12 GB and the 16 GB cards; how far depends on how much offloads.

Video on 12 GB.

  • Wan 2.2 5B fits in fp16 at its native 1280 × 704.
  • Wan 2.2 14B runs as two fp8 halves with the 4-step LoRAs at 832 × 480, with 32 GB of RAM. At 1280 × 720, Wan’s own script ran out of memory on a 12 GB card before the first frame [5].
  • HunyuanVideo 1.5 fits in fp8 at 480p.

The RTX 3080 10 GB.

The original 3080 has 10 GB. It’s fast (760 GB/s) and sits between the tiers: everything on the 8 GB list fits with room, and most of the 12 GB models marked Fits here fit too. Krea 2 and Flux.1 Dev in fp8 offload.

12 or 16 GB?

If you’re buying, 16 GB fits Krea 2, Flux.1 Dev and HunyuanVideo at 720p without offloading. An RTX 5060 Ti 16 GB is slower than a 5070 whenever a model fits in 12 GB, and faster when it doesn’t.

When it goes wrong.

torch.OutOfMemoryError: CUDA out of memory
The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)
Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
It suddenly got much slower
The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram
An old launch flag. Remove it from the shortcut or the .bat file. Current ComfyUI doesn’t need memory flags on NVIDIA.
The first image takes minutes, the next ones don’t
The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.

Questions.

Can a 12 GB card run Flux Dev?

Flux.1 Dev, yes: in fp8 or as a GGUF Q6 to Q8, with 32 GB of system RAM. Flux.2 Dev needs 64 GB of RAM and patience on 12 GB.

Can I run Wan 2.2 14B on 12 GB?

Yes, at 832 × 480 with the two fp8 halves and the 4-step lightx2v LoRAs, and 32 GB of RAM. Wan 2.2 5B fits outright.

RTX 5070 or RTX 5060 Ti 16 GB for ComfyUI?

The 5070 is faster on anything that fits in 12 GB. The 5060 Ti 16 GB fits more: Krea 2 in fp8, Flux.1 Dev and 720p video without offloading. With nvfp4 files the 5070 closes much of that gap.

Is the RTX 3080 10 GB still good for this?

Yes for SDXL, Flux.2 Klein 4B in fp8, Z-Image and Wan 2.2 5B, and it’s quick. Krea 2 and Flux.1 Dev in fp8 offload on 10 GB.

Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] Krea 2 Turbo on an RTX 3060 12 GB, MediaPixel, [3] Flux.2 Klein and MageFlow on an RTX 3060 12 GB, MediaPixel, [4] Krea 2 files, Comfy-Org on Hugging Face, [5] Wan 2.2 A14B on an RTX 3060 12 GB, Wan2.2 issue #144, [6] Dynamic VRAM in ComfyUI, Comfy blog, [7] Steam Hardware Survey, August 2026, [8] System memory fallback, NVIDIA.

HEISS UI

Sized for 12 GB.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. It reads your card and does the fitting for you.

  • The model that fits your 12 GB is marked. Pick a first model and the size that suits this card is already chosen. One tap downloads it with everything it needs.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.