Short answer
An RTX 3060 12 GB runs SDXL, Flux.2 Klein 4B at full precision, Z-Image, Qwen-Image 2.1 and Wan 2.2 5B from its own memory, and Krea 2 Turbo in fp8 just fits: 28 seconds at 1280 × 720 in one test. Wan 2.2 14B and the 20 GB models run with part of the model in system RAM.
It’s a slow card with a lot of room. Give it 32 GB of system RAM, or big models load for minutes.
What fits in 12 GB.
Fits runs from the card, Tight fits at the default size, Offloads keeps part of the model in system RAM and streams it in: slower, but it works. Sizes are the files to download.
- SDXLfp16 · 6.9 GBFits
With room for LoRAs and a second image in the batch.
- Flux.2 Klein 4Bbf16 · 7.8 GBFits
Full precision, 4 steps. The fp8 file (4.1 GB) takes about 9 s.
- MageFlow Turboint8 · 4.2 GBFits
About 5 s at 1024 px.
- Z-Image Turboint8 · 6.2 GBFits
The bf16 file (12.3 GB) is tight.
- Qwen-Image 2.1int8 · 7.3 GBFits
Its 9.4 GB text encoder runs first and then makes room.
- Chromafp8 · 9.2 GBFits
Or a GGUF Q5 to Q8.
- Flux.2 Klein 9BGGUF Q8_0 · 10 GBTight
Or Q5_K_M (7 GB) for more room.
- Krea 2 Turbofp8 · 13.1 GBTight
Peaked at 11.8 GB at 1920 × 1080 in one test, with 64 GB of RAM behind it.
- Flux.1 Devfp8 · about 12 GBTight
Runs, but needs 32 GB of RAM to load in reasonable time.
- Wan 2.2 5Bfp16 · 10 GBFits
1280 × 704 is its native size; shorter clips at first.
- HunyuanVideo 1.5fp8 · 8.3 GBFits
480p. 720p offloads.
- Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads
832 × 480 with the 4-step LoRAs. Not 720p.
- Ideogram 4fp8 · 2 × 9.3 GBOffloads
It runs two models side by side.
- Qwen-Image, HiDream, SD 3.5 Large15 to 20 GBOffloads
Qwen-Image as a GGUF Q4_K_M (13.1 GB) is the gentlest way in.
- MiniMax H316 to 21 GBOffloads
Comfy says H3 runs “on a GPU like the RTX 3060” with its offloading [8]. Give it 64 GB of RAM and time.
- Flux.2 Devfp8 · 35.5 GBNo
Plus an 18 GB text encoder. Only with 64 GB of RAM, and slowly.
Set up ComfyUI for a 3060.
Install or update ComfyUI
ComfyUI Desktop or the portable NVIDIA build (Python 3.13, CUDA 13.0, 20 series and newer). Update the NVIDIA driver first.
Keep the default memory handling
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.On a 3060, add nothing unless the desktop stutters during a render. Then add
--reserve-vram 1torun_nvidia_gpu.bat, or set Reserved VRAM (GB) in ComfyUI Desktop’s server settings.run_nvidia_gpu.bat.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --reserve-vram 1 pause
Stop the silent slowdown
On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the
python.exeComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.Pagefile and RAM
Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with
The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.With 16 GB of RAM, stick to the models marked Fits. Anything bigger wants 32 GB.
Files for the best start.
Flux.2 Klein 4B at full precision is what 12 GB buys over 8 GB. These are the files from Comfy-Org’s template:
- Model7.8 GBDownload
flux-2-klein-4b.safetensorsComfyUI/models/diffusion_models/ - Text encoder8.0 GBDownload
qwen_3_4b.safetensorsComfyUI/models/text_encoders/ - VAE0.3 GBDownload
flux2-vae.safetensorsComfyUI/models/vae/
For Krea 2 Turbo, the fp8 model with its fp8 text encoder:
- Model13.1 GBDownload
krea2_turbo_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ - Text encoder5.2 GBDownload
qwen3vl_4b_fp8_scaled.safetensorsComfyUI/models/text_encoders/ - VAE0.3 GBDownload
qwen_image_vae.safetensorsComfyUI/models/vae/
How long it takes.
| Model | Settings | Time |
|---|---|---|
| MageFlow Turbo | int8, 1024 px, 4 steps | 5.1 s[1] |
| Flux.2 Klein 4B | fp8, 1024 px, 4 steps | 9.2 s[1] |
| Krea 2 Turbo | fp8, 1280 × 720, 6 steps | 28.4 s[2] |
| Krea 2 Turbo | fp8, 1920 × 1080, 8 steps | 88.2 s[2] |
Medians and averages from the testers, Windows, 64 GB of RAM. Both peaked just under 12 GB of VRAM.
There’s no clean SDXL number for the 3060 12 GB in the big benchmark threads. The cards around it: an RTX 3060 Ti takes 13.3 s and an RTX 4060 15.9 s for SDXL at 1024 px and 20 steps [3]. The 3060 has fewer cores than both, so expect a little longer. For video there’s no reliable 2026 number either; one 3060 owner’s Wan 2.1 clip at about 720p and 81 frames took around two and a half hours [6], which is why the settings below stay at 480p and 4 steps.
Video on a 3060.
- Wan 2.2 5B fits in fp16. Start at 832 × 480 and 49 frames, then go up.
- Wan 2.2 14B runs as two fp8 halves with the 4-step lightx2v LoRAs, at 832 × 480. It needs 32 GB of RAM; the halves are 14.3 GB each.
- Not 720p with the 14B. Wan’s own script ran out of memory at 1280 × 720 on a 3060 12 GB, in the text encoder, before it rendered a frame [5].
- HunyuanVideo 1.5 in fp8 fits at 480p.
Give it 32 GB of RAM.
This is the 3060’s real limit. A 3060 owner with 16 GB of RAM waited 10 to 20 minutes for Flux.1 Dev in fp8 to load, with Windows frozen and the mouse stuttering, while SDXL ran fine on the same PC [4]. The card had room; the RAM didn’t.
| System RAM | What runs well |
|---|---|
| 16 GB | Everything marked Fits, one model at a time. |
| 32 GB | Krea 2, Flux.1 Dev, Wan 2.2 14B at 480p, Ideogram 4. |
| 64 GB | Qwen-Image, HiDream, MiniMax H3, longer Wan 2.2 clips. |
Which 3060 you have.
NVIDIA also sold an RTX 3060 with 8 GB, on a narrower bus (240 GB/s instead of 360). It’s slower and has the 8 GB limits. Task Manager › Performance › GPU shows Dedicated GPU memory: 12.0 GB is the one this page is about. With 8 GB, read the 8 GB guide instead.
When it goes wrong.
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
Worth upgrading?
Against an RTX 3060 Ti, the 12 GB card fits more and the Ti is faster on what fits, so a sideways move doesn’t pay. The step that changes what runs is 16 GB: an RTX 5060 Ti 16 GB fits Krea 2 with room, computes fp8 and nvfp4 natively, and runs Flux.1 Dev fp8 in about 26 seconds with SageAttention [9]. Before buying a card, 32 GB of RAM is the cheaper fix.
Questions.
Can the RTX 3060 12 GB run Flux?
Yes. Flux.2 Klein 4B runs at full precision, in about 9 seconds in fp8. Flux.1 Dev runs in fp8 or as a GGUF Q8, as long as the computer has 32 GB of RAM. Flux.2 Dev is too big without 64 GB of RAM.
Can it run Wan 2.2?
Yes. Wan 2.2 5B fits in fp16. Wan 2.2 14B runs as two fp8 halves with the 4-step LoRAs at 832 × 480, with 32 GB of RAM. Skip 720p on the 14B.
Is 16 GB of RAM enough?
For SDXL, Flux.2 Klein, Z-Image and Qwen-Image 2.1, yes. For Flux.1 Dev, Krea 2 and Wan 2.2 14B, get 32 GB. With 16 GB, one owner waited 10 to 20 minutes for Flux to load.
Why is my 3060 slower than the benchmarks?
Check it’s the 12 GB model, that the driver is current, and that Shared GPU memory in Task Manager stays near zero during a render. If it climbs, the card is full and Windows is lending it RAM. Lower the size or use a smaller file.
Does HEISS UI mark Krea 2 on a 3060?
On 12 GB it marks Flux.2 Klein 4B at full precision as the best fit. Krea 2’s compact version is listed at 16 GB, so it shows its size without a mark, and it still runs, as the test on this page shows.
Sources: [1] Flux.2 Klein and MageFlow on an RTX 3060 12 GB, MediaPixel, [2] Krea 2 Turbo on an RTX 3060 12 GB, MediaPixel, [3] SDXL GPU benchmark thread, ComfyUI discussion #2970, [4] Flux on an RTX 3060 with 16 GB of RAM, ComfyUI issue #12334, [5] Wan 2.2 A14B on an RTX 3060 12 GB, Wan2.2 issue #144, [6] Wan on an RTX 3060, Popular AI, [7] Wan 2.2 in ComfyUI, Comfy docs, [8] MiniMax H3 in ComfyUI, Comfy blog, [9] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [10] Dynamic VRAM in ComfyUI, Comfy blog, [11] System memory fallback, NVIDIA, [12] Steam Hardware Survey, August 2026.