Short answer
SDXL and the new few-step models run entirely on the card: about 13 seconds for an SDXL image, and Flux.2 Klein 4B takes about 9 seconds on its slower sister card. Krea 2, Flux.1 Dev and Wan 2.2 14B run too, with part of the model in system RAM, as long as the computer has 32 GB of it.
The limit is memory, not speed. Whenever a model fits in 8 GB, the 3060 Ti is quicker than an RTX 3060 12 GB.
What fits in 8 GB.
Fits means it runs from the card. Tight means it fits at the default size and can run out at larger sizes or batches. Offloads means ComfyUI keeps part of the model in system RAM and streams it in as it goes: it works, only slower. The sizes are the files to download.
- SDXLfp16 · 6.9 GBFits
The whole checkpoint sits on the card at 1024 px, Pony, Illustrious and NoobAI merges included.
- Flux.2 Klein 4Bfp8 · 4.1 GBFits
4 steps at 1024 px. The Qwen3 4B text encoder runs first and then makes room.
- Z-Image Turboint8 · 6.2 GBFits
8 steps. The bf16 file (12.3 GB) runs too, with offloading.
- MageFlow Turboint8 · 4.2 GBFits
4 steps, and the quickest model here.
- SD 1.5, Anima, Sana2 to 5 GBFits
Small models with room to spare.
- Flux.2 Klein 9BGGUF Q4_K_M · 5.9 GBTight
The model fits. Its Qwen3 8B text encoder (8.7 GB in fp8) loads first and then moves out.
- ChromaGGUF Q4_K_M · 5.6 GBTight
The fp8 file (9.2 GB) runs with offloading.
- Flux.1 DevGGUF Q4_K_S · 6.8 GBTight
Or the fp8 model file (about 12 GB) with offloading. Either way, give it 32 GB of RAM.
- Wan 2.2 5Bfp16 · 10 GB, or GGUFTight
ComfyUI’s docs say it fits 8 GB with its own offloading. Start at 832 × 480 and 33 to 49 frames.
- Krea 2 Turbofp8 · 13.1 GBOffloads
Works with 32 GB of RAM. Its GGUF versions don’t load in ComfyUI-GGUF yet.
- Qwen-Image 2.1int8 · 7.3 GBOffloads
The model nearly fits; with its 9.4 GB text encoder the whole set needs system RAM. No GGUF route yet.
- HunyuanVideo 1.5fp8 · 8.3 GBOffloads
480p clips.
- Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads
Two models, one after the other. Works with 32 GB of RAM and the 4-step LoRAs; 64 GB is smoother.
- Qwen-Image, HiDream, SD 3.5 Large15 to 20 GBOffloads
Slow on 8 GB. Worth it with 64 GB of RAM, not with 16.
- Flux.2 Devfp8 · 35.5 GBNo
Plus an 18 GB text encoder. It needs a 24 to 32 GB card, or 64 GB of RAM and a lot of patience.
- MiniMax H316 to 21 GBNo
Its 32B text encoder alone is 15.7 GB at its smallest. Comfy runs it on a 12 GB RTX 3060 with heavy offloading.
Set up ComfyUI for a 3060 Ti.
Install or update ComfyUI
On Windows, ComfyUI Desktop or the portable build for NVIDIA both work. The portable one ships Python 3.13 and CUDA 13.0 and supports the 20 series and newer. Update the NVIDIA driver first: the portable build won’t start on an old one.
Leave the memory flags alone
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.The one flag worth adding on an 8 GB card is
--reserve-vram 1, if the desktop stutters while ComfyUI renders. In the portable build it goes at the end of the first line ofrun_nvidia_gpu.bat:run_nvidia_gpu.bat.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --reserve-vram 1 pause
In ComfyUI Desktop the same setting is Reserved VRAM (GB) in its server settings.
Make it fail fast instead of crawling
On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the
python.exeComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.Give Windows pagefile room
Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with
The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.Download the files
Pick the precision from the table above: ComfyUI’s own fp8 and int8 files first, a GGUF when the model is too big for your RAM or has no fp8 file. The files for the two best starts are just below.
Expect a slow first run
The first image after starting ComfyUI reads the model from disk and can take a minute or more. The next ones start warm.
Files for the best two starts.
Flux.2 Klein 4B in fp8 is the strongest image model that sits fully on the card. It needs three files:
- Model4.1 GBDownload
flux-2-klein-4b-fp8.safetensorsComfyUI/models/diffusion_models/ - Text encoder8.0 GBDownload
qwen_3_4b.safetensorsComfyUI/models/text_encoders/ - VAE0.3 GBDownload
flux2-vae.safetensorsComfyUI/models/vae/
SDXL is one file with everything inside, and the biggest world of LoRAs:
- Checkpoint6.9 GBDownload
RealVisXL_V5.0_fp16.safetensorsComfyUI/models/checkpoints/
Settings that work.
Starting points that work, from the models’ ComfyUI templates and HEISS UI’s defaults. On 8 GB, keep the batch at one image.
| Model | Steps | CFG | Sampler | Size |
|---|---|---|---|---|
| SDXL | 25 | 7 | dpmpp_2m · karras | 1024 × 1024 |
| Flux.2 Klein 4B | 4 | 1 | euler · simple | 1024 × 1024 |
| Z-Image Turbo | 8 | 1 | res_multistep · simple | 1024 × 1024 |
| Krea 2 Turbo | 8 | 1 | euler · simple | 1024 × 1024 |
| Wan 2.2 5B | 20 | 5 | uni_pc · simple | 832 × 480, 33 to 49 frames |
| Wan 2.2 14B, 4-step | 4 | 1 | euler · simple | 832 × 480, 33 to 81 frames |
How long it takes.
Measured times from people who posted their setup. Nobody has published a clean 3060 Ti number for the newer models yet, so the table borrows from the two cards around it and says so.
| Card | Model and settings | Time |
|---|---|---|
| RTX 3060 Ti | SDXL, 1024 px, 20 steps | 13.3 s[1] |
| RTX 3070 8 GB | SDXL, 1024 px, 20 steps | 10.9 s[1] |
| RTX 3060 12 GB | Flux.2 Klein 4B fp8, 1024 px, 4 stepsMeasured on the 12 GB card; no 3060 Ti test yet. | 9.2 s[2] |
| RTX 3060 12 GB | Krea 2 Turbo fp8, 1280 × 720, 6 stepsFits in 12 GB; on 8 GB it offloads and takes longer. | 28.4 s[11] |
| RTX 4060 8 GB | Wan 2.2 14B, 4 steps, 480 × 480, 33 framesWith 32 GB of RAM. | 111 s[3] |
For Flux.1 Dev and Krea 2 on an 8 GB card there’s no reliable 2026 number yet. Both offload, so how long they take depends as much on your RAM and SSD as on the card.
RAM decides the rest.
Anything marked Offloads lives partly in system RAM. That makes RAM the real limit on an 8 GB card, and 16 GB is what most people have.
| System RAM | What it changes |
|---|---|
| 16 GB | Everything marked Fits runs well. Offloaded models load slowly and lean on the pagefile. An RTX 3060 owner with 16 GB waited 10 to 20 minutes for a 12 GB Flux file to load, with Windows frozen meanwhile [8]. |
| 32 GB | Krea 2, Flux.1 Dev, Qwen-Image 2.1 and Wan 2.2 14B with the 4-step LoRAs become practical. |
| 64 GB | The 20 GB models and longer Wan 2.2 clips run without the pagefile getting involved. |
Flux Dev in fp8 took 10 to 20 minutes to start on a 3060 with 16 GB of RAM. The mouse stuttered and Windows froze. SDXL on the same PC was fine.
Going even a little past the card’s memory into shared GPU memory made every step crawl. Setting Prefer No Sysmem Fallback turned it into a clear error.
When it goes wrong.
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
3060 Ti or 3060 12 GB?
They’re the two most common RTX 30 cards on Steam, and they trade places depending on the model.
| RTX 3060 Ti | RTX 3060 12 GB | |
|---|---|---|
| Graphics memory | 8 GB | 12 GB |
| Bandwidth | 448 GB/s | 360 GB/s |
| CUDA cores | 4864 | 3584 |
| SDXL, 1024 px | 13.3 s[1] | Slower, no clean number |
| Flux.2 Klein 4B | fp8 only | Full precision fits |
| Krea 2 Turbo fp8 | Offloads | Tight 28.4 s[11] |
| Wan 2.2 5B | With offloading | fp16 fits |
Steam share, August 2026: RTX 3060 3.92%, RTX 3060 Ti 2.16%.
For image and video models the 12 GB card fits more, and the Ti is faster on what fits. If you have a 3060 Ti, there’s no reason to swap sideways. The upgrade that changes what runs is 16 GB: see the RTX 5060 Ti 16 GB.
Questions.
Can an RTX 3060 Ti run Flux?
Yes. Flux.2 Klein 4B runs fully on the card in fp8, at 4 steps. Flux.1 Dev runs as a GGUF Q4 or with the fp8 file offloading, which needs 32 GB of system RAM. Flux.2 Dev is too big for it.
Can an RTX 3060 Ti make video?
Yes, short clips. Wan 2.2 5B fits with ComfyUI’s offloading, and Wan 2.2 14B runs with 32 GB of RAM and the 4-step LoRAs. Keep it around 480p and 33 to 81 frames.
Do I still need --lowvram?
No. Since March 2026 ComfyUI manages graphics memory by itself on NVIDIA cards, and --lowvram does nothing while that’s on. Add --reserve-vram 1 only if the desktop stutters during a render.
Is 16 GB of RAM enough for a 3060 Ti?
For the models that fit the card, yes. For anything that offloads, get 32 GB. With 16 GB, a 12 GB Flux file took 10 to 20 minutes to load on an RTX 3060 and froze Windows while it did.
GGUF or fp8 on 8 GB?
Start with ComfyUI’s own fp8 or int8 files. ComfyUI recommends them, and they’re usually faster. Use a GGUF, Q4_K_M to Q8_0, when a model is too big for your RAM or has no fp8 file. The ComfyUI-GGUF pack doesn’t load Krea 2, Ideogram 4, MiniMax H3 or Qwen-Image 2.1 yet.
Why did my renders suddenly get much slower?
The card filled up and Windows started lending it system RAM, which shows as Shared GPU memory in Task Manager. Lower the size, close other apps that use the GPU, and set Prefer No Sysmem Fallback for ComfyUI in the NVIDIA Control Panel so it stops with an error instead.
Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] Flux.2 Klein and MageFlow on an RTX 3060 12 GB, MediaPixel, [3] Wan 2.2 on an RTX 4060 8 GB, lilting channel, [4] Dynamic VRAM in ComfyUI, Comfy blog, [5] ComfyUI launch flags, cli_args.py, [6] Wan 2.2 in ComfyUI, Comfy docs, [7] System memory fallback, NVIDIA, [8] Flux on an RTX 3060 with 16 GB of RAM, ComfyUI issue #12334, [9] Shared GPU memory slowdown, ComfyUI discussion #14092, [10] Steam Hardware Survey, August 2026, [11] Krea 2 Turbo on an RTX 3060 12 GB, MediaPixel, [12] Paging file error on a 16 GB laptop, ComfyUI issue #7550.