Hardware guide

What runs on an RTX 3060 Ti.

An RTX 3060 Ti has 8 GB of graphics memory. That runs SDXL, Flux.2 Klein 4B, Z-Image Turbo and Wan 2.2 5B well in ComfyUI, and Krea 2, Flux.1 Dev and Wan 2.2 14B with part of the model in system RAM. Here’s the setup, the files for each model, measured times, and what to do when something doesn’t fit.

Updated 29 Sep 20268 min read

Memory
8 GB GDDR6GDDR6X on later cards
Bandwidth
448 GB/s608 GB/s with GDDR6X
Architecture
Ampereno fp8 or fp4 compute
System RAM
32 GBfor models that offload

Short answer

SDXL and the new few-step models run entirely on the card: about 13 seconds for an SDXL image, and Flux.2 Klein 4B takes about 9 seconds on its slower sister card. Krea 2, Flux.1 Dev and Wan 2.2 14B run too, with part of the model in system RAM, as long as the computer has 32 GB of it.

The limit is memory, not speed. Whenever a model fits in 8 GB, the 3060 Ti is quicker than an RTX 3060 12 GB.

What fits in 8 GB.

Fits means it runs from the card. Tight means it fits at the default size and can run out at larger sizes or batches. Offloads means ComfyUI keeps part of the model in system RAM and streams it in as it goes: it works, only slower. The sizes are the files to download.

  • SDXLfp16 · 6.9 GBFits

    The whole checkpoint sits on the card at 1024 px, Pony, Illustrious and NoobAI merges included.

  • Flux.2 Klein 4Bfp8 · 4.1 GBFits

    4 steps at 1024 px. The Qwen3 4B text encoder runs first and then makes room.

  • Z-Image Turboint8 · 6.2 GBFits

    8 steps. The bf16 file (12.3 GB) runs too, with offloading.

  • MageFlow Turboint8 · 4.2 GBFits

    4 steps, and the quickest model here.

  • SD 1.5, Anima, Sana2 to 5 GBFits

    Small models with room to spare.

  • Flux.2 Klein 9BGGUF Q4_K_M · 5.9 GBTight

    The model fits. Its Qwen3 8B text encoder (8.7 GB in fp8) loads first and then moves out.

  • ChromaGGUF Q4_K_M · 5.6 GBTight

    The fp8 file (9.2 GB) runs with offloading.

  • Flux.1 DevGGUF Q4_K_S · 6.8 GBTight

    Or the fp8 model file (about 12 GB) with offloading. Either way, give it 32 GB of RAM.

  • Wan 2.2 5Bfp16 · 10 GB, or GGUFTight

    ComfyUI’s docs say it fits 8 GB with its own offloading. Start at 832 × 480 and 33 to 49 frames.

  • Krea 2 Turbofp8 · 13.1 GBOffloads

    Works with 32 GB of RAM. Its GGUF versions don’t load in ComfyUI-GGUF yet.

  • Qwen-Image 2.1int8 · 7.3 GBOffloads

    The model nearly fits; with its 9.4 GB text encoder the whole set needs system RAM. No GGUF route yet.

  • HunyuanVideo 1.5fp8 · 8.3 GBOffloads

    480p clips.

  • Wan 2.2 14Bfp8 · 2 × 14.3 GBOffloads

    Two models, one after the other. Works with 32 GB of RAM and the 4-step LoRAs; 64 GB is smoother.

  • Qwen-Image, HiDream, SD 3.5 Large15 to 20 GBOffloads

    Slow on 8 GB. Worth it with 64 GB of RAM, not with 16.

  • Flux.2 Devfp8 · 35.5 GBNo

    Plus an 18 GB text encoder. It needs a 24 to 32 GB card, or 64 GB of RAM and a lot of patience.

  • MiniMax H316 to 21 GBNo

    Its 32B text encoder alone is 15.7 GB at its smallest. Comfy runs it on a 12 GB RTX 3060 with heavy offloading.

Set up ComfyUI for a 3060 Ti.

  1. Install or update ComfyUI

    On Windows, ComfyUI Desktop or the portable build for NVIDIA both work. The portable one ships Python 3.13 and CUDA 13.0 and supports the 20 series and newer. Update the NVIDIA driver first: the portable build won’t start on an old one.

  2. Leave the memory flags alone

    Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).

    That changes the old advice. --lowvram now does nothing while Dynamic VRAM is on, and --novram, --highvram and --gpu-only switch it off. --normalvram is gone, and an old shortcut that still passes it stops ComfyUI with unrecognized arguments.

    The one flag worth adding on an 8 GB card is --reserve-vram 1, if the desktop stutters while ComfyUI renders. In the portable build it goes at the end of the first line of run_nvidia_gpu.bat:

    run_nvidia_gpu.bat
    .\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --reserve-vram 1
    pause

    In ComfyUI Desktop the same setting is Reserved VRAM (GB) in its server settings.

  3. Make it fail fast instead of crawling

    On Windows, the NVIDIA driver can quietly borrow system RAM when the card is full. Task Manager shows it as Shared GPU memory, and a render that would have stopped with an error crawls on, many times slower, instead. To get the error instead: NVIDIA Control Panel › Manage 3D settings › Program Settings, add the python.exe ComfyUI runs with, and set CUDA - Sysmem Fallback Policy to Prefer No Sysmem Fallback. Restart ComfyUI.

  4. Give Windows pagefile room

    Models that don’t fit the card pass through system RAM, and Windows needs pagefile room while it loads them. Leave the pagefile on System managed size on your fastest SSD. If ComfyUI stops with The paging file is too small for this operation to complete, set a custom size of 32 GB or more under System › About › Advanced system settings › Performance › Advanced › Virtual memory.

  5. Download the files

    Pick the precision from the table above: ComfyUI’s own fp8 and int8 files first, a GGUF when the model is too big for your RAM or has no fp8 file. The files for the two best starts are just below.

  6. Expect a slow first run

    The first image after starting ComfyUI reads the model from disk and can take a minute or more. The next ones start warm.

Files for the best two starts.

Flux.2 Klein 4B in fp8 is the strongest image model that sits fully on the card. It needs three files:

  • Modelflux-2-klein-4b-fp8.safetensorsComfyUI/models/diffusion_models/
    4.1 GBDownload
  • Text encoderqwen_3_4b.safetensorsComfyUI/models/text_encoders/
    8.0 GBDownload
  • VAEflux2-vae.safetensorsComfyUI/models/vae/
    0.3 GBDownload

SDXL is one file with everything inside, and the biggest world of LoRAs:

  • CheckpointRealVisXL_V5.0_fp16.safetensorsComfyUI/models/checkpoints/
    6.9 GBDownload

Settings that work.

Starting points that work, from the models’ ComfyUI templates and HEISS UI’s defaults. On 8 GB, keep the batch at one image.

ModelStepsCFGSamplerSize
SDXL257dpmpp_2m · karras1024 × 1024
Flux.2 Klein 4B41euler · simple1024 × 1024
Z-Image Turbo81res_multistep · simple1024 × 1024
Krea 2 Turbo81euler · simple1024 × 1024
Wan 2.2 5B205uni_pc · simple832 × 480, 33 to 49 frames
Wan 2.2 14B, 4-step41euler · simple832 × 480, 33 to 81 frames

How long it takes.

Measured times from people who posted their setup. Nobody has published a clean 3060 Ti number for the newer models yet, so the table borrows from the two cards around it and says so.

CardModel and settingsTime
RTX 3060 TiSDXL, 1024 px, 20 steps13.3 s[1]
RTX 3070 8 GBSDXL, 1024 px, 20 steps10.9 s[1]
RTX 3060 12 GBFlux.2 Klein 4B fp8, 1024 px, 4 stepsMeasured on the 12 GB card; no 3060 Ti test yet.9.2 s[2]
RTX 3060 12 GBKrea 2 Turbo fp8, 1280 × 720, 6 stepsFits in 12 GB; on 8 GB it offloads and takes longer.28.4 s[11]
RTX 4060 8 GBWan 2.2 14B, 4 steps, 480 × 480, 33 framesWith 32 GB of RAM.111 s[3]

For Flux.1 Dev and Krea 2 on an 8 GB card there’s no reliable 2026 number yet. Both offload, so how long they take depends as much on your RAM and SSD as on the card.

RAM decides the rest.

Anything marked Offloads lives partly in system RAM. That makes RAM the real limit on an 8 GB card, and 16 GB is what most people have.

System RAMWhat it changes
16 GBEverything marked Fits runs well. Offloaded models load slowly and lean on the pagefile. An RTX 3060 owner with 16 GB waited 10 to 20 minutes for a 12 GB Flux file to load, with Windows frozen meanwhile [8].
32 GBKrea 2, Flux.1 Dev, Qwen-Image 2.1 and Wan 2.2 14B with the 4-step LoRAs become practical.
64 GBThe 20 GB models and longer Wan 2.2 clips run without the pagefile getting involved.

Flux Dev in fp8 took 10 to 20 minutes to start on a 3060 with 16 GB of RAM. The mouse stuttered and Windows froze. SDXL on the same PC was fine.

ComfyUI issue #12334

Going even a little past the card’s memory into shared GPU memory made every step crawl. Setting Prefer No Sysmem Fallback turned it into a clear error.

ComfyUI discussion #14092

When it goes wrong.

torch.OutOfMemoryError: CUDA out of memory
The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)
Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
It suddenly got much slower
The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram
An old launch flag. Remove it from the shortcut or the .bat file. Current ComfyUI doesn’t need memory flags on NVIDIA.
The first image takes minutes, the next ones don’t
The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.

3060 Ti or 3060 12 GB?

They’re the two most common RTX 30 cards on Steam, and they trade places depending on the model.

RTX 3060 TiRTX 3060 12 GB
Graphics memory8 GB12 GB
Bandwidth448 GB/s360 GB/s
CUDA cores48643584
SDXL, 1024 px13.3 s[1]Slower, no clean number
Flux.2 Klein 4Bfp8 onlyFull precision fits
Krea 2 Turbo fp8OffloadsTight 28.4 s[11]
Wan 2.2 5BWith offloadingfp16 fits

Steam share, August 2026: RTX 3060 3.92%, RTX 3060 Ti 2.16%.

For image and video models the 12 GB card fits more, and the Ti is faster on what fits. If you have a 3060 Ti, there’s no reason to swap sideways. The upgrade that changes what runs is 16 GB: see the RTX 5060 Ti 16 GB.

Questions.

Can an RTX 3060 Ti run Flux?

Yes. Flux.2 Klein 4B runs fully on the card in fp8, at 4 steps. Flux.1 Dev runs as a GGUF Q4 or with the fp8 file offloading, which needs 32 GB of system RAM. Flux.2 Dev is too big for it.

Can an RTX 3060 Ti make video?

Yes, short clips. Wan 2.2 5B fits with ComfyUI’s offloading, and Wan 2.2 14B runs with 32 GB of RAM and the 4-step LoRAs. Keep it around 480p and 33 to 81 frames.

Do I still need --lowvram?

No. Since March 2026 ComfyUI manages graphics memory by itself on NVIDIA cards, and --lowvram does nothing while that’s on. Add --reserve-vram 1 only if the desktop stutters during a render.

Is 16 GB of RAM enough for a 3060 Ti?

For the models that fit the card, yes. For anything that offloads, get 32 GB. With 16 GB, a 12 GB Flux file took 10 to 20 minutes to load on an RTX 3060 and froze Windows while it did.

GGUF or fp8 on 8 GB?

Start with ComfyUI’s own fp8 or int8 files. ComfyUI recommends them, and they’re usually faster. Use a GGUF, Q4_K_M to Q8_0, when a model is too big for your RAM or has no fp8 file. The ComfyUI-GGUF pack doesn’t load Krea 2, Ideogram 4, MiniMax H3 or Qwen-Image 2.1 yet.

Why did my renders suddenly get much slower?

The card filled up and Windows started lending it system RAM, which shows as Shared GPU memory in Task Manager. Lower the size, close other apps that use the GPU, and set Prefer No Sysmem Fallback for ComfyUI in the NVIDIA Control Panel so it stops with an error instead.

Sources: [1] SDXL GPU benchmark thread, ComfyUI discussion #2970, [2] Flux.2 Klein and MageFlow on an RTX 3060 12 GB, MediaPixel, [3] Wan 2.2 on an RTX 4060 8 GB, lilting channel, [4] Dynamic VRAM in ComfyUI, Comfy blog, [5] ComfyUI launch flags, cli_args.py, [6] Wan 2.2 in ComfyUI, Comfy docs, [7] System memory fallback, NVIDIA, [8] Flux on an RTX 3060 with 16 GB of RAM, ComfyUI issue #12334, [9] Shared GPU memory slowdown, ComfyUI discussion #14092, [10] Steam Hardware Survey, August 2026, [11] Krea 2 Turbo on an RTX 3060 12 GB, MediaPixel, [12] Paging file error on a 16 GB laptop, ComfyUI issue #7550.

HEISS UI

Sized for your 3060 Ti.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. It reads your card and does the fitting for you.

  • The model that fits your 8 GB is marked. Pick a first model and the size that suits this card is already chosen. One tap downloads it with everything it needs.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.