Image model

How to run SDXL locally.

SDXL runs in ComfyUI from one 6.9 GB checkpoint that already holds its text encoders and VAE. It fits an 8 GB graphics card, runs on 6 GB with some help from system RAM, and on a Mac with 16 GB. Most people run a fine-tune of it: RealVisXL for photos, Pony or Illustrious for anime, and Lightning or DMD2 versions when speed matters.

Updated 29 Sep 202614 min read

Maker
Stability AI
Released
Jul 202326 July. Turbo in Nov 2023
Licence
OpenRAIL++-MTurbo, DMD2 non-commercial. Pony, Illustrious: Fair AI
Memory
8 GB6 GB with offloading. Mac 16 GB

Base, fine-tunes, speed files.

SDXL 1.0 came out on 26 July 2023. It’s a 3.5B model with two text encoders, CLIP-L and OpenCLIP bigG, and it makes images at about one megapixel. It has a huge pool of fine-tunes, LoRAs and ControlNets, and it makes 1024-pixel images on modest hardware.

Every file below has the same architecture and loads the same way. What changes is the settings.

  • SDXL 1.0 base. Stability’s original. Fine for testing, plain next to what came after.
  • Photo fine-tunes such as RealVisXL V5.0. Same size, much better photos. Its author allows commercial use.
  • Anime fine-tunes. Pony Diffusion V6 XL, Illustrious XL (September 2024) and NoobAI-XL, which builds on Illustrious. Thousands of merges on Civitai descend from these three. Each wants its own prompt style and settings.
  • Speed files. SDXL Turbo, SDXL-Lightning, Hyper-SD, DMD2 and LCM. They come as full checkpoints or as LoRAs, and cut an image from 25 steps to between 1 and 8. Many fine-tunes ship a Lightning or DMD2 version, like RealVisXL V5.0 Lightning.
  • V-prediction files such as NoobAI-XL V-Pred. They need one extra node, covered below.

Files you need.

One file. An SDXL checkpoint holds the model, both text encoders and the VAE, and goes in models/checkpoints. Pick one of these to start:

  • Photo RealVisXL_V5.0_fp16.safetensors ComfyUI/models/checkpoints/
    6.9 GB Download
  • Photo, fast RealVisXL_V5.0_Lightning_fp16.safetensors ComfyUI/models/checkpoints/ five steps instead of thirty
    6.9 GB Download
  • Base sd_xl_base_1.0.safetensors ComfyUI/models/checkpoints/
    6.9 GB Download
  • Anime Illustrious-XL-v0.1.safetensors ComfyUI/models/checkpoints/ Pony V6 is on Civitai only; most people use a merge of either
    6.9 GB Download
  • Anime, v-pred NoobAI-XL-Vpred-v1.0.safetensors ComfyUI/models/checkpoints/
    7.1 GB Download
ComfyUI/models
models/
├── checkpoints/
│   └── RealVisXL_V5.0_fp16.safetensors
├── loras/
│   └── sdxl_lightning_4step_lora.safetensors   (optional)
└── vae/
    └── sdxl_vae.safetensors                    (optional)

Speed LoRAs

Instead of a separate Lightning checkpoint, you can add a speed LoRA to any SDXL model. They go in models/loras.

  • LoRA sdxl_lightning_4step_lora.safetensors ComfyUI/models/loras/ 2- and 8-step versions in the same repo
    0.4 GB Download
  • LoRA dmd2_sdxl_4step_lora_fp16.safetensors ComfyUI/models/loras/ non-commercial licence
    0.4 GB Download

The VAE and black images

The VAE is baked into every checkpoint, so you don’t need a separate one. The well-known SDXL black image comes from the original VAE overflowing in fp16. ComfyUI avoids that on its own: it runs the SDXL VAE in bf16 where the GPU supports it and in fp32 otherwise, never in fp16 unless you start it with --fp16-vae. If you use that flag, or come from AUTOMATIC1111 with the same problem, load madebyollin’s fp16-fix VAE (0.3 GB, MIT) with a Load VAE node.

  • VAE sdxl_vae.safetensors (fp16 fix) ComfyUI/models/vae/ only needed with --fp16-vae
    0.3 GB Download

fp8 and GGUF

Not needed. The 6.9 GB fp16 file already fits 8 GB. Starting ComfyUI with --fast or --fp8_e4m3fn-unet turns most SDXL fine-tunes into black images, so leave those off for SDXL. GGUF versions exist but are rare, and there is no official one.

What fits your computer.

An SDXL checkpoint at 1024 × 1024, 25 steps. Stability’s own figure at launch was 8 GB of video memory. ComfyUI moves parts of the model to system RAM when the card is smaller, and falls back to a tiled VAE decode when memory runs out at the end.

  • 6 GBTight

    Works with offloading, as on an RTX 2060 6 GB. A Lightning or DMD2 file keeps the wait short.

  • 8 GBFits

    Any SDXL checkpoint, plus a LoRA or two. An RTX 4060 8 GB makes a 20-step image in about 16 seconds.

  • 12 GBFits

    Room for LoRAs, a ControlNet and a hires pass on top.

  • 16 GBFits

    Everything above, with room for batches.

  • 24 GBFits

    Base and refiner together, larger batches and upscales.

  • Mac 16 GBFits

    Runs on Apple Silicon with 16 GB. An M1 Pro 16 GB took about 5.5 seconds per step at 1024 pixels.

  • Mac 32 GB+Fits

    Room for LoRAs and ControlNets. Max chips are much faster: an M2 Max reached 1.5 to 3 steps per second.

Set it up.

  1. Download a checkpoint

    One file from the list above, into ComfyUI/models/checkpoints. Then press R in ComfyUI or restart it so the loader lists the file.

  2. Open the SDXL template

    In ComfyUI’s template browser, pick SDXL Simple. It runs the base model and then the refiner. For a fine-tune, delete the refiner’s loader, its two prompt nodes and the second sampler, then set the first KSampler (Advanced) to end_at_step 10000 and return_with_leftover_noise disable. Or build it by hand: Load Checkpoint, two CLIP Text Encode nodes, Empty Latent Image at 1024 × 1024, KSampler, VAE Decode, Save Image.

  3. For Pony: add clip skip

    Put a CLIP Set Last Layer node between Load Checkpoint and both prompt nodes, set to -2. This is ComfyUI’s clip skip 2. Pony V6 needs it; without it you get blobs.

  4. For v-pred files: add the sampling node

    Put ModelSamplingDiscrete between the checkpoint and the sampler, with sampling v_prediction and zsnr on. NoobAI’s makers add RescaleCFG at 0.2 after it. Newer v-pred files carry a marker that ComfyUI reads on its own; many merges don’t.

  5. Add LoRAs with a node

    Use Load LoRA between the checkpoint and the rest. ComfyUI ignores <lora:name:1> text in the prompt, which is AUTOMATIC1111 syntax.

  6. Run

    The first run loads the file and takes longer. After that an 8 GB card makes an image in well under half a minute.

Settings that work.

Same model shape, very different settings. Use the block for the file you have. Where a fine-tune’s author gives settings, these are theirs.

SDXL 1.0 and most fine-tunes

Steps
25
CFG
7 to 8
Sampler
euler or dpmpp_2m
Scheduler
normal or karras
Size
1024 × 1024
Negative
yes
Clip skip
none
VAE
from checkpoint

Comfy’s SDXL Simple template uses 25 steps, CFG 8, euler and normal. dpmpp_2m with karras at CFG 7 is the other common default. Both are fine.

RealVisXL V5.0

Steps
30+
Sampler
dpmpp_sde
Scheduler
karras
Hires fix
0.1 to 0.3 denoise

The RealVisXL V5.0 card recommends DPM++ SDE Karras at 30 steps or more, or DPM++ 2M Karras at 50 or more.

RealVisXL V5.0 Lightning

Steps
5
CFG
1 to 2
Sampler
dpmpp_sde
Scheduler
karras

From the V5.0 Lightning card, which also suggests a hires pass of 3 steps at 0.5 denoise. At CFG 1 the negative prompt has no effect.

Pony Diffusion V6 XL

Steps
25
CFG
7
Sampler
euler_ancestral
Scheduler
normal
Clip skip
-2
Size
1024 × 1024
Prompt starts
score_9, score_8_up, …
Negative
yes

Pony is trained on tags. Start the prompt with the score chain (score_9, score_8_up, score_7_up), then source_anime or another source_ tag, then your tags. Clip skip 2 is required, the author says, “otherwise you will be getting low quality blobs”.

Illustrious and NoobAI (eps)

Steps
20 to 28
CFG
5 to 7.5
Sampler
euler_ancestral
Clip skip
1 or 2

Sampler, steps and CFG are from the Illustrious v0.1 card, which says nothing about clip skip. Merges disagree: some say Illustrious was trained without it, many use 2. Follow the page of your merge. Illustrious and NoobAI don’t use Pony’s score chain.

NoobAI-XL V-Pred

Steps
28 to 35
CFG
4 to 5
Sampler
euler
Scheduler
normal
Sampling
v_prediction, zsnr
RescaleCFG
0.2
Size
832 × 1216
Negative
yes

The V-Pred card says Euler only: “Other samplers will not work properly”. Karras schedulers don’t suit v-prediction either.

Speed files

FileStepsCFGSamplerScheduler
Lightning2, 4 or 8, as in the name1eulersgm_uniform
Hyper-SDas in the name1ddimsgm_uniform
DMD241lcmsgm_uniform
Turbo11euler_ancestralSDTurboScheduler

Lightning’s steps must match the file: the 4-step file wants 4. Hyper’s 12-step CFG LoRA is the exception and wants CFG 5 to 8. DMD2’s card gives diffusers code only (LCM, 4 steps, no guidance); lcm with sgm_uniform is the ComfyUI match. Turbo makes 512 × 512 images.

At CFG 1 these ignore the negative prompt. A higher CFG gives burnt or abstract images: the advice in ByteDance’s Lightning discussions is CFG 1 to 2 for any Lightning, DMD2 or Hyper file. Don’t stack a Lightning LoRA on a checkpoint that’s already Lightning.

Sizes

SDXL was trained at about one megapixel in these shapes, listed in Comfy’s template. Other sizes near one megapixel work too. For bigger images, generate at one of these and upscale.

ShapeWideTall
Square1024 × 10241024 × 1024
Near 4:31152 × 896896 × 1152
Near 3:21216 × 832832 × 1216
Near 16:91344 × 768768 × 1344
Near 21:91536 × 640640 × 1536

The refiner

SDXL launched with a second model, sd_xl_refiner_1.0 (6.1 GB), to polish the last steps. Today’s fine-tunes make finished images on their own and most people skip it. The refiner can’t use SDXL LoRAs or the base model’s ControlNets. For more detail, a hires pass or an upscale does more.

LoRAs across SDXL, Pony and Illustrious.

Pony, Illustrious and NoobAI are all SDXL underneath, so any SDXL LoRA loads on any of them without an error. Loading isn’t the same as working well.

  • Same branch works best. Pony LoRAs on Pony merges, Illustrious LoRAs on Illustrious merges. NoobAI descends from Illustrious, so Illustrious LoRAs mostly work on NoobAI eps.
  • Across branches is hit and miss. A Pony LoRA on Illustrious often loses its look, and Pony’s score_ and source_ tags mean nothing to Illustrious.
  • SD 1.5 LoRAs don’t work. The log fills with lora key not loaded and the image doesn’t change.
  • Pony V7 is a different model. V6 LoRAs don’t work on Pony V7, which is built on AuraFlow.

How fast.

GPUSpeedTime
RTX 40906.2 to 7.6 it/s3.1 to 3.6 s[1]
RTX 30903.6 it/s6.2 s[1]
RTX 4070 12 GB3.2 it/s7.1 s[1]
RTX 30702.3 it/s10.9 s[1]
RTX 3060 Ti2.1 it/s13.3 s[1]
RTX 3060 12 GB1.5 it/s13.8 to 14.9 s[1]
RTX 4060 8 GB1.7 to 1.8 it/sabout 16 s[1]
RTX 4060 Laptop1.8 it/s18.5 s[1]
GTX 10703.2 s/it72 s[1]

SDXL 1.0 in ComfyUI’s default workflow, 1024 × 1024, 20 steps, second run, from ComfyUI’s benchmark thread. A 4-step Lightning or DMD2 file samples in about a fifth of the steps. On a Mac, 2023 reports put an M1 Pro 16 GB at about 5.5 seconds per step.

When it goes wrong.

Black images after starting ComfyUI with --fast
The fp8 mode breaks most SDXL fine-tunes. Remove --fast and --fp8_e4m3fn-unet.
Black images now and then, often on the second run
A 2026 bug in ComfyUI’s dynamic VRAM with reused models. Start ComfyUI with --disable-dynamic-vram until it’s fixed. The console often shows RuntimeWarning: invalid value encountered in cast.
Black or noise images on a Mac
A PyTorch 2.10+ problem on M3 Ultra, M5 and others. Start ComfyUI with --disable-smart-memory, or use torch 2.9.0.
Black images with --fp16-vae, or in AUTOMATIC1111
The original SDXL VAE overflows in fp16. Load the fp16-fix VAE, or drop the flag.
Black images on Illustrious with clip skip
One report shows Illustrious going black with CLIP Set Last Layer at -2, and even at -1. Bypass the node.
Grey, washed-out or noisy images from an anime model
It’s a v-prediction file. Add ModelSamplingDiscrete with v_prediction and zsnr on, and use euler with the normal scheduler.
Abstract or burnt images from a Lightning, DMD2 or Hyper file
CFG is too high. Use 1 to 2, and steps that match the file.
mat1 and mat2 shapes cannot be multiplied (2x2560 and 2816x1280)
A ControlNet that doesn’t match: a base-model ControlNet on the refiner, or an SD 1.5 ControlNet on SDXL.
lora key not loaded
The LoRA is for a different model, usually SD 1.5 on SDXL.
Value not in list: ckpt_name
The file isn’t in models/checkpoints, or ComfyUI hasn’t seen it yet. Move it there and press R.
Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.
Not an error. ComfyUI decodes in tiles and the image still comes out.
ERROR: Could not detect model type
The file is a diffusion model only, not a full checkpoint. Load it with Load Diffusion Model instead of Load Checkpoint.

The 8-step Lightning model only gives me abstract images. What am I doing wrong?

Hugging Face, ByteDance SDXL-Lightning

My Pony LoRA has no effect at all in ComfyUI, though it’s in the prompt.

GitHub, ComfyUI issue #5168

Questions.

How much VRAM does SDXL need?

8 GB runs any SDXL checkpoint at 1024 × 1024. 6 GB works with part of the model in system RAM, which is slower. On a Mac, 16 GB of unified memory is enough.

Do I need the SDXL refiner?

No. Current fine-tunes like RealVisXL, Pony and Illustrious make finished images without it, and the refiner can’t use SDXL LoRAs. A hires pass or an upscale adds more detail.

Why are my SDXL images black?

In ComfyUI the usual causes are the --fast or fp8 flags, the 2026 dynamic VRAM bug (start with --disable-dynamic-vram) and PyTorch 2.10+ on some Macs (start with --disable-smart-memory). The classic fp16 VAE overflow only happens with --fp16-vae; the fp16-fix VAE solves it.

What clip skip do Pony and Illustrious need?

Pony V6 needs clip skip 2, which is CLIP Set Last Layer at -2 in ComfyUI. Illustrious is less clear: its card doesn’t say, and merges recommend 1 or 2. Follow the page of the merge you use.

Do Pony LoRAs work on Illustrious?

They load, since both are SDXL, but they often lose their look, and Pony’s score tags do nothing on Illustrious. Use LoRAs made for the same branch. Illustrious LoRAs mostly work on NoobAI.

What resolution should I use for SDXL?

About one megapixel in the trained shapes: 1024 × 1024, 1152 × 896, 1216 × 832, 1344 × 768 or 1536 × 640, and the same turned on their side.

How many steps for SDXL Lightning?

Exactly the number in the file name: 4 for the 4-step file, 8 for the 8-step one, with CFG 1, euler and sgm_uniform. RealVisXL V5.0 Lightning is its own tune and wants 5 steps with DPM++ SDE Karras at CFG 1 to 2.

Sources: SDXL 1.0 model card, Stability’s SDXL 1.0 announcement, RealVisXL V5.0, RealVisXL V5.0 Lightning, SDXL-Lightning, Hyper-SD, DMD2, Pony Diffusion V6 XL, Illustrious XL v0.1, NoobAI-XL V-Pred, ComfyUI SDXL examples, ComfyUI GPU benchmark thread [1], issue #4572, #15452, #10681, #7521, #1316, #948.

HEISS UI

Or just pick the file.

HEISS UI runs SDXL and everything built on it on the ComfyUI you already have: Pony, Illustrious, NoobAI, Lightning and the rest.

  • Each fine-tune set up right. Pony, Illustrious, v-prediction and speed files get the setting they need, read from the file name.
  • A first model that fits. On an empty studio it’s one tap away, with the version for your GPU or Mac marked.
  • Your LoRAs, sorted. The ones made for this model come first, with their trigger words.
  • A phone studio. Prompt, browse and share from the couch while the computer renders.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.