Base, fine-tunes, speed files.
SDXL 1.0 came out on 26 July 2023. It’s a 3.5B model with two text encoders, CLIP-L and OpenCLIP bigG, and it makes images at about one megapixel. It has a huge pool of fine-tunes, LoRAs and ControlNets, and it makes 1024-pixel images on modest hardware.
Every file below has the same architecture and loads the same way. What changes is the settings.
- SDXL 1.0 base. Stability’s original. Fine for testing, plain next to what came after.
- Photo fine-tunes such as RealVisXL V5.0. Same size, much better photos. Its author allows commercial use.
- Anime fine-tunes. Pony Diffusion V6 XL, Illustrious XL (September 2024) and NoobAI-XL, which builds on Illustrious. Thousands of merges on Civitai descend from these three. Each wants its own prompt style and settings.
- Speed files. SDXL Turbo, SDXL-Lightning, Hyper-SD, DMD2 and LCM. They come as full checkpoints or as LoRAs, and cut an image from 25 steps to between 1 and 8. Many fine-tunes ship a Lightning or DMD2 version, like RealVisXL V5.0 Lightning.
- V-prediction files such as NoobAI-XL V-Pred. They need one extra node, covered below.
Files you need.
One file. An SDXL checkpoint holds the model, both text encoders and the VAE, and goes in models/checkpoints. Pick one of these to start:
-
Photo6.9 GB Download
RealVisXL_V5.0_fp16.safetensorsComfyUI/models/checkpoints/ -
Photo, fast6.9 GB Download
RealVisXL_V5.0_Lightning_fp16.safetensorsComfyUI/models/checkpoints/ five steps instead of thirty -
Base6.9 GB Download
sd_xl_base_1.0.safetensorsComfyUI/models/checkpoints/ -
Anime6.9 GB Download
Illustrious-XL-v0.1.safetensorsComfyUI/models/checkpoints/ Pony V6 is on Civitai only; most people use a merge of either -
Anime, v-pred7.1 GB Download
NoobAI-XL-Vpred-v1.0.safetensorsComfyUI/models/checkpoints/
models/
├── checkpoints/
│ └── RealVisXL_V5.0_fp16.safetensors
├── loras/
│ └── sdxl_lightning_4step_lora.safetensors (optional)
└── vae/
└── sdxl_vae.safetensors (optional)
Speed LoRAs
Instead of a separate Lightning checkpoint, you can add a speed LoRA to any SDXL model. They go in models/loras.
-
LoRA0.4 GB Download
sdxl_lightning_4step_lora.safetensorsComfyUI/models/loras/ 2- and 8-step versions in the same repo -
LoRA0.4 GB Download
dmd2_sdxl_4step_lora_fp16.safetensorsComfyUI/models/loras/ non-commercial licence
The VAE and black images
The VAE is baked into every checkpoint, so you don’t need a separate one. The well-known SDXL black image comes from the original VAE overflowing in fp16. ComfyUI avoids that on its own: it runs the SDXL VAE in bf16 where the GPU supports it and in fp32 otherwise, never in fp16 unless you start it with --fp16-vae. If you use that flag, or come from AUTOMATIC1111 with the same problem, load madebyollin’s fp16-fix VAE (0.3 GB, MIT) with a Load VAE node.
-
VAE0.3 GB Download
sdxl_vae.safetensors (fp16 fix)ComfyUI/models/vae/ only needed with --fp16-vae
fp8 and GGUF
Not needed. The 6.9 GB fp16 file already fits 8 GB. Starting ComfyUI with --fast or --fp8_e4m3fn-unet turns most SDXL fine-tunes into black images, so leave those off for SDXL. GGUF versions exist but are rare, and there is no official one.
What fits your computer.
An SDXL checkpoint at 1024 × 1024, 25 steps. Stability’s own figure at launch was 8 GB of video memory. ComfyUI moves parts of the model to system RAM when the card is smaller, and falls back to a tiled VAE decode when memory runs out at the end.
- 6 GBTight
Works with offloading, as on an RTX 2060 6 GB. A Lightning or DMD2 file keeps the wait short.
- 8 GBFits
Any SDXL checkpoint, plus a LoRA or two. An RTX 4060 8 GB makes a 20-step image in about 16 seconds.
- 12 GBFits
Room for LoRAs, a ControlNet and a hires pass on top.
- 16 GBFits
Everything above, with room for batches.
- 24 GBFits
Base and refiner together, larger batches and upscales.
- Mac 16 GBFits
Runs on Apple Silicon with 16 GB. An M1 Pro 16 GB took about 5.5 seconds per step at 1024 pixels.
- Mac 32 GB+Fits
Room for LoRAs and ControlNets. Max chips are much faster: an M2 Max reached 1.5 to 3 steps per second.
- 8 GB graphics cards
- 12 GB cards
- RTX 3060 Ti
- RTX 3060 12 GB
- Laptop GPUs
- Mac with Apple Silicon
- AMD Radeon
Set it up.
-
Download a checkpoint
One file from the list above, into
ComfyUI/models/checkpoints. Then press R in ComfyUI or restart it so the loader lists the file. -
Open the SDXL template
In ComfyUI’s template browser, pick SDXL Simple. It runs the base model and then the refiner. For a fine-tune, delete the refiner’s loader, its two prompt nodes and the second sampler, then set the first KSampler (Advanced) to
end_at_step10000 andreturn_with_leftover_noisedisable. Or build it by hand: Load Checkpoint, two CLIP Text Encode nodes, Empty Latent Image at 1024 × 1024, KSampler, VAE Decode, Save Image. -
For Pony: add clip skip
Put a CLIP Set Last Layer node between Load Checkpoint and both prompt nodes, set to
-2. This is ComfyUI’s clip skip 2. Pony V6 needs it; without it you get blobs. -
For v-pred files: add the sampling node
Put ModelSamplingDiscrete between the checkpoint and the sampler, with sampling
v_predictionandzsnron. NoobAI’s makers add RescaleCFG at 0.2 after it. Newer v-pred files carry a marker that ComfyUI reads on its own; many merges don’t. -
Add LoRAs with a node
Use Load LoRA between the checkpoint and the rest. ComfyUI ignores
<lora:name:1>text in the prompt, which is AUTOMATIC1111 syntax. -
Run
The first run loads the file and takes longer. After that an 8 GB card makes an image in well under half a minute.
Settings that work.
Same model shape, very different settings. Use the block for the file you have. Where a fine-tune’s author gives settings, these are theirs.
SDXL 1.0 and most fine-tunes
- Steps
- 25
- CFG
- 7 to 8
- Sampler
- euler or dpmpp_2m
- Scheduler
- normal or karras
- Size
- 1024 × 1024
- Negative
- yes
- Clip skip
- none
- VAE
- from checkpoint
Comfy’s SDXL Simple template uses 25 steps, CFG 8, euler and normal. dpmpp_2m with karras at CFG 7 is the other common default. Both are fine.
RealVisXL V5.0
- Steps
- 30+
- Sampler
- dpmpp_sde
- Scheduler
- karras
- Hires fix
- 0.1 to 0.3 denoise
The RealVisXL V5.0 card recommends DPM++ SDE Karras at 30 steps or more, or DPM++ 2M Karras at 50 or more.
RealVisXL V5.0 Lightning
- Steps
- 5
- CFG
- 1 to 2
- Sampler
- dpmpp_sde
- Scheduler
- karras
From the V5.0 Lightning card, which also suggests a hires pass of 3 steps at 0.5 denoise. At CFG 1 the negative prompt has no effect.
Pony Diffusion V6 XL
- Steps
- 25
- CFG
- 7
- Sampler
- euler_ancestral
- Scheduler
- normal
- Clip skip
- -2
- Size
- 1024 × 1024
- Prompt starts
- score_9, score_8_up, …
- Negative
- yes
Pony is trained on tags. Start the prompt with the score chain (score_9, score_8_up, score_7_up), then source_anime or another source_ tag, then your tags. Clip skip 2 is required, the author says, “otherwise you will be getting low quality blobs”.
Illustrious and NoobAI (eps)
- Steps
- 20 to 28
- CFG
- 5 to 7.5
- Sampler
- euler_ancestral
- Clip skip
- 1 or 2
Sampler, steps and CFG are from the Illustrious v0.1 card, which says nothing about clip skip. Merges disagree: some say Illustrious was trained without it, many use 2. Follow the page of your merge. Illustrious and NoobAI don’t use Pony’s score chain.
NoobAI-XL V-Pred
- Steps
- 28 to 35
- CFG
- 4 to 5
- Sampler
- euler
- Scheduler
- normal
- Sampling
- v_prediction, zsnr
- RescaleCFG
- 0.2
- Size
- 832 × 1216
- Negative
- yes
The V-Pred card says Euler only: “Other samplers will not work properly”. Karras schedulers don’t suit v-prediction either.
Speed files
| File | Steps | CFG | Sampler | Scheduler |
|---|---|---|---|---|
| Lightning | 2, 4 or 8, as in the name | 1 | euler | sgm_uniform |
| Hyper-SD | as in the name | 1 | ddim | sgm_uniform |
| DMD2 | 4 | 1 | lcm | sgm_uniform |
| Turbo | 1 | 1 | euler_ancestral | SDTurboScheduler |
Lightning’s steps must match the file: the 4-step file wants 4. Hyper’s 12-step CFG LoRA is the exception and wants CFG 5 to 8. DMD2’s card gives diffusers code only (LCM, 4 steps, no guidance); lcm with sgm_uniform is the ComfyUI match. Turbo makes 512 × 512 images.
At CFG 1 these ignore the negative prompt. A higher CFG gives burnt or abstract images: the advice in ByteDance’s Lightning discussions is CFG 1 to 2 for any Lightning, DMD2 or Hyper file. Don’t stack a Lightning LoRA on a checkpoint that’s already Lightning.
Sizes
SDXL was trained at about one megapixel in these shapes, listed in Comfy’s template. Other sizes near one megapixel work too. For bigger images, generate at one of these and upscale.
| Shape | Wide | Tall |
|---|---|---|
| Square | 1024 × 1024 | 1024 × 1024 |
| Near 4:3 | 1152 × 896 | 896 × 1152 |
| Near 3:2 | 1216 × 832 | 832 × 1216 |
| Near 16:9 | 1344 × 768 | 768 × 1344 |
| Near 21:9 | 1536 × 640 | 640 × 1536 |
The refiner
SDXL launched with a second model, sd_xl_refiner_1.0 (6.1 GB), to polish the last steps. Today’s fine-tunes make finished images on their own and most people skip it. The refiner can’t use SDXL LoRAs or the base model’s ControlNets. For more detail, a hires pass or an upscale does more.
LoRAs across SDXL, Pony and Illustrious.
Pony, Illustrious and NoobAI are all SDXL underneath, so any SDXL LoRA loads on any of them without an error. Loading isn’t the same as working well.
- Same branch works best. Pony LoRAs on Pony merges, Illustrious LoRAs on Illustrious merges. NoobAI descends from Illustrious, so Illustrious LoRAs mostly work on NoobAI eps.
- Across branches is hit and miss. A Pony LoRA on Illustrious often loses its look, and Pony’s
score_andsource_tags mean nothing to Illustrious. - SD 1.5 LoRAs don’t work. The log fills with
lora key not loadedand the image doesn’t change. - Pony V7 is a different model. V6 LoRAs don’t work on Pony V7, which is built on AuraFlow.
How fast.
| GPU | Speed | Time |
|---|---|---|
| RTX 4090 | 6.2 to 7.6 it/s | 3.1 to 3.6 s[1] |
| RTX 3090 | 3.6 it/s | 6.2 s[1] |
| RTX 4070 12 GB | 3.2 it/s | 7.1 s[1] |
| RTX 3070 | 2.3 it/s | 10.9 s[1] |
| RTX 3060 Ti | 2.1 it/s | 13.3 s[1] |
| RTX 3060 12 GB | 1.5 it/s | 13.8 to 14.9 s[1] |
| RTX 4060 8 GB | 1.7 to 1.8 it/s | about 16 s[1] |
| RTX 4060 Laptop | 1.8 it/s | 18.5 s[1] |
| GTX 1070 | 3.2 s/it | 72 s[1] |
SDXL 1.0 in ComfyUI’s default workflow, 1024 × 1024, 20 steps, second run, from ComfyUI’s benchmark thread. A 4-step Lightning or DMD2 file samples in about a fifth of the steps. On a Mac, 2023 reports put an M1 Pro 16 GB at about 5.5 seconds per step.
When it goes wrong.
- Black images after starting ComfyUI with
--fast - The fp8 mode breaks most SDXL fine-tunes. Remove
--fastand--fp8_e4m3fn-unet. - Black images now and then, often on the second run
- A 2026 bug in ComfyUI’s dynamic VRAM with reused models. Start ComfyUI with
--disable-dynamic-vramuntil it’s fixed. The console often showsRuntimeWarning: invalid value encountered in cast. - Black or noise images on a Mac
- A PyTorch 2.10+ problem on M3 Ultra, M5 and others. Start ComfyUI with
--disable-smart-memory, or use torch 2.9.0. - Black images with
--fp16-vae, or in AUTOMATIC1111 - The original SDXL VAE overflows in fp16. Load the fp16-fix VAE, or drop the flag.
- Black images on Illustrious with clip skip
- One report shows Illustrious going black with CLIP Set Last Layer at -2, and even at -1. Bypass the node.
- Grey, washed-out or noisy images from an anime model
- It’s a v-prediction file. Add ModelSamplingDiscrete with
v_predictionand zsnr on, and use euler with the normal scheduler. - Abstract or burnt images from a Lightning, DMD2 or Hyper file
- CFG is too high. Use 1 to 2, and steps that match the file.
mat1 and mat2 shapes cannot be multiplied (2x2560 and 2816x1280)- A ControlNet that doesn’t match: a base-model ControlNet on the refiner, or an SD 1.5 ControlNet on SDXL.
lora key not loaded- The LoRA is for a different model, usually SD 1.5 on SDXL.
Value not in list: ckpt_name- The file isn’t in
models/checkpoints, or ComfyUI hasn’t seen it yet. Move it there and press R. Ran out of memory when regular VAE decoding, retrying with tiled VAE decoding.- Not an error. ComfyUI decodes in tiles and the image still comes out.
ERROR: Could not detect model type- The file is a diffusion model only, not a full checkpoint. Load it with Load Diffusion Model instead of Load Checkpoint.
The 8-step Lightning model only gives me abstract images. What am I doing wrong?
My Pony LoRA has no effect at all in ComfyUI, though it’s in the prompt.
Questions.
How much VRAM does SDXL need?
8 GB runs any SDXL checkpoint at 1024 × 1024. 6 GB works with part of the model in system RAM, which is slower. On a Mac, 16 GB of unified memory is enough.
Do I need the SDXL refiner?
No. Current fine-tunes like RealVisXL, Pony and Illustrious make finished images without it, and the refiner can’t use SDXL LoRAs. A hires pass or an upscale adds more detail.
Why are my SDXL images black?
In ComfyUI the usual causes are the --fast or fp8 flags, the 2026 dynamic VRAM bug (start with --disable-dynamic-vram) and PyTorch 2.10+ on some Macs (start with --disable-smart-memory). The classic fp16 VAE overflow only happens with --fp16-vae; the fp16-fix VAE solves it.
What clip skip do Pony and Illustrious need?
Pony V6 needs clip skip 2, which is CLIP Set Last Layer at -2 in ComfyUI. Illustrious is less clear: its card doesn’t say, and merges recommend 1 or 2. Follow the page of the merge you use.
Do Pony LoRAs work on Illustrious?
They load, since both are SDXL, but they often lose their look, and Pony’s score tags do nothing on Illustrious. Use LoRAs made for the same branch. Illustrious LoRAs mostly work on NoobAI.
What resolution should I use for SDXL?
About one megapixel in the trained shapes: 1024 × 1024, 1152 × 896, 1216 × 832, 1344 × 768 or 1536 × 640, and the same turned on their side.
How many steps for SDXL Lightning?
Exactly the number in the file name: 4 for the 4-step file, 8 for the 8-step one, with CFG 1, euler and sgm_uniform. RealVisXL V5.0 Lightning is its own tune and wants 5 steps with DPM++ SDE Karras at CFG 1 to 2.
Sources: SDXL 1.0 model card, Stability’s SDXL 1.0 announcement, RealVisXL V5.0, RealVisXL V5.0 Lightning, SDXL-Lightning, Hyper-SD, DMD2, Pony Diffusion V6 XL, Illustrious XL v0.1, NoobAI-XL V-Pred, ComfyUI SDXL examples, ComfyUI GPU benchmark thread [1], issue #4572, #15452, #10681, #7521, #1316, #948.