Large, Turbo or Medium.
Stability released SD 3.5 in October 2024. All three versions use three text encoders: CLIP-L, OpenCLIP bigG (CLIP-G) and T5-XXL.
- Large is the 8B model, made for images up to about one megapixel. The best quality of the three.
- Large Turbo is Large distilled to four steps. Same size, much faster, a little less detail.
- Medium is a 2.5B model with a newer layout (MMDiT-X). It handles 0.25 to 2 megapixels and fits smaller cards.
Stability notes that SD 3.5 gives more variation from seed to seed than earlier models, which suits exploring and makes one exact look harder to repeat. For anatomy, the Medium card recommends Skip Layer Guidance, covered under settings.
Files you need.
The easy way: one file
Comfy-Org packs the model, all three text encoders and the VAE into one checkpoint. It goes in models/checkpoints and loads with Load Checkpoint.
-
Large14.9 GB Download
sd3.5_large_fp8_scaled.safetensorsComfyUI/models/checkpoints/ model and encoders in fp8 -
Medium11.6 GB Download
sd3.5_medium_incl_clips_t5xxlfp8scaled.safetensorsComfyUI/models/checkpoints/ fp16 model, T5 in fp8
Both carry an fp8 T5, which doesn’t run on a Mac. See the Mac notes below.
Separate files
Stability’s own files hold the model without the text encoders, so the encoders come separately. This is also the route for Large Turbo and for GGUF.
-
Model16.5 GB Download
sd3.5_large.safetensorsComfyUI/models/checkpoints/ gated. Turbo and Medium (5.1 GB) in their own repos -
Text encoder0.2 GB Download
clip_l.safetensorsComfyUI/models/text_encoders/ -
Text encoder1.4 GB Download
clip_g.safetensorsComfyUI/models/text_encoders/ -
Text encoder5.2 GB Download
t5xxl_fp8_e4m3fn_scaled.safetensorsComfyUI/models/text_encoders/ or t5xxl_fp16.safetensors, 9.8 GB, which a Mac needs
models/
├── checkpoints/
│ └── sd3.5_large.safetensors
└── text_encoders/
├── clip_l.safetensors
├── clip_g.safetensors
└── t5xxl_fp8_e4m3fn_scaled.safetensors
Smaller files: GGUF
city96’s GGUF builds load through the ComfyUI-GGUF nodes and go in models/unet. There is a Large Turbo set with the same sizes. Large has no K-quants; Medium does.
| Model | Q8_0 | Q5 | Q4 |
|---|---|---|---|
| Large | 8.8 GB | 6.3 GB (Q5_1) | 4.8 GB (Q4_0) |
| Medium | 2.9 GB | 2.1 GB (Q5_K_M) | 1.8 GB (Q4_K_M) |
A GGUF holds the model only. You also need the three text encoders and a VAE.
What fits your computer.
At 1024 × 1024. T5-XXL is bigger than Medium itself, and ComfyUI moves it out of the way before sampling, so system RAM matters too: Comfy’s examples ask for 32 GB of RAM or more for the fp16 T5.
- 6 GBTight
Medium as GGUF, with the fp8 T5 offloaded to system RAM.
- 8 GBFits
Medium, or Large as GGUF Q4_0 (4.8 GB). One Krita user on an RTX 4060 8 GB found SD 3.5 GGUF about twice as fast as Flux GGUF.
- 12 GBFits
Medium at full precision, or Large as GGUF Q8_0 (8.8 GB).
- 16 GBFits
Large fp8 all-in-one. Comfy estimates 14.9 GB.
- 24 GBFits
Large at full precision with the fp16 T5.
- Mac 16 GBTight
Medium as GGUF with the fp16 T5. No fp8 files.
- Mac 32 GB+Tight
Large as GGUF Q8_0 with the fp16 T5, about 20 GB of files together.
Set it up.
-
Update ComfyUI
Medium needs a ComfyUI from after its release; older versions stop with a size mismatch error. ComfyUI Desktop updates itself. In a manual install:
Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download
An all-in-one checkpoint into
models/checkpoints, or a model file plus the three encoders. For Stability’s files, accept the licence on Hugging Face first. -
Open the template
In the template browser, SD3.5 Simple loads
sd3.5_large_fp8_scaledwith one Load Checkpoint node. Swap in the Medium file if you use that. -
For separate files, add the encoders
Keep Load Checkpoint for the model and VAE, and add a TripleCLIPLoader with
clip_l,clip_gand the T5. For GGUF, use Unet Loader (GGUF), the TripleCLIPLoader and Load VAE. -
For Medium, add Skip Layer Guidance
Put SkipLayerGuidanceSD3 between the model and the sampler. Stability recommends it for Medium’s structure and anatomy.
-
Run
Use EmptySD3LatentImage for the size. The first run loads the T5 and takes a while.
Settings that work.
Large
- Steps
- 28
- CFG
- 3.5
- Sampler
- euler
- Scheduler
- sgm_uniform
- Size
- 1024 × 1024
- Shift
- 3 (default)
- Negative
- yes
- Latent
- EmptySD3LatentImage
28 steps at 3.5 is from Stability’s Large card. Comfy’s SD3.5 Simple template uses 20 steps at CFG 4, which is faster and close. ComfyUI sets SD 3.5’s shift to 3 on its own.
Large Turbo
- Steps
- 4
- CFG
- 1.2
- Sampler
- euler
- Scheduler
- sgm_uniform
ComfyUI’s examples say “set steps to 4 and cfg to 1.2”. Stability’s card uses no guidance at all, which is CFG 1 in ComfyUI.
Medium
- Steps
- 40
- CFG
- 4.5
- Guidance
- SkipLayerGuidanceSD3
- Size
- 0.25 to 2 MP
From the Medium card. Keep prompts under 256 T5 tokens: the card warns of artefacts at the image edges beyond that.
Without T5
T5 is optional. Load only CLIP-L and CLIP-G with a DualCLIPLoader and SD 3.5 still works, with less memory and worse prompt following and text in images. At low CFG the difference is smaller.
When it goes wrong.
Trying to convert Float8_e4m3fn to the MPS backend but it does not have support for that dtype.- An fp8 file on a Mac. Use GGUF or full-precision weights and
t5xxl_fp16. size mismatch for joint_blocks.0.x_block.adaLN_modulation.1.weight- ComfyUI is too old for Medium. Update it.
clip missing: ['text_projection.weight']- Harmless. It shows after some updates and the image comes out fine.
Value not in list: clip_name2: 'clip_l.safetensors'- An encoder file is missing from
models/text_encoders. Download it and press R. - Black images with a GGUF
- Usually a wrong or half-downloaded VAE. On a Mac, Turbo GGUF black images at sizes other than 1024 were a PyTorch bug fixed in 2.6.
- Mosaic-like images from Medium on a Mac
- In the one report, the answer was a CFG set too high. Go back to 4.5 or lower.
unknown model architecture: 'sd3'- The GGUF was opened in LM Studio or llama.cpp. These files are for ComfyUI-GGUF only.
- Out of memory before sampling starts
- The T5 encoder. Use the fp8 T5 (not on a Mac), or run without T5.
Which VAE am I meant to use with the Large GGUF? The repo only has the model.
Medium runs out of memory on a 16 GB card, though the model is small.
Questions.
How much VRAM does SD 3.5 Large need?
About 15 GB for Comfy-Org’s fp8 all-in-one file, which fits a 16 GB card. As GGUF, Large runs on 8 GB (Q4_0, 4.8 GB) or 12 GB (Q8_0, 8.8 GB).
Can I run SD 3.5 without T5?
Yes. Load only CLIP-L and CLIP-G with a DualCLIPLoader. It uses less memory, but prompt following and text in images get worse.
Can I use SD 3.5 commercially?
Yes, under the Stability AI Community License, as long as your yearly revenue is below 1 million US dollars. Above that you need an Enterprise licence. You own what you generate.
Which VAE do I need for the SD 3.5 GGUF?
The SD 3.5 VAE from the vae folder of Stability’s gated repo, or one saved out of an all-in-one checkpoint with a VAE Save node.
Why does SD 3.5 fail on my Mac?
SD 3.5’s fp8 files stop with a Float8 error on Apple Silicon, and both Comfy-Org all-in-one files and the usual T5 download are fp8. Use GGUF or full-precision model files with t5xxl_fp16.
Large or Medium?
Large for quality on 12 GB or more, with GGUF. Medium on smaller cards and for sizes up to 2 megapixels. Give Medium 40 steps and Skip Layer Guidance.
Sources: SD 3.5 Large card, Large Turbo card, Medium card, Stability’s announcement, Community License, ComfyUI SD3 examples, Comfy blog on Medium, Comfy-Org fp8 files, city96 GGUF, ComfyUI issue #12202, Krita AI Diffusion #1328.