Image model

How to run SD 3.5 locally.

SD 3.5 runs in ComfyUI most easily from Comfy-Org’s all-in-one fp8 checkpoints: 14.9 GB for Large, 11.6 GB for Medium, text encoders and VAE included. Large wants about 15 GB of video memory, or a GGUF on 8 to 12 GB cards. It’s free for commercial use below 1 million dollars of yearly revenue.

Updated 29 Sep 202610 min read

Maker
Stability AI
Released
Oct 2024Large and Turbo 22 Oct, Medium 29 Oct
Licence
Stability CommunityFree below $1M yearly revenue
Memory
15 GBLarge fp8. GGUF from 8 GB

Large, Turbo or Medium.

Stability released SD 3.5 in October 2024. All three versions use three text encoders: CLIP-L, OpenCLIP bigG (CLIP-G) and T5-XXL.

  • Large is the 8B model, made for images up to about one megapixel. The best quality of the three.
  • Large Turbo is Large distilled to four steps. Same size, much faster, a little less detail.
  • Medium is a 2.5B model with a newer layout (MMDiT-X). It handles 0.25 to 2 megapixels and fits smaller cards.

Stability notes that SD 3.5 gives more variation from seed to seed than earlier models, which suits exploring and makes one exact look harder to repeat. For anatomy, the Medium card recommends Skip Layer Guidance, covered under settings.

Files you need.

The easy way: one file

Comfy-Org packs the model, all three text encoders and the VAE into one checkpoint. It goes in models/checkpoints and loads with Load Checkpoint.

  • Large sd3.5_large_fp8_scaled.safetensors ComfyUI/models/checkpoints/ model and encoders in fp8
    14.9 GB Download
  • Medium sd3.5_medium_incl_clips_t5xxlfp8scaled.safetensors ComfyUI/models/checkpoints/ fp16 model, T5 in fp8
    11.6 GB Download

Both carry an fp8 T5, which doesn’t run on a Mac. See the Mac notes below.

Separate files

Stability’s own files hold the model without the text encoders, so the encoders come separately. This is also the route for Large Turbo and for GGUF.

  • Model sd3.5_large.safetensors ComfyUI/models/checkpoints/ gated. Turbo and Medium (5.1 GB) in their own repos
    16.5 GB Download
  • Text encoder clip_l.safetensors ComfyUI/models/text_encoders/
    0.2 GB Download
  • Text encoder clip_g.safetensors ComfyUI/models/text_encoders/
    1.4 GB Download
  • Text encoder t5xxl_fp8_e4m3fn_scaled.safetensors ComfyUI/models/text_encoders/ or t5xxl_fp16.safetensors, 9.8 GB, which a Mac needs
    5.2 GB Download
ComfyUI/models
models/
├── checkpoints/
│   └── sd3.5_large.safetensors
└── text_encoders/
    ├── clip_l.safetensors
    ├── clip_g.safetensors
    └── t5xxl_fp8_e4m3fn_scaled.safetensors

Smaller files: GGUF

city96’s GGUF builds load through the ComfyUI-GGUF nodes and go in models/unet. There is a Large Turbo set with the same sizes. Large has no K-quants; Medium does.

ModelQ8_0Q5Q4
Large8.8 GB6.3 GB (Q5_1)4.8 GB (Q4_0)
Medium2.9 GB2.1 GB (Q5_K_M)1.8 GB (Q4_K_M)

A GGUF holds the model only. You also need the three text encoders and a VAE.

What fits your computer.

At 1024 × 1024. T5-XXL is bigger than Medium itself, and ComfyUI moves it out of the way before sampling, so system RAM matters too: Comfy’s examples ask for 32 GB of RAM or more for the fp16 T5.

  • 6 GBTight

    Medium as GGUF, with the fp8 T5 offloaded to system RAM.

  • 8 GBFits

    Medium, or Large as GGUF Q4_0 (4.8 GB). One Krita user on an RTX 4060 8 GB found SD 3.5 GGUF about twice as fast as Flux GGUF.

  • 12 GBFits

    Medium at full precision, or Large as GGUF Q8_0 (8.8 GB).

  • 16 GBFits

    Large fp8 all-in-one. Comfy estimates 14.9 GB.

  • 24 GBFits

    Large at full precision with the fp16 T5.

  • Mac 16 GBTight

    Medium as GGUF with the fp16 T5. No fp8 files.

  • Mac 32 GB+Tight

    Large as GGUF Q8_0 with the fp16 T5, about 20 GB of files together.

Set it up.

  1. Update ComfyUI

    Medium needs a ComfyUI from after its release; older versions stop with a size mismatch error. ComfyUI Desktop updates itself. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download

    An all-in-one checkpoint into models/checkpoints, or a model file plus the three encoders. For Stability’s files, accept the licence on Hugging Face first.

  3. Open the template

    In the template browser, SD3.5 Simple loads sd3.5_large_fp8_scaled with one Load Checkpoint node. Swap in the Medium file if you use that.

  4. For separate files, add the encoders

    Keep Load Checkpoint for the model and VAE, and add a TripleCLIPLoader with clip_l, clip_g and the T5. For GGUF, use Unet Loader (GGUF), the TripleCLIPLoader and Load VAE.

  5. For Medium, add Skip Layer Guidance

    Put SkipLayerGuidanceSD3 between the model and the sampler. Stability recommends it for Medium’s structure and anatomy.

  6. Run

    Use EmptySD3LatentImage for the size. The first run loads the T5 and takes a while.

Settings that work.

Large

Steps
28
CFG
3.5
Sampler
euler
Scheduler
sgm_uniform
Size
1024 × 1024
Shift
3 (default)
Negative
yes
Latent
EmptySD3LatentImage

28 steps at 3.5 is from Stability’s Large card. Comfy’s SD3.5 Simple template uses 20 steps at CFG 4, which is faster and close. ComfyUI sets SD 3.5’s shift to 3 on its own.

Large Turbo

Steps
4
CFG
1.2
Sampler
euler
Scheduler
sgm_uniform

ComfyUI’s examples say “set steps to 4 and cfg to 1.2”. Stability’s card uses no guidance at all, which is CFG 1 in ComfyUI.

Medium

Steps
40
CFG
4.5
Guidance
SkipLayerGuidanceSD3
Size
0.25 to 2 MP

From the Medium card. Keep prompts under 256 T5 tokens: the card warns of artefacts at the image edges beyond that.

Without T5

T5 is optional. Load only CLIP-L and CLIP-G with a DualCLIPLoader and SD 3.5 still works, with less memory and worse prompt following and text in images. At low CFG the difference is smaller.

When it goes wrong.

Trying to convert Float8_e4m3fn to the MPS backend but it does not have support for that dtype.
An fp8 file on a Mac. Use GGUF or full-precision weights and t5xxl_fp16.
size mismatch for joint_blocks.0.x_block.adaLN_modulation.1.weight
ComfyUI is too old for Medium. Update it.
clip missing: ['text_projection.weight']
Harmless. It shows after some updates and the image comes out fine.
Value not in list: clip_name2: 'clip_l.safetensors'
An encoder file is missing from models/text_encoders. Download it and press R.
Black images with a GGUF
Usually a wrong or half-downloaded VAE. On a Mac, Turbo GGUF black images at sizes other than 1024 were a PyTorch bug fixed in 2.6.
Mosaic-like images from Medium on a Mac
In the one report, the answer was a CFG set too high. Go back to 4.5 or lower.
unknown model architecture: 'sd3'
The GGUF was opened in LM Studio or llama.cpp. These files are for ComfyUI-GGUF only.
Out of memory before sampling starts
The T5 encoder. Use the fp8 T5 (not on a Mac), or run without T5.

Which VAE am I meant to use with the Large GGUF? The repo only has the model.

Hugging Face, city96 SD 3.5 Large GGUF

Medium runs out of memory on a 16 GB card, though the model is small.

Hugging Face, SD 3.5 Medium

Questions.

How much VRAM does SD 3.5 Large need?

About 15 GB for Comfy-Org’s fp8 all-in-one file, which fits a 16 GB card. As GGUF, Large runs on 8 GB (Q4_0, 4.8 GB) or 12 GB (Q8_0, 8.8 GB).

Can I run SD 3.5 without T5?

Yes. Load only CLIP-L and CLIP-G with a DualCLIPLoader. It uses less memory, but prompt following and text in images get worse.

Can I use SD 3.5 commercially?

Yes, under the Stability AI Community License, as long as your yearly revenue is below 1 million US dollars. Above that you need an Enterprise licence. You own what you generate.

Which VAE do I need for the SD 3.5 GGUF?

The SD 3.5 VAE from the vae folder of Stability’s gated repo, or one saved out of an all-in-one checkpoint with a VAE Save node.

Why does SD 3.5 fail on my Mac?

SD 3.5’s fp8 files stop with a Float8 error on Apple Silicon, and both Comfy-Org all-in-one files and the usual T5 download are fp8. Use GGUF or full-precision model files with t5xxl_fp16.

Large or Medium?

Large for quality on 12 GB or more, with GGUF. Medium on smaller cards and for sizes up to 2 megapixels. Give Medium 40 steps and Skip Layer Guidance.

Sources: SD 3.5 Large card, Large Turbo card, Medium card, Stability’s announcement, Community License, ComfyUI SD3 examples, Comfy blog on Medium, Comfy-Org fp8 files, city96 GGUF, ComfyUI issue #12202, Krita AI Diffusion #1328.

HEISS UI

Three encoders, one list.

HEISS UI runs SD 3.5 Large, Turbo and Medium on the ComfyUI you already have. Pick the file and the setup shows which parts are missing.

  • All three text encoders, listed. Each one with its size and a button, fetched in one go.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • Your LoRAs, sorted. The ones made for this model come first, with their trigger words.
  • Failures that explain themselves. When a run runs out of memory, it says so and offers the fix.

The VAE is gated at Stability, so it stays a manual download. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.