Image model

How to run Sana locally.

Sana isn’t built into ComfyUI: it runs through one of two custom node packs, and each loads its own Gemma 2 2B encoder and DC-AE VAE. The 1.6B model at 1024 px needs about 9 GB of VRAM, the 4K one 18 GB. Sana Sprint makes an image in two steps, and one of the packs runs it on a Mac.

Updated 29 Sep 20269 min read

Maker
NVIDIAwith MIT Han Lab
Released
Nov 2024SANA 1.5 and Sprint in Mar 2025
Licence
Apache 2.0since 31 July 2026. Gemma: own terms
Memory
About 9 GB1.6B at 1024 px. 4K needs 18 GB

Which model.

Sana is NVIDIA’s efficient image model: a linear diffusion transformer with a deep-compression autoencoder that shrinks images 32 times per side, so large sizes stay cheap. The family has grown in three steps.

  • Sana 1.0 (late 2024): 0.6B and 1.6B at 1024 px, plus 1.6B models trained for 2K and 4K.
  • SANA 1.5 (March 2025): 1.6B and 4.8B at 1024 px.
  • Sana Sprint (March 2025): 0.6B and 1.6B distilled for one to four steps.

Start with SANA 1.5 1.6B for quality or Sprint for speed. In ComfyUI, results are a little behind NVIDIA’s own code, which samples with a Flow-DPM solver ComfyUI doesn’t have, and text in images is weaker than SDXL’s.

Two node packs.

ComfyUI has no Sana loader of its own. Two packs fill the gap, and they read different files.

  • ComfyUI_ExtraModels, NVIDIA’s fork loads the original .pth checkpoints and samples with ComfyUI’s own KSampler, so you get live previews and ComfyUI’s samplers. It runs on NVIDIA GPUs or the CPU, not on a Mac.
  • ComfyUI-SANA wraps the diffusers pipeline and loads diffusers folders. It runs on Apple Silicon, NVIDIA and CPU, with its own generate node instead of a KSampler.

Files you need.

For ExtraModels

Pick a preset whose name starts with Efficient-Large-Model/ in the SanaCheckpointLoader and it downloads the checkpoint itself. The Gemma encoder and the VAE download the same way.

ModelCheckpointResolution
SANA 1.5 1.6B6.4 GB1024 px
SANA 1.5 4.8B18.9 GB1024 px
Sprint 1.6B6.5 GB1024 px
Sprint 0.6B2.4 GB1024 px
Sana 1.6B6.4 GB1024 px
Sana 1.6B 2K / 4K6.5 / 6.6 GB2048 / 4096 px
Sana 0.6B2.4 GB1024 px

Plus the Gemma 2 2B encoder (5.2 GB) and the DC-AE 1.1 VAE (1.2 GB), shared by all of them.

For ComfyUI-SANA

Whole diffusers folders, each with its own encoder and VAE inside, placed in ComfyUI/models/diffusers/<name>. Download only the variant you need: a full snapshot of the 1.6B repo is about 22 GB of duplicates.

  • Sprint Sana_Sprint_0.6B_1024px_diffusers ComfyUI/models/diffusers/ folder: transformer, text encoder, VAE
    7.7 GB Download
  • SANA 1.5 SANA1.5_1.6B_1024px_diffusers ComfyUI/models/diffusers/ folder: transformer, text encoder, VAE
    9.7 GB Download
Terminal, in the ComfyUI folder
hf download Efficient-Large-Model/Sana_Sprint_0.6B_1024px_diffusers --local-dir models/diffusers/Sana_Sprint_0.6B_1024px_diffusers

What fits your computer.

At 1024 × 1024 unless noted. The 18 GB for 4K is NVIDIA’s figure; the rest comes from user reports and file sizes.

  • 8 GBTight

    The 0.6B models. The 1.6B at 1024 px needs about 8.7 GB with VAE offloading, and an 8 GB RTX 4060 laptop ran out of memory.

  • 12 GBFits

    Sana 1.6B, SANA 1.5 1.6B and Sprint at 1024 px.

  • 16 GBFits

    All the 1.6B models with room to spare. The 4K model needs more.

  • 24 GBFits

    The 4K workflow (18 GB) and SANA 1.5 4.8B, which runs on an RTX 3090.

  • MacVia pack

    Only through ComfyUI-SANA with device set to mps. Its README says Sprint 0.6B runs comfortably on Apple Silicon; no timings given.

Set it up.

With ExtraModels (NVIDIA GPUs)

  1. Remove the old pack

    If custom_nodes has city96’s ComfyUI_ExtraModels, delete or move that folder first. Both use the same folder name.

  2. Install NVIDIA’s fork

    Manager can’t find it by name, so clone it and install its requirements with ComfyUI’s own Python. ExtraModels’ VAE loader also needs diffusers.

    Terminal, in ComfyUI/custom_nodes
    git clone https://github.com/lawrence-cj/ComfyUI_ExtraModels.git
    pip install -r ComfyUI_ExtraModels/requirements.txt diffusers
  3. Load NVIDIA’s workflow

    Restart ComfyUI and open Sana_FlowEuler.json (or the 2K, 4K, SANA-1.5 or Sprint one) from NVlabs/Sana’s ComfyUI folder.

  4. Pick presets

    In SanaCheckpointLoader, choose an Efficient-Large-Model/… preset. In GemmaLoader, use Efficient-Large-Model/gemma-2-2b-it: it’s the same model as Google’s, without the gated access.

  5. Swap the latent node

    EmptySanaLatentImage fails on current ComfyUI. Use ComfyUI’s EmptyHunyuanImageLatent instead: same 32 channels, same 1/32 scale. Keep the size a multiple of 32 and matched to the model.

  6. Set CFG and run

    CFG 2 for Sana 1.0, 4.5 for SANA 1.5. The first run downloads the checkpoint, Gemma and the VAE.

With ComfyUI-SANA (Mac, or anywhere)

Terminal, in ComfyUI/custom_nodes
git clone https://github.com/geoffitect/ComfyUI-SANA.git
pip install -r ComfyUI-SANA/requirements.txt

Download a diffusers folder as shown above, restart ComfyUI and wire SANA Model Loader into SANA Generate and a Save Image node. On a Mac, set the loader’s device to mps. The original .pth files don’t load here.

Settings that work.

These follow NVIDIA’s own ComfyUI workflows for ExtraModels.

Sana 1.0 and SANA 1.5

Steps
28
CFG
2SANA 1.5: 4.5
Sampler
euler
Scheduler
normal
Size
1024 × 10242K and 4K models: their own size
Size step
32
Negative
yes
Latent
EmptyHunyuanImageLatent

Sprint

Steps
2
CFG
14.5 in ScmModelSampling
Sampler
scm
Scheduler
sgm_uniform

Sprint’s guidance goes in through the ScmModelSampling node, while the KSampler stays at CFG 1. Negative prompts do nothing on Sprint. In ComfyUI-SANA, the README suggests about 20 steps at guidance 4.5 for regular models and 2 steps for Sprint. Community tests find CFG 2 to 7 useful depending on style, and longer, more detailed prompts help.

How fast.

GPUModelSizeTime
H100Sprint1024 px0.1 s[1]
RTX 4090Sprint1024 px0.3 s[1]

NVIDIA’s figures from its own code, not ComfyUI.

When it goes wrong.

Grey, black or yellow images
city96’s ExtraModels. Replace it with lawrence-cj/ComfyUI_ExtraModels and pick the Efficient-Large-Model/… presets.
Access to model google/gemma-2-2b-it is restricted or 401 Client Error … Cannot access gated repo
The workflow points at Google’s gated repo. Use Efficient-Large-Model/gemma-2-2b-it in GemmaLoader.
The checkpoint you are trying to load has model type gemma2 but Transformers does not recognize this architecture
Update transformers in ComfyUI’s own Python (python_embeded in the portable build).
ExtraVAELoader: No module named 'diffusers'
Install diffusers into ComfyUI’s Python environment.
Value not in list: vae_type: 'dcae-f32c32-sana-1.1-diffusers'
The node pack is too old. Update NVIDIA’s fork.
size mismatch for pos_embed
The preset and resolution don’t match the checkpoint, or the pack is old. Use the 2K or 4K preset with its own size, and update the fork.
TypeError: 'int' object is not subscriptable in KSampler
Another city96 ExtraModels symptom. Switch to NVIDIA’s fork.
Input type (torch.cuda.HalfTensor) and weight type (torch.HalfTensor) should be the same
One loader is on the CPU and another on the GPU. Put them on the same device.

Sana doesn’t work, I only get a grey image.

GitHub, city96/ComfyUI_ExtraModels

The Gemma models aren’t downloading for the Sana workflow.

GitHub, city96/ComfyUI_ExtraModels

Questions.

Does ComfyUI support Sana natively?

No. Sana needs a custom node pack: NVIDIA’s fork of ComfyUI_ExtraModels (lawrence-cj) for NVIDIA GPUs and the CPU, or ComfyUI-SANA (geoffitect), which also runs on Apple Silicon.

Why does Sana give me a grey image in ComfyUI?

You’re most likely using city96’s original ExtraModels, which the Manager installs. Replace it with NVIDIA’s fork at github.com/lawrence-cj/ComfyUI_ExtraModels and pick the Efficient-Large-Model presets.

What CFG should Sana use?

NVIDIA’s ComfyUI workflows use CFG 2 for the Sana 1.0 models (1024 px, 2K and 4K) and 4.5 for SANA 1.5, both at 28 steps with euler and the normal scheduler. Sprint runs 2 steps with CFG 4.5 set in ScmModelSampling.

Does Sana run on a Mac?

Through ComfyUI-SANA, yes: it wraps diffusers and runs on MPS. The ExtraModels pack is NVIDIA-or-CPU only. Its README names Sprint 0.6B as comfortable on Apple Silicon.

Can I use Sana commercially?

Since 31 July 2026 the image model weights are Apache 2.0; before that they were non-commercial. The Gemma 2 2B encoder is under Google’s Gemma Terms of Use, and the model cards still describe the models as intended for research.

How do I fix “Access to model google/gemma-2-2b-it is restricted”?

Point GemmaLoader at Efficient-Large-Model/gemma-2-2b-it, an ungated copy of the same encoder. Logging in to Hugging Face and accepting Google’s terms also works.

Sources: NVlabs/Sana repository [1], NVIDIA’s ComfyUI workflows, Sana 1.6B model card, ComfyUI_ExtraModels fork, ComfyUI-SANA, settings and quality thread, memory thread, VAE type thread.

HEISS UI

Custom nodes, one button.

Sana isn’t built into ComfyUI. HEISS UI installs the node pack it needs and runs Sana from the same prompt box as everything else.

  • The node pack, installed for you. One Install button, with a pack that also runs on a Mac.
  • Presets with plain names. NVIDIA’s models show up by name and fetch themselves on first use.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • A phone studio. Prompt, browse and share from the couch while the computer renders.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.