Image model

How to run Chroma locally.

Chroma runs in ComfyUI with three files: the model, the T5-XXL text encoder and the Flux.1 VAE. The fp8 model is 9.2 GB and fits a 12 GB graphics card, and GGUF brings it down to 8 GB. It’s Apache 2.0, with a real CFG and a negative prompt that works.

Updated 29 Sep 20268 min read

Maker
lodestones
Released
Aug 2025Chroma1-HD, the finished model
Licence
Apache 2.0commercial use allowed
Memory
8 GB and upas GGUF. fp8 from 12 GB

What Chroma is.

Chroma is a community model by lodestones, funded by donations. It started from Flux.1 Schnell and was retrained into an 8.9B model: a 3.3B layer was swapped for a much smaller one, and the T5 padding is masked out. It uses T5-XXL only, with no CLIP-L, and the Flux.1 VAE.

  • Chroma1-HD is the one to use, uploaded in August 2025.
  • Chroma1-Base is the base checkpoint, published shortly before HD.
  • Chroma1-Flash is a faster version with the guidance baked in. lodestones points to its delta weights for low-step use.
  • Chroma1-Radiance is a different, pixel-space model with no VAE and its own nodes. It’s still in progress and not covered here.

Unlike Flux.1 Dev and Schnell, Chroma uses a real CFG, so the negative prompt works. It prefers natural sentences to tag lists, in the negative prompt too.

Files you need.

If you already run Flux.1, you have the T5 and the VAE. Chroma loads its T5 through a single Load CLIP node with type chroma.

  • Model Chroma1-HD-fp8mixed.safetensors ComfyUI/models/diffusion_models/ full precision: Chroma1-HD.safetensors, 17.8 GB
    9.2 GB Download
  • Text encoder t5xxl_fp8_e4m3fn_scaled.safetensors ComfyUI/models/text_encoders/ or t5xxl_fp16.safetensors, 9.8 GB
    5.2 GB Download
  • VAE ae.safetensors ComfyUI/models/vae/
    0.3 GB Download

The full-precision model is at lodestones/Chroma1-HD. Users find a plain fp8 conversion noticeably worse than full precision; Comfy’s fp8mixed file, a Q8 GGUF or the full file are the ones to use.

ComfyUI/models
models/
├── diffusion_models/
│   └── Chroma1-HD-fp8mixed.safetensors
├── text_encoders/
│   └── t5xxl_fp8_e4m3fn_scaled.safetensors
└── vae/
    └── ae.safetensors

Smaller files: GGUF

silveroxides’ GGUF builds load through the ComfyUI-GGUF nodes and go in diffusion_models.

ModelQ8_0Q6_KQ5_K_MQ4_K_M
Chroma1-HD9.7 GB7.7 GB6.7 GB5.6 GB

Q8_0 is close to the original. Q6_K and Q5_K_M are the usual middle ground.

What fits your computer.

Chroma1-HD at 1024 × 1024 with the fp8 T5. ComfyUI runs the encoder first and moves it aside, so the model file sets the limit.

  • 6 GBTight

    Q4_K_M GGUF (5.6 GB), with the T5 in system RAM.

  • 8 GBOffloads

    Q5_K_M or Q6_K GGUF. Users report it working on 8 GB with offloading.

  • 12 GBFits

    The fp8mixed model or Q8_0 GGUF. One 12 GB user ran the full-precision model too.

  • 16 GBFits

    fp8mixed or Q8_0 comfortably.

  • 24 GBFits

    Full precision (17.8 GB), with the T5 moved aside during sampling.

  • MacUntested

    Use the full-precision file or a GGUF: a Mac can’t compute in fp8. We found no Mac reports for Chroma.

Set it up.

  1. Download the three files

    Model, T5 and VAE from the list above. None of them is gated.

  2. Put them in their folders

    Model in diffusion_models, T5 in text_encoders, VAE in vae. Then restart ComfyUI so the loaders list them.

  3. Open the Chroma template

    In ComfyUI’s template browser, pick Chroma: Text to Image. lodestones also publishes a ComfyUI workflow of their own on the model page.

  4. Check the loaders

    Load Diffusion Model gets the Chroma file. Load CLIP gets the T5 with type chroma. Load VAE gets ae.safetensors. For a GGUF, swap the model loader for Unet Loader (GGUF).

  5. Leave the shift and padding alone

    ModelSamplingAuraFlow at shift 1 is the value the creator intended. T5TokenizerOptions sets the T5 padding; Comfy’s note says a min_padding of 1 is the official way and 0 works too.

  6. Write both prompts and run

    Describe the image in sentences, and write the negative prompt in sentences too, like the template’s “This low quality greyscale unfinished sketch is inaccurate and flawed.”

Settings that work.

These are from lodestones’ own ComfyUI workflow for Chroma1-HD.

Steps
26
CFG
3.8
Sampler
euler
Scheduler
BetaSamplingScheduler
Alpha, beta
0.45, 0.45
Shift
1
Size
1152 × 1152
Negative
yes, in sentences

Comfy’s Chroma template is close: 26 steps, CFG 3.5, the plain beta scheduler and 1024 × 1024. The diffusers example on the model card uses 40 steps at guidance 3.0. Keep CFG between about 3 and 5; higher burns the image. lodestones says any ODE sampler should work, with beta or sigmoid schedules for fewer steps. Stay under about 2 megapixels, where quality starts to fall off.

How fast.

GPUFilesTime
RTX 4070 12 GBQ8_0 GGUF2.6 s per step[1]
RTX 4070 12 GBQ6_K GGUF3.1 s per step[1]
RTX 4070 12 GBfp8 scaled, fp8_e4m3fn_fast1.3 s per step[1]
RTX 50901152 px, 40 steps, 10 LoRAs28 s per image[2]

User reports. The 1.3 s figure used torch.compile and SageAttention as well.

When it goes wrong.

Messy or burnt images
CFG is too high. Stay around 3 to 5. For more contrast without burning, users try CFGNorm or Skimmed CFG.
Soft, low-detail images from an fp8 file
A plain fp8 conversion loses more than you’d expect. Switch to Comfy’s fp8mixed file, a Q8_0 GGUF or full precision.
The prompt is only half followed
Chroma reads sentences better than tags. Rewrite tag lists as a description, and do the same in the negative prompt.
Perturbed Attention Guidance does nothing
PAG only works on UNet models like SDXL. Chroma is a transformer, so the node has no effect.
Faces and details fall apart at large sizes
Above about 2 megapixels, Chroma degrades. Generate near 1 megapixel and upscale.

My image is a mess at CFG 5 and looks better at 12. Isn’t that backwards?

Hugging Face, lodestones/Chroma

What’s the minimum VRAM for Chroma? Can it run on 8 or 12 GB?

Hugging Face, lodestones/Chroma

Questions.

Which text encoder does Chroma use?

Only T5-XXL, the same file as Flux.1, loaded with a single Load CLIP node set to type chroma. Chroma has no CLIP-L. The VAE is the Flux.1 ae.safetensors.

Can I use Chroma commercially?

Yes. Chroma1-HD is Apache 2.0, which allows commercial use.

Does the negative prompt work in Chroma?

Yes. Chroma uses a real CFG, around 3.5 to 3.8, so the negative prompt has an effect. Write it in sentences, like the positive prompt.

What’s the difference between Chroma1-HD, Base and Flash?

HD is the finished model to use. Base is the base checkpoint, published shortly before HD. Flash has the guidance baked in for fewer steps. Radiance is a separate pixel-space model with its own nodes.

Do Flux LoRAs work on Chroma?

Some do, since Chroma started from Flux.1 Schnell. One user reports Flux Dev LoRAs at strength 1 with CFG 4 to 7 and 40 steps. Results vary by LoRA.

Does Chroma run on a Mac?

It should run in ComfyUI on Apple Silicon with the full-precision file (17.8 GB) or a GGUF, since a Mac can’t compute in fp8. We found no Mac reports to confirm speed or memory.

Sources: Chroma1-HD model card, lodestones’ ComfyUI workflow, Comfy-Org Chroma1-HD files, fp8 vs bf16 thread [1], RTX 5090 thread [2], CFG thread, sampler thread, ComfyUI issue #12177.

HEISS UI

Chroma, set up in one pick.

HEISS UI runs Chroma on the ComfyUI you already have. Pick the file and it brings the text encoder and VAE, often ones you already have from Flux.1.

  • Negative prompts that work. Chroma uses real guidance, so the negative prompt box is there and does something.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • One gallery for everything. Search by prompt, model or LoRA, star the keepers, and bring old ComfyUI, AUTOMATIC1111 or Forge folders along.

Chroma1-Radiance isn’t built in; it runs as your own workflow. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.