Which version.
Anima is a 2B model from CircleStone Labs, made with Comfy Org and built on NVIDIA’s Cosmos-Predict2. It was trained on several million anime images and about 800,000 non-anime artworks, with no photos. All versions are the same size.
- Turbo (
anima-turbo-v1.1) is distilled for 8 to 12 steps at CFG 1. The author recommends starting here: slightly behind Aesthetic on average, and much faster. - Aesthetic (
v1.1, orv1.0bwithout the merged style LoRAs) is the quality fine-tune for finished pictures. - Base (
anima-base-v1.0) has a plain, neutral style. Train LoRAs on this one.
Files you need.
Everything is in the model’s own repository. Only bf16 is official.
-
Model4.2 GB Download
anima-turbo-v1.1.safetensorsComfyUI/models/diffusion_models/ or anima-aesthetic-v1.1 / anima-base-v1.0, 4.2 GB each -
Text encoder1.2 GB Download
qwen_3_06b_base.safetensorsComfyUI/models/text_encoders/ -
VAE0.3 GB Download
qwen_image_vae.safetensorsComfyUI/models/vae/ the Qwen-Image VAE. You may have it already
models/
├── diffusion_models/
│ └── anima-turbo-v1.1.safetensors
├── text_encoders/
│ └── qwen_3_06b_base.safetensors
└── vae/
└── qwen_image_vae.safetensors
Smaller files: GGUF
Community GGUF builds load through the ComfyUI-GGUF Unet Loader. They matter for 4 GB cards; above that, the full file fits anyway.
| Version | Q8_0 | Q6_K | Q5_K_M | Q4_K_M |
|---|---|---|---|---|
| Base v1.0 | 2.2 GB | 1.7 GB | 1.6 GB | 1.4 GB |
The same uploader has Aesthetic and Turbo builds. Its page names the NVIDIA Cosmos licence; CircleStone’s licence applies.
What fits your computer.
At 1024 × 1024. Anima is small, but its attention makes it slower per image than SDXL at the same size.
- 4 GBSlow
A GTX 970 runs the Q8_0 GGUF at under 7.5 s per step with fp16 compute.
- 6 GBFits
The full bf16 model. Comfy’s template puts the whole set at 5.6 GB.
- 8 GBFits
Full model with room for LoRAs.
- 12 GBFits
Comfortable, including hires passes.
- Older GPUsUpdate
GTX 10 and 16 series and RTX 20 cards lack bf16. Before February 2026 ComfyUI ran Anima in fp32 on them, very slowly or with a CUBLAS error. Current ComfyUI uses fp16 there.
- MacUntested
No Apple Silicon reports yet. The weights are bf16, not fp8, so they should load on current PyTorch.
Set it up.
-
Update ComfyUI
Anima is native. Errors about
embed_tokenssize mismatches or'conv_in.weight'almost always mean an old ComfyUI. In a manual install:Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the three files
A model, the Qwen3 0.6B encoder and the Qwen-Image VAE from the list above.
-
Put them in their folders
Model in
diffusion_models, encoder intext_encoders, VAE invae. Restart ComfyUI. -
Open the template
Pick Anima Base v1: Text to Image in the template browser. Load Diffusion Model gets your Anima file, Load CLIP the Qwen3 encoder. Its type is
stable_diffusionin the template; cosmos or qwen make no difference in current ComfyUI. -
Match the sampler to the file
The template is set for Base: 30 steps, CFG 4,
er_sde,simple. For Turbo, change it to 8 to 12 steps and CFG 1. -
Prompt with tags or sentences
Danbooru tags, plain English or both. The template’s prompts show the quality tags; the section below has the rest.
Settings that work.
Base and Aesthetic
- Steps
- 30 to 50
- CFG
- 4 to 5
- Sampler
- er_sde
- Scheduler
- simple
- Size
- 1024 × 1024512² to 1536²
- Negative
- yes
- Text encoder
- Load CLIP
- Latent
- EmptyLatentImage
Turbo
- Steps
- 8 to 12
- CFG
- 1
- Sampler
- euler or er_sde
- Negative
- noneignored at CFG 1
The card calls er_sde a reasonable default: neutral style, flat colours, sharp lines. euler is a bit more creative and suits Turbo and Aesthetic, euler_a gives softer lines. For a painterly look the author likes the beta57 scheduler from the RES4LYF pack.
Prompting.
Lowercase tags with spaces, not underscores (score tags keep theirs). Order: quality and safety, then 1girl or 1boy, character, series, artist, the rest. Put @ before artist names, or they barely register. Weights need to be higher than on SDXL, for example (chibi:2). For plain English, write at least two sentences and name each character before describing them.
masterpiece, best quality, score_7, safe,
worst quality, low quality, score_1, score_2, score_3, artist name, blurry, jpeg artifacts, chromatic aberration
Leave the score_ tags out with Aesthetic. The safety tags are safe, sensitive, nsfw and explicit. Putting ye-pop or deviantart on the first line switches to the non-anime art styles.
How fast.
| GPU | Setup | Speed |
|---|---|---|
| Intel Arc B580 | 1216 × 832 | 1.3 it/s[1] |
| Intel Arc B580 | 1216 × 832, torch.compile | 2 it/s[1] |
| RTX 4060 Ti | vs SDXL | about 3× slower[2] |
| GTX 970 4 GB | Q8_0 GGUF, 1 MP | under 7.5 s/it[3] |
bf16 unless noted. The GTX 970 ran with fp16 compute. Turbo at 8 steps and CFG 1 needs a fraction of the model passes that 30 steps at CFG 4 do, since CFG above 1 runs two passes per step.
When it goes wrong.
size mismatch for model.embed_tokens.weight: copying a param with shape torch.Size([151936, 1024])- ComfyUI is too old to know the Qwen3 0.6B encoder. Update it; the files are fine.
'conv_in.weight'when loading the model- Also an old ComfyUI. Update.
RuntimeError: CUDA error: CUBLAS_STATUS_NOT_SUPPORTEDon an older card- A GPU without bf16. Update ComfyUI (fixed in February 2026), or start it with
--fp16-unet. - Hours per image on a GTX 1060 or similar
- Same cause: fp32 fallback on a card without bf16. Update ComfyUI.
- TAESD preview fails with a size mismatch
- Use
taew2_1for previews. lora key not loaded: lora_te_layers_…- A kohya LoRA with text encoder layers. Fixed in ComfyUI; update.
- Plain, generic-looking pictures from Base
- Base has a neutral default style by design. Add @artist tags, or use Aesthetic or Turbo.
It’s only a 2B model, so why is it slower than SDXL?
Which clip type is right: stable diffusion, cosmos or qwen?
Questions.
Which Anima file should I download?
Start with Turbo: 8 to 12 steps at CFG 1, and only slightly behind Aesthetic on average. Use Aesthetic for finished pictures and Base for training LoRAs or a neutral style.
What sampler and settings does Anima use?
Base and Aesthetic: 30 to 50 steps, CFG 4 to 5, er_sde with the simple scheduler, which is what the model card and Comfy’s templates use. Turbo: 8 to 12 steps at CFG 1, where euler also works well.
Why is Anima slower than SDXL?
It’s a diffusion transformer, and its attention costs more per step than SDXL’s UNet. On an RTX 4060 Ti it takes about three times as long. Turbo, torch.compile and int8 conversions are the usual ways to speed it up.
Can I sell images made with Anima?
Yes. The weights are non-commercial, but the licence lets you use outputs for any purpose, including commercial ones. Running Anima as a paid service needs a licence from CircleStone Labs.
Which clip type do I pick for the Anima text encoder?
Comfy’s template uses stable_diffusion. In current ComfyUI, cosmos and qwen give the same result, so it doesn’t matter.
Does Anima run on an old GPU?
Yes, once ComfyUI is up to date. Cards without bf16, such as the GTX 10 and 16 series, used to fall back to fp32 and take minutes or hours per image; since February 2026 ComfyUI uses fp16 on them. A GTX 970 with 4 GB runs the Q8_0 GGUF.
Sources: Anima model card, licence, ComfyUI Anima tutorial, Arc B580 report [1], SDXL comparison [2], GTX 970 and 1060 thread [3], ComfyUI issue #12477, sampler and hires thread.