Which model.
All three share one architecture: a 2B flow transformer from Shanghai AI Laboratory’s Alpha-VLLM team, with Gemma 2 2B as its text encoder and the 16-channel Flux.1 VAE.
- Lumina Image 2.0 (January 2025) is the general base model.
- Neta Lumina (v1.0, July 2025) is Neta.art Lab’s anime fine-tune, trained on more than 13 million anime images. It reads Danbooru tags and plain text in Chinese, English and Japanese.
- NetaYume Lumina is a community fine-tune of Neta Lumina by duongve, now at v4.0 (December 2025). Comfy’s anime template uses its v3.5, and v4 also understands JSON and XML structured prompts.
For anime, start with NetaYume. For anything else, Lumina Image 2.0.
Files you need.
The easy way is one all-in-one checkpoint that goes in checkpoints. Split files use about 2 GB less VRAM.
All-in-one
-
NetaYume10.6 GB Download
NetaYume_v4_all_in_one.safetensorsComfyUI/models/checkpoints/ Comfy’s template uses NetaYumev35_pretrained_all_in_one -
Neta Lumina10.6 GB Download
neta-lumina-v1.0-all-in-one.safetensorsComfyUI/models/checkpoints/ -
Lumina 2.010.6 GB Download
lumina_2.safetensorsComfyUI/models/checkpoints/
Split files (Lumina Image 2.0)
-
Model5.2 GB Download
lumina_2_model_bf16.safetensorsComfyUI/models/diffusion_models/ NetaYume and Neta Lumina have their own in Unet/ folders, 5.2 GB each -
Text encoder5.2 GB Download
gemma_2_2b_fp16.safetensorsComfyUI/models/text_encoders/ the same file for all three -
VAE0.3 GB Download
ae.safetensorsComfyUI/models/vae/ the Flux.1 VAE
models/
├── checkpoints/
│ └── NetaYume_v4_all_in_one.safetensors (all-in-one)
├── diffusion_models/
│ └── lumina_2_model_bf16.safetensors (or split files)
├── text_encoders/
│ └── gemma_2_2b_fp16.safetensors
└── vae/
└── ae.safetensors
Smaller files: GGUF
ComfyUI-GGUF loads Lumina 2 models through its Unet Loader. The Gemma 2 encoder stays a safetensors file.
| Model | Q8_0 | Q5_K_M | Q4_K_M |
|---|---|---|---|
| Lumina 2.0 | 2.8 GB | 1.8 GB | 1.5 GB |
| NetaYume v2 plus | 3.1 GB | 2.3 GB | 2.1 GB |
calcuis also has a Neta Lumina Q4_K_M at 2.0 GB. No official fp8 release exists.
What fits your computer.
At 1024 × 1024. These models compute in bf16 only; forcing fp16 gives black images.
- 6 GBTight
A GGUF model with the 5.2 GB Gemma encoder partly in system RAM.
- 8 GBFits
Split files. The NetaYume author puts them at about 8 GB, and Neta’s card asks for 8 GB or more.
- 12 GBFits
The all-in-one checkpoint, about 10 GB in use.
- 16 GBFits
All-in-one with room for LoRAs and larger sizes.
- RTX 20 and olderSlow
No bf16 on these cards, so ComfyUI runs in fp32. An RTX 2060 took 14 s per step.
- MacLikely
An early Apple Silicon crash was fixed in ComfyUI in February 2025. No Mac timings found.
Set it up.
-
Update ComfyUI
Lumina 2 has been native since February 2025. Use an official, current ComfyUI; old or repackaged bundles fail with “Could not detect model type”.
Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download a checkpoint
For a first run, one all-in-one file from the list above into
checkpoints. Restart ComfyUI. -
Open a workflow
For NetaYume, pick NetaYume Lumina Text to Image in the template browser. For Lumina Image 2.0, drag in the workflow image from ComfyUI’s Lumina 2 example.
-
Check the shift
Both workflows put ModelSamplingAuraFlow before the KSampler: shift 6 for Lumina Image 2.0, 4 in the NetaYume template.
-
Keep the system prompt
The prompt box starts with a fixed sentence ending in
<Prompt Start>. Write your prompt after it and leave the sentence as it is. For Lumina, the CLIP Text Encode (Lumina 2) node can add it for you instead. -
Run
The first run loads Gemma and the model. The next ones are quicker.
The system prompt.
Lumina 2 was trained with a system prompt in front of every caption. It still draws without one, but Comfy’s workflows and the official demo use it. ComfyUI’s plain text encoder sends exactly what you type, so the prefix has to be in the box. These are the ones the Comfy workflows use.
You are an assistant designed to generate superior images with the superior degree of image-text alignment based on textual prompts or user prompts. <Prompt Start>
You are an assistant designed to generate high quality anime images based on textual prompts. <Prompt Start>
You are an assistant designed to generate low-quality images based on textual prompts <Prompt Start>
For NetaYume, put artist tags as @name, mix Danbooru tags with sentences, and see Neta’s prompt book for more. For readable text in the image, NetaYume’s notes suggest a different prefix that tells the model to render quoted text verbatim.
Settings that work.
Lumina Image 2.0
- Steps
- 36Comfy example: 25. Card: 50
- CFG
- 4
- Sampler
- res_multistep
- Scheduler
- simple
- Shift
- 6ModelSampling
AuraFlow - Size
- 1024 × 1024
- Negative
- yes
- System prompt
- superior
Comfy’s example workflow runs 25 steps, and its own note says the official way to sample is shift 6 with 36 steps. Alpha-VLLM’s diffusers card uses 50. To mimic the official CFG truncation and normalisation, ComfyUI has a RenormCFG node.
Neta Lumina and NetaYume
- Steps
- 30NetaYume card: 40 to 50
- CFG
- 4 to 5.5
- Sampler
- res_multistep
- Scheduler
- linear_quadratictemplate: simple
- Shift
- 4ModelSampling
AuraFlow - Size
- 1024 × 1024or 1024 × 1536
- Negative
- yes, with prefix
- System prompt
- anime
The Neta and NetaYume cards recommend linear_quadratic; Comfy’s NetaYume template uses simple, which also works. euler_ancestral is the cards’ other sampler. There’s no working few-step version of these models; a community CFG-distill LoRA runs at 20 steps and CFG 1.
When it goes wrong.
ERROR: Could not detect model type of: …lumina_2.safetensors- ComfyUI is too old, or a repackaged bundle. Use a current official ComfyUI and Comfy’s example workflow.
repeat(): Not supported for complex yeton a Mac- Fixed in ComfyUI in February 2025. Update.
- Black images with
--fp16-unet - Lumina 2 overflows in fp16. Remove the flag and let it run in bf16.
- NetaYume draws odd styles or garbage
- SageAttention doesn’t agree with it. Start ComfyUI without SageAttention.
- ComfyUI crashes in the KSampler with a very long prompt
- A known issue with extremely long prompts. Shorten it.
- Results don’t match the official demo
- Use shift 6 at 36 steps and keep the system prompt. The noise also differs between CPU and GPU.
What are the recommended settings for ComfyUI? The card only lists diffusers options.
How much video memory does NetaYume need?
Questions.
Do I need the system prompt for Lumina Image 2.0?
It works without one, but it was trained with one, and matching the official demo needs it. In a plain CLIP Text Encode node, type it before your prompt, ending with <Prompt Start>. The CLIP Text Encode (Lumina 2) node adds it for you.
What’s the difference between Neta Lumina and NetaYume?
Neta Lumina is Neta.art Lab’s anime fine-tune of Lumina Image 2.0. NetaYume is a community fine-tune of Neta Lumina with more training, now at v4. Both use the same files around the model and the same anime system prompt.
How much VRAM do NetaYume and Lumina 2 need?
About 8 GB with the split files and about 10 GB with the all-in-one checkpoint, according to the NetaYume author. GGUF model files bring that down further.
Why are my Lumina 2 images black?
Lumina 2 can’t compute in fp16. Remove --fp16-unet or any setting that forces fp16 and let ComfyUI use bf16. Cards without bf16 fall back to fp32, which works but is slow.
Which scheduler should NetaYume use?
The Neta Lumina and NetaYume cards recommend res_multistep with linear_quadratic. Comfy’s template uses simple, which also works.
Can I use Lumina Image 2.0 commercially?
Lumina Image 2.0 and Neta Lumina are Apache 2.0. NetaYume is Apache 2.0 on Hugging Face, but its Civitai permissions don’t include selling images. The Gemma 2 2B encoder is under Google’s Gemma Terms of Use.
Sources: Lumina Image 2.0 model card, Neta Lumina model card, NetaYume model card, ComfyUI Lumina 2 example, demo comparison thread, fp16 and RTX 2060 thread, NetaYume sampling thread, ComfyUI issue #6700.