Base or Turbo.
ERNIE-Image is Baidu’s 8B single-stream diffusion transformer, public since April 2026. It shares the Flux.2 VAE and latent, and it’s good at posters and text. Both versions use the same files around the model.
- ERNIE-Image is the full model. Baidu runs it at 50 steps and CFG 4, and it follows a negative prompt.
- ERNIE-Image-Turbo is distilled for 8 steps at CFG 1, with no negative prompt. It’s the faster place to start.
Files you need.
Comfy-Org’s repack has everything in bf16. There’s no fp8 or int8 version from Comfy-Org.
-
Model16.1 GB Download
ernie-image-turbo.safetensorsComfyUI/models/diffusion_models/ or ernie-image.safetensors, the full model, 16.1 GB -
Text encoder7.7 GB Download
ministral-3-3b.safetensorsComfyUI/models/text_encoders/ -
VAE0.3 GB Download
flux2-vae.safetensorsComfyUI/models/vae/ -
Optional6.9 GB Download
ernie-image-prompt-enhancer.safetensorsComfyUI/models/text_encoders/ a language model that rewrites prompts. Not a second encoder
models/
├── diffusion_models/
│ └── ernie-image-turbo.safetensors
├── text_encoders/
│ ├── ministral-3-3b.safetensors
│ └── ernie-image-prompt-enhancer.safetensors (optional)
└── vae/
└── flux2-vae.safetensors
Smaller files: GGUF
unsloth publishes both versions as GGUF, built with ComfyUI-GGUF’s tools. Load them with the ComfyUI-GGUF Unet Loader. The Ministral 3 encoder has no GGUF that ComfyUI-GGUF reads yet, so it stays at 7.7 GB.
Q8_0 is close to the original. Q5_K_M and Q6_K are the usual middle ground.
What fits your computer.
At 1024 × 1024. The encoder runs first and ComfyUI moves it out before sampling, so the model file sets the limit. Only Baidu’s 24 GB figure is official; the other rows follow from the file sizes.
- 8 GBTight
GGUF Q4_K_M (5.0 GB). The 7.7 GB encoder partly waits in system RAM.
- 12 GBFits
GGUF Q5_K_M or Q6_K (5.9 to 6.8 GB).
- 16 GBFits
GGUF Q8_0 (8.7 GB). The bf16 file is 16.1 GB and offloads part of itself.
- 24 GBFits
The bf16 files, as Baidu intends.
- RTX 20 and olderSlow
ComfyUI runs ERNIE-Image in bf16 or fp32 only. Cards without bf16 fall back to fp32, which is slower and needs more memory.
- MacUntested
No ComfyUI reports on Apple Silicon yet. The bf16 files avoid the fp8 problems Macs have with other models.
Set it up.
-
Update ComfyUI
ERNIE-Image needs a ComfyUI from mid-April 2026 or later, with
EmptyFlux2LatentImageand the ERNIE text encoder. In a manual install:Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the files
Model, encoder and VAE from the list above. Leave the prompt enhancer out if you don’t want it.
-
Open the template
Pick Ernie Image Turbo: Text To Image or Ernie Image: Text to Image. Load CLIP gets
ministral-3-3bwith typeflux2. -
Add the shift
The template has no shift node, so ComfyUI uses its default of 3. Baidu trained and runs with 4. Put a ModelSamplingSD3 node between Load Diffusion Model and the KSampler and set
shiftto 4. -
Set the steps
Turbo stays at 8 steps and CFG 1. For the full model, the template uses 20 steps and notes “Try 50 steps” for the original setting; Baidu’s card uses 50.
-
Skip or keep the enhancer, then run
If the template’s enhancer nodes are active and you’d rather not use them, select them and press Ctrl+B to bypass them, so your prompt goes straight to the encoder.
Settings that work.
ERNIE-Image
- Steps
- 50template: 20
- CFG
- 4
- Sampler
- euler
- Scheduler
- simple
- Shift
- 4ModelSamplingSD3. Default: 3
- Size
- 1024 × 1024
- Negative
- yes
- Latent
- EmptyFlux2LatentImage
ERNIE-Image-Turbo
- Steps
- 8
- CFG
- 1
- Shift
- 4
- Negative
- none
Turbo uses the same sampler, scheduler and latent. Baidu lists these sizes: 1024 × 1024, 848 × 1264, 1264 × 848, 768 × 1376, 1376 × 768, 896 × 1200 and 1200 × 896. When people complain about anatomy, the usual advice is one of Baidu’s sizes, the full step count and the enhancer.
The prompt enhancer.
The enhancer is a separate 3B language model. The template runs it through TextGenerate to expand a short prompt into a detailed one, then encodes the result with Ministral 3 as usual. Baidu’s own examples use it.
When it goes wrong.
- “Dual clips not working” after loading the enhancer as a second encoder
- The enhancer isn’t an encoder. Use one Load CLIP with
ministral-3-3b, typeflux2, and run the enhancer throughTextGenerateor not at all. - A stall of several minutes at “model initializing” after the enhancer
- The enhancer is still in memory. Add an unload or garbage-collect node between the enhancer and the sampler.
- Missing
EmptyFlux2LatentImageor an unknown text encoder type - ComfyUI is too old for ERNIE-Image. Update it.
- Results look off next to Baidu’s examples
- Check the shift is 4, not the default 3, and that the full model runs 50 steps at one of Baidu’s sizes.
- Out of memory with the bf16 file on 16 GB
- Use a GGUF build: Q8_0 is 8.7 GB and close to the original.
Is the prompt enhancer required, and why does it change what I asked for?
Which timestep shift should ERNIE-Image use? Baidu’s answer: 4.
Questions.
Do I need the ERNIE-Image prompt enhancer?
No. It’s an optional language model that rewrites your prompt before encoding. It helps short prompts, but its default system prompt can change what you asked for, and long, detailed prompts don’t need it.
What shift should ERNIE-Image use?
4. Baidu’s scheduler config and one of its developers both give 4. ComfyUI defaults to 3 when no shift node is present, so add ModelSamplingSD3 with shift 4.
How many steps does ERNIE-Image need?
The full model runs at 50 steps and CFG 4 on Baidu’s card; Comfy’s template uses 20 to save time. Turbo needs 8 steps at CFG 1.
How much VRAM does ERNIE-Image need?
Baidu says it runs on consumer GPUs with 24 GB. With unsloth’s GGUF builds, from 5.0 GB for Q4_K_M to 8.7 GB for Q8_0, it fits 8 to 16 GB cards.
Can I use ERNIE-Image commercially?
Yes. ERNIE-Image and ERNIE-Image-Turbo are both released under Apache 2.0.
Sources: ERNIE-Image model card, ERNIE-Image-Turbo model card, scheduler config (shift 4), ComfyUI ERNIE-Image tutorial, Comfy-Org ERNIE-Image files, dual clip thread, enhancer stall thread, unsloth GGUF.