Image model

How to run ERNIE-Image locally.

ERNIE-Image runs natively in ComfyUI with three files: the 16.1 GB model, a Ministral 3 3B text encoder and the Flux.2 VAE. Baidu says 24 GB of VRAM; GGUF builds from 5.0 GB bring it to smaller cards. It’s Apache 2.0, and the Turbo version needs eight steps.

Updated 29 Sep 20268 min read

Maker
Baidu
Released
Apr 2026Base and Turbo, in ComfyUI from 14 April
Licence
Apache 2.0both versions, commercial use allowed
Memory
24 GBBaidu’s figure. GGUF from 5.0 GB

Base or Turbo.

ERNIE-Image is Baidu’s 8B single-stream diffusion transformer, public since April 2026. It shares the Flux.2 VAE and latent, and it’s good at posters and text. Both versions use the same files around the model.

  • ERNIE-Image is the full model. Baidu runs it at 50 steps and CFG 4, and it follows a negative prompt.
  • ERNIE-Image-Turbo is distilled for 8 steps at CFG 1, with no negative prompt. It’s the faster place to start.

Files you need.

Comfy-Org’s repack has everything in bf16. There’s no fp8 or int8 version from Comfy-Org.

  • Model ernie-image-turbo.safetensors ComfyUI/models/diffusion_models/ or ernie-image.safetensors, the full model, 16.1 GB
    16.1 GB Download
  • Text encoder ministral-3-3b.safetensors ComfyUI/models/text_encoders/
    7.7 GB Download
  • VAE flux2-vae.safetensors ComfyUI/models/vae/
    0.3 GB Download
  • Optional ernie-image-prompt-enhancer.safetensors ComfyUI/models/text_encoders/ a language model that rewrites prompts. Not a second encoder
    6.9 GB Download
ComfyUI/models
models/
├── diffusion_models/
│   └── ernie-image-turbo.safetensors
├── text_encoders/
│   ├── ministral-3-3b.safetensors
│   └── ernie-image-prompt-enhancer.safetensors   (optional)
└── vae/
    └── flux2-vae.safetensors

Smaller files: GGUF

unsloth publishes both versions as GGUF, built with ComfyUI-GGUF’s tools. Load them with the ComfyUI-GGUF Unet Loader. The Ministral 3 encoder has no GGUF that ComfyUI-GGUF reads yet, so it stays at 7.7 GB.

VersionQ8_0Q6_KQ5_K_MQ4_K_M
Turbo8.7 GB6.8 GB5.9 GB5.0 GB
Full8.7 GB6.8 GB5.9 GB5.0 GB

Q8_0 is close to the original. Q5_K_M and Q6_K are the usual middle ground.

What fits your computer.

At 1024 × 1024. The encoder runs first and ComfyUI moves it out before sampling, so the model file sets the limit. Only Baidu’s 24 GB figure is official; the other rows follow from the file sizes.

  • 8 GBTight

    GGUF Q4_K_M (5.0 GB). The 7.7 GB encoder partly waits in system RAM.

  • 12 GBFits

    GGUF Q5_K_M or Q6_K (5.9 to 6.8 GB).

  • 16 GBFits

    GGUF Q8_0 (8.7 GB). The bf16 file is 16.1 GB and offloads part of itself.

  • 24 GBFits

    The bf16 files, as Baidu intends.

  • RTX 20 and olderSlow

    ComfyUI runs ERNIE-Image in bf16 or fp32 only. Cards without bf16 fall back to fp32, which is slower and needs more memory.

  • MacUntested

    No ComfyUI reports on Apple Silicon yet. The bf16 files avoid the fp8 problems Macs have with other models.

Set it up.

  1. Update ComfyUI

    ERNIE-Image needs a ComfyUI from mid-April 2026 or later, with EmptyFlux2LatentImage and the ERNIE text encoder. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the files

    Model, encoder and VAE from the list above. Leave the prompt enhancer out if you don’t want it.

  3. Open the template

    Pick Ernie Image Turbo: Text To Image or Ernie Image: Text to Image. Load CLIP gets ministral-3-3b with type flux2.

  4. Add the shift

    The template has no shift node, so ComfyUI uses its default of 3. Baidu trained and runs with 4. Put a ModelSamplingSD3 node between Load Diffusion Model and the KSampler and set shift to 4.

  5. Set the steps

    Turbo stays at 8 steps and CFG 1. For the full model, the template uses 20 steps and notes “Try 50 steps” for the original setting; Baidu’s card uses 50.

  6. Skip or keep the enhancer, then run

    If the template’s enhancer nodes are active and you’d rather not use them, select them and press Ctrl+B to bypass them, so your prompt goes straight to the encoder.

Settings that work.

ERNIE-Image

Steps
50template: 20
CFG
4
Sampler
euler
Scheduler
simple
Shift
4ModelSamplingSD3. Default: 3
Size
1024 × 1024
Negative
yes
Latent
EmptyFlux2LatentImage

ERNIE-Image-Turbo

Steps
8
CFG
1
Shift
4
Negative
none

Turbo uses the same sampler, scheduler and latent. Baidu lists these sizes: 1024 × 1024, 848 × 1264, 1264 × 848, 768 × 1376, 1376 × 768, 896 × 1200 and 1200 × 896. When people complain about anatomy, the usual advice is one of Baidu’s sizes, the full step count and the enhancer.

The prompt enhancer.

The enhancer is a separate 3B language model. The template runs it through TextGenerate to expand a short prompt into a detailed one, then encodes the result with Ministral 3 as usual. Baidu’s own examples use it.

When it goes wrong.

“Dual clips not working” after loading the enhancer as a second encoder
The enhancer isn’t an encoder. Use one Load CLIP with ministral-3-3b, type flux2, and run the enhancer through TextGenerate or not at all.
A stall of several minutes at “model initializing” after the enhancer
The enhancer is still in memory. Add an unload or garbage-collect node between the enhancer and the sampler.
Missing EmptyFlux2LatentImage or an unknown text encoder type
ComfyUI is too old for ERNIE-Image. Update it.
Results look off next to Baidu’s examples
Check the shift is 4, not the default 3, and that the full model runs 50 steps at one of Baidu’s sizes.
Out of memory with the bf16 file on 16 GB
Use a GGUF build: Q8_0 is 8.7 GB and close to the original.

Is the prompt enhancer required, and why does it change what I asked for?

Hugging Face, Comfy-Org ERNIE-Image

Which timestep shift should ERNIE-Image use? Baidu’s answer: 4.

Hugging Face, baidu/ERNIE-Image

Questions.

Do I need the ERNIE-Image prompt enhancer?

No. It’s an optional language model that rewrites your prompt before encoding. It helps short prompts, but its default system prompt can change what you asked for, and long, detailed prompts don’t need it.

What shift should ERNIE-Image use?

4. Baidu’s scheduler config and one of its developers both give 4. ComfyUI defaults to 3 when no shift node is present, so add ModelSamplingSD3 with shift 4.

How many steps does ERNIE-Image need?

The full model runs at 50 steps and CFG 4 on Baidu’s card; Comfy’s template uses 20 to save time. Turbo needs 8 steps at CFG 1.

How much VRAM does ERNIE-Image need?

Baidu says it runs on consumer GPUs with 24 GB. With unsloth’s GGUF builds, from 5.0 GB for Q4_K_M to 8.7 GB for Q8_0, it fits 8 to 16 GB cards.

Can I use ERNIE-Image commercially?

Yes. ERNIE-Image and ERNIE-Image-Turbo are both released under Apache 2.0.

Sources: ERNIE-Image model card, ERNIE-Image-Turbo model card, scheduler config (shift 4), ComfyUI ERNIE-Image tutorial, Comfy-Org ERNIE-Image files, dual clip thread, enhancer stall thread, unsloth GGUF.

HEISS UI

ERNIE-Image, without the wiring.

HEISS UI runs ERNIE-Image and ERNIE-Image Turbo on the ComfyUI you already have. Pick the file and write a prompt.

  • Turbo or full, spotted for you. Each gets settings that work, so there’s nothing to look up.
  • Missing parts, shown first. Each one listed with its size and a button. Get all checks free space, and downloads resume and are verified.
  • Start from a picture. Add a start image and a slider sets how far the result may move from it.
  • Upscale and compare. One click to upscale with SeedVR2, then drag a slider to see what changed.

For Baidu’s own setting, raise the steps in Advanced. Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.