What’s different.
Ideogram 4 is the first Ideogram model with open weights, released on 3 June 2026. It’s a 9.3B diffusion transformer trained from scratch, strong at text in images, posters and layouts. Three things set it apart from other local models.
- Two model files. Guidance runs through a second, unconditional model instead of a negative prompt. You need both, and both are 9.3 GB in fp8.
- JSON prompts. It was trained only on structured JSON captions. Plain sentences work, but they trip the safety filter far more often.
- A filter in the weights. Blocked prompts return a grey image that reads “Image blocked by safety filter”. It lives in the model itself, so ComfyUI can’t turn it off, and a different text encoder doesn’t change that.
Files you need.
Comfy-Org repackaged everything for ComfyUI. Ideogram only released fp8 (and an nf4 build ComfyUI can’t load), so there is no bf16 file.
-
Model9.3 GB Download
ideogram4_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ or ideogram4_int8_convrot, 9.6 GB -
Unconditional9.3 GB Download
ideogram4_unconditional_fp8_scaled.safetensorsComfyUI/models/diffusion_models/ required. Match the main model’s format -
Text encoder10.6 GB Download
qwen3vl_8b_fp8_scaled.safetensorsComfyUI/models/text_encoders/ -
VAE0.3 GB Download
flux2-vae.safetensorsComfyUI/models/vae/
The VAE is the same one Flux.2 uses. Comfy-Org also ships nvfp4_mixed versions of both models at 5.5 GB each. They save memory, but running both in nvfp4 costs a lot of quality; using nvfp4 only for the unconditional model is the gentler trade.
models/
├── diffusion_models/
│ ├── ideogram4_fp8_scaled.safetensors
│ └── ideogram4_unconditional_fp8_scaled.safetensors
├── text_encoders/
│ └── qwen3vl_8b_fp8_scaled.safetensors
└── vae/
└── flux2-vae.safetensors
GGUF builds exist (molbal/ideogram-4-gguf, Q8_0 at 10.1 GB per model), but ComfyUI-GGUF doesn’t load Ideogram 4 yet. Stay with the safetensors files for now.
What fits your computer.
At 1024 × 1024. The encoder runs first and is moved out before sampling. Then both models run on every step, so the question is whether 18.6 GB of fp8 fits at once.
- 6 GBOffloads
Reported working on an RTX 3060 6 GB laptop with 16 GB of RAM, after a ComfyUI fix for “Buffer too small”. Slow.
- 8 GBOffloads
The nvfp4 files ran on an RTX 3070 and a 3050 laptop. Keep the main model in fp8 if you can and put only the unconditional in nvfp4.
- 12 GBOffloads
fp8 main model with an nvfp4 unconditional, the rest in system RAM.
- 16 GBOffloads
Both fp8 files, with part of them in system RAM. A 4060 Ti 16 GB is among the speed reports.
- 24 GBFits
Both fp8 models at once, 18.6 GB together.
- MacNot yet
The fp8 files fail on Apple Silicon with a Float8_e4m3fn error, there’s no bf16 release and the GGUFs aren’t recognised. A community patch exists; no official fix.
Set it up.
-
Update ComfyUI
Ideogram 4 needs ComfyUI 0.24.0 or newer. For low-memory cards, take a version with the “Buffer too small” fix (PR #14372) as well. ComfyUI Desktop updates itself; in a manual install:
Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the four files
Main model, unconditional model, text encoder and VAE from the list above. Take the main and unconditional files in the same format, both fp8 or both int8.
-
Put them in their folders
Both models in
diffusion_models, the encoder intext_encoders, the VAE invae. Restart ComfyUI so the loaders see them. -
Open the template
In the template browser, pick Ideogram v4: Text to Image (or the Int8 one). It has two Load Diffusion Model nodes, one per model, feeding a DualModelGuider.
-
Check the loaders and the polish step
Load CLIP gets the Qwen3-VL 8B file with type
ideogram4. CFGOverride should read 3, start 0.7, end 1. The first template shipped with 0.9, which gives soft, blurry images. -
Write a JSON prompt and run
Paste a JSON prompt into the text box (see below). The first run loads everything and takes longer.
JSON prompts.
Ideogram 4 learned from captions with a fixed structure: a short overall description, then a background and a list of elements with bounding boxes on a 0 to 1000 grid, including any text to render. A small one looks like this:
{
"aspect_ratio": "1:1",
"high_level_description": "A minimalist concert poster.",
"compositional_deconstruction": {
"background": "Deep navy water, soft ripples.",
"elements": [
{
"type": "obj",
"bbox": [150, 150, 550, 850],
"desc": "A large flat red sun."
},
{
"type": "text",
"bbox": [750, 100, 900, 900],
"text": "NIGHT TIDE",
"desc": "Bold white sans-serif."
}
]
}
}
Writing these by hand gets old. Ideogram publishes the system prompt it uses to turn a plain idea into JSON, so any chat model can do it for you. In ComfyUI, the Ideogram 4 Prompt Builder KJ node from KJNodes builds them visually.
Settings that work.
Ideogram 4 samples through SamplerCustomAdvanced with its own Ideogram4Scheduler. There’s no negative prompt: the unconditional model takes that role.
- Steps
- 20
- CFG
- 7DualModelGuider
- Polish
- CFG 3 from 0.7CFGOverride
- Sampler
- euler
- Scheduler
- own nodeIdeogram4Scheduler
- mu / std
- 0 / 1.75template: 0.5 / 1.75
- Size
- 1024 × 1024up to 2048, multiples of 16
- Negative
- none
Ideogram’s code defines three presets. The Comfy template uses mu 0.5 at 20 steps; Ideogram’s own 20-step preset uses 0. Both are in use.
| Preset | Steps | CFG 7 then 3 | mu | std |
|---|---|---|---|---|
| Turbo | 12 | 11 + 1 | 0.5 | 1.75 |
| Default | 20 | 18 + 2 | 0 | 1.75 |
| Quality | 48 | 45 + 3 | 0 | 1.5 |
From sampler_configs.py in Ideogram’s repository. Ideogram’s best quality is Quality at 2048 × 2048.
How fast.
| GPU | Attention | Time |
|---|---|---|
| RTX 5090 | PyTorch default | 155 s[1] |
| RTX 5090 | Flash attention | 89 s[1] |
| RTX 5090 | SageAttention | 64 s[1] |
fp8, 20 steps at 2048 × 2048. Flash attention through a KJNodes patch; SageAttention from a build that supports head size 256. Standard SageAttention builds don’t, and ComfyUI falls back to PyTorch attention.
When it goes wrong.
- A grey image: “Image blocked by safety filter”
- The filter inside the model flagged the prompt, often wrongly for plain text. Rewrite it as JSON; there’s no switch to turn it off.
- Soft or blurry results with the template
- CFGOverride starts at 0.9 in older copies. Set start to 0.7 and use plain
euler. ValueError: Buffer too small: needs … bytes, but only has …- A ComfyUI and comfy-aimdo mismatch on low-memory cards. Update ComfyUI, or start it with
--disable-dynamic-vram. TypeError: Trying to convert Float8_e4m3fn to the MPS backend but it does not have support for that dtype.- Apple Silicon can’t load the fp8 files. There’s no official fix yet; see ComfyUI issue #14358 for a community patch.
Error running sage attention: Unsupported head_dim: 256- Standard SageAttention can’t handle Ideogram 4; ComfyUI falls back to PyTorch attention, which is slower but fine.
- Missing nodes such as
Ideogram4SchedulerorDualModelGuider - ComfyUI is older than 0.24.0. Update it.
Do we really need both the main model and the unconditional one? That’s a lot of disk for one model.
A prompt as simple as a red apple came back blocked by the safety filter.
Questions.
Why does Ideogram 4 need two model files?
It uses asymmetric guidance: the main model follows your prompt and a separate unconditional model supplies the pass that a negative prompt would normally drive. A DualModelGuider blends the two on every step, so both files have to be present.
How do I get around “Image blocked by safety filter”?
Write the prompt as structured JSON, which the model was trained on and which is flagged far less often. The filter is part of the weights, so ComfyUI has no setting to turn it off.
Does an abliterated or heretic text encoder remove the filter?
No. The refusal sits in the diffusion model, not the encoder. DreamFast, who publishes the heretic Qwen3-VL encoders, said one makes little difference for Ideogram 4.
Can I use Ideogram 4 commercially?
Not under the open licence. The Ideogram 4 Non-Commercial License allows personal use but excludes using images in or to promote anything that earns money. Commercial use needs an agreement with Ideogram.
Does Ideogram 4 run on a Mac?
Not in plain ComfyUI yet. Ideogram released fp8 weights only, and Apple Silicon can’t load fp8, so ComfyUI stops with a Float8_e4m3fn error. A community patch is discussed in ComfyUI issue #14358.
How much VRAM does Ideogram 4 need?
24 GB holds both fp8 models at once. Smaller cards work by offloading to system RAM: it has run on a 6 GB laptop GPU with 16 GB of RAM, slowly. nvfp4 files for the unconditional model save memory with little quality loss.
Sources: Ideogram 4 repository, prompting guide, Comfy blog: Ideogram 4 in ComfyUI, Comfy-Org Ideogram-4 files, 6 GB report, RTX 5090 timings [1], ComfyUI issue #14358, Heretic issue #351.