Short answer
An RTX 5090 runs every model HEISS UI supports, most of them at full precision, and computes fp8, mxfp8 and nvfp4 directly. Flux.1 Dev takes about 8 seconds, Flux.2 Klein about one. Only Flux.2 Dev in fp8 and the biggest MiniMax H3 files go past 32 GB, and those offload a little.
The setup is where people get stuck: it needs a PyTorch built for CUDA 12.8 or newer.
Set up ComfyUI for a 5090.
Use a current ComfyUI
ComfyUI’s own README asks for a PyTorch built for CUDA 13.0 on the 20 series and newer [3]. The current ComfyUI Desktop and the portable NVIDIA build ship that. An install from 2024 with a CUDA 12.1 or 12.4 PyTorch doesn’t know the card and stops with:
ErrorCUDA error: no kernel image is available for execution on the device
Download a fresh portable build or update Desktop rather than patching the old Python.
Keep the default memory handling
Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).
That changes the old advice.
--lowvramnow does nothing while Dynamic VRAM is on, and--novram,--highvramand--gpu-onlyswitch it off.--normalvramis gone, and an old shortcut that still passes it stops ComfyUI withunrecognized arguments.On 32 GB there’s nothing to add.
--use-sage-attentiongave the fastest Flux result in the benchmark thread, once the SageAttention package was installed in ComfyUI’s Python.Get 64 GB of RAM if you want Flux.2 Dev
Flux.2 Dev in fp8 is 35.5 GB with an 18 GB text encoder. It runs on a 5090 with part of it in system RAM, and 64 GB makes that smooth.
What runs on 32 GB.
Fits runs from the card. Offloads keeps part of the model in system RAM: it works, slower.
- Krea 2 Turbo and Rawbf16 · 26.3 GBFits
Full precision, or mxfp8 (13.5 GB) and nvfp4 (7.7 GB) for speed.
- Qwen-Imagefp8 20.4 GB · nvfp4 19.8 GBFits
bf16 (40.9 GB) offloads.
- HiDream I1fp8 · 17.1 GBFits
bf16 (34.2 GB) offloads a little.
- Flux.1 Devbf16 · 23.8 GBFits
About 8 s in fp8.
- Flux.2 Klein, Z-Image, Qwen-Image 2.1, Ideogram 4, Chroma, SD 3.5, ERNIEfull precisionFits
bf16 files, with room left over.
- Wan 2.2 14Bfp16 · 2 × 28.6 GBFits
One half at a time, at full precision.
- Wan 2.2 5B, HunyuanVideo 1.5fp16Fits
720p and longer clips.
- MiniMax H3pruned int8 · 21 GBFits
The pruned bf16 (40.2 GB) and the full int8 (34 GB) offload.
- Flux.2 Devfp8 · 35.5 GBOffloads
A little over 32 GB, plus its text encoder. 64 GB of RAM.
Which files to pick.
Blackwell computes three small formats natively, and Comfy-Org ships several models in each [4]:
- bf16 or fp16: the reference. On 32 GB it fits for most models; pick it when quality is the point.
- fp8 and mxfp8: half the size, close to the reference, and faster.
_mxfp8files are made for RTX 50. - nvfp4: a quarter of the size and the fastest. Worth comparing on your own prompts before switching for good.
How long it takes.
| Model | Settings | Time |
|---|---|---|
| Flux.1 Dev fp8 | 1024 px, 20 steps | 8 to 8.8 s[1] |
| Flux.1 Dev fp8 | same, with ComfyUI speed-ups | 5.5 s[1] |
| Flux.2 Klein 4B | distilled, 4 steps | about 1.2 s[2] |
| Flux.2 Klein 9B | distilled, 4 steps | about 2 s[2] |
| Flux.2 Klein 4B / 9B | base, 50 steps | 17 s / 35 s[2] |
There are no clean published 5090 numbers for Krea 2, Qwen-Image or Wan 2.2 in ComfyUI yet.
5090 or 4090?
| RTX 5090 | RTX 4090 | |
|---|---|---|
| Graphics memory | 32 GB | 24 GB |
| Bandwidth | 1792 GB/s | 1008 GB/s |
| nvfp4, mxfp8 | Yes | No |
| Flux.1 Dev fp8, 20 steps | 8 to 8.8 s[1] | 11.3 s[1] |
| Krea 2 at full precision | Fits | Offloads |
| Power | 575 W | 450 W |
If your models fit in 24 GB, a 4090 gets close. The 5090 earns its price with full-precision Krea 2, Flux.2 Dev with less offloading, Wan 2.2 14B at fp16, and nvfp4.
When it goes wrong.
no kernel image is available for execution on the device- The PyTorch in this ComfyUI predates the RTX 50 series. Install a current ComfyUI Desktop or portable build (CUDA 13.0).
torch.OutOfMemoryError: CUDA out of memory- The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)- Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
- It suddenly got much slower
- The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram- An old launch flag. Remove it from the shortcut or the
.batfile. Current ComfyUI doesn’t need memory flags on NVIDIA. - The first image takes minutes, the next ones don’t
- The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.
Questions.
Why does my RTX 5090 say no kernel image is available?
The ComfyUI install has a PyTorch from before the RTX 50 series. Install the current ComfyUI Desktop or the portable NVIDIA build, which ship CUDA 13.0, instead of patching the old one.
How much system RAM does a 5090 need?
32 GB covers everything that fits the card. For Flux.2 Dev in fp8, bf16 Qwen-Image or the larger MiniMax H3 files, get 64 GB.
Is nvfp4 worth it on a 5090?
For speed, yes: it’s the fastest format the card has, and a quarter the size of bf16. The 5090 has room for full precision too, so compare both on your own prompts.
What about the RTX 5090 Laptop GPU?
It has 24 GB, not 32, and far less power. It runs what the 24 GB desktop cards run, more slowly.
Sources: [1] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [2] FLUX.2 Klein in ComfyUI, Comfy blog, [3] ComfyUI README, [4] Krea 2 files, Comfy-Org on Hugging Face, [5] FLUX.2 on RTX GPUs, NVIDIA blog, [6] MiniMax H3 in ComfyUI, Comfy blog, [7] Dynamic VRAM in ComfyUI, Comfy blog, [8] Steam Hardware Survey, August 2026.