Hardware guide

What runs on an RTX 5090.

An RTX 5090 runs every model HEISS UI supports, most at full precision, and computes fp8, mxfp8 and nvfp4 natively. The part people trip over is the setup: it needs a PyTorch built for CUDA 12.8 or newer, which current ComfyUI Desktop and the portable build include.

Updated 29 Sep 20264 min read

Memory
32 GB GDDR7
Bandwidth
1792 GB/s512-bit
Architecture
Blackwellfp8, mxfp8, nvfp4
System RAM
64 GBfor Flux.2 Dev in fp8

Short answer

An RTX 5090 runs every model HEISS UI supports, most of them at full precision, and computes fp8, mxfp8 and nvfp4 directly. Flux.1 Dev takes about 8 seconds, Flux.2 Klein about one. Only Flux.2 Dev in fp8 and the biggest MiniMax H3 files go past 32 GB, and those offload a little.

The setup is where people get stuck: it needs a PyTorch built for CUDA 12.8 or newer.

Set up ComfyUI for a 5090.

  1. Use a current ComfyUI

    ComfyUI’s own README asks for a PyTorch built for CUDA 13.0 on the 20 series and newer [3]. The current ComfyUI Desktop and the portable NVIDIA build ship that. An install from 2024 with a CUDA 12.1 or 12.4 PyTorch doesn’t know the card and stops with:

    Error
    CUDA error: no kernel image is available for execution on the device

    Download a fresh portable build or update Desktop rather than patching the old Python.

  2. Keep the default memory handling

    Since March 2026, ComfyUI manages graphics memory by itself on NVIDIA cards. It calls this Dynamic VRAM: model weights stream into the card as they’re needed, and the rest waits in system RAM without going to the pagefile. It’s on by default in current ComfyUI, on Windows and Linux (not WSL).

    That changes the old advice. --lowvram now does nothing while Dynamic VRAM is on, and --novram, --highvram and --gpu-only switch it off. --normalvram is gone, and an old shortcut that still passes it stops ComfyUI with unrecognized arguments.

    On 32 GB there’s nothing to add. --use-sage-attention gave the fastest Flux result in the benchmark thread, once the SageAttention package was installed in ComfyUI’s Python.

  3. Get 64 GB of RAM if you want Flux.2 Dev

    Flux.2 Dev in fp8 is 35.5 GB with an 18 GB text encoder. It runs on a 5090 with part of it in system RAM, and 64 GB makes that smooth.

What runs on 32 GB.

Fits runs from the card. Offloads keeps part of the model in system RAM: it works, slower.

Which files to pick.

Blackwell computes three small formats natively, and Comfy-Org ships several models in each [4]:

  • bf16 or fp16: the reference. On 32 GB it fits for most models; pick it when quality is the point.
  • fp8 and mxfp8: half the size, close to the reference, and faster. _mxfp8 files are made for RTX 50.
  • nvfp4: a quarter of the size and the fastest. Worth comparing on your own prompts before switching for good.

How long it takes.

ModelSettingsTime
Flux.1 Dev fp81024 px, 20 steps8 to 8.8 s[1]
Flux.1 Dev fp8same, with ComfyUI speed-ups5.5 s[1]
Flux.2 Klein 4Bdistilled, 4 stepsabout 1.2 s[2]
Flux.2 Klein 9Bdistilled, 4 stepsabout 2 s[2]
Flux.2 Klein 4B / 9Bbase, 50 steps17 s / 35 s[2]

There are no clean published 5090 numbers for Krea 2, Qwen-Image or Wan 2.2 in ComfyUI yet.

5090 or 4090?

RTX 5090RTX 4090
Graphics memory32 GB24 GB
Bandwidth1792 GB/s1008 GB/s
nvfp4, mxfp8YesNo
Flux.1 Dev fp8, 20 steps8 to 8.8 s[1]11.3 s[1]
Krea 2 at full precisionFitsOffloads
Power575 W450 W

If your models fit in 24 GB, a 4090 gets close. The 5090 earns its price with full-precision Krea 2, Flux.2 Dev with less offloading, Wan 2.2 14B at fp16, and nvfp4.

When it goes wrong.

no kernel image is available for execution on the device
The PyTorch in this ComfyUI predates the RTX 50 series. Install a current ComfyUI Desktop or portable build (CUDA 13.0).
torch.OutOfMemoryError: CUDA out of memory
The model, the image size and the batch didn’t fit together. Generate one image at a time, go down a size, or use the fp8, int8 or GGUF file. When it fails at the very end, the VAE decode ran out: swap VAE Decode for VAE Decode (Tiled). For video, fewer frames help more than a smaller size.
The paging file is too small for this operation to complete. (os error 1455)
Windows ran out of virtual memory while loading a big file. Give the pagefile 32 GB or more, or add RAM.
It suddenly got much slower
The card is full and Windows is borrowing system RAM. Check Shared GPU memory in Task Manager. Close other GPU apps, lower the size, and set Prefer No Sysmem Fallback so it fails fast instead.
unrecognized arguments: --normalvram
An old launch flag. Remove it from the shortcut or the .bat file. Current ComfyUI doesn’t need memory flags on NVIDIA.
The first image takes minutes, the next ones don’t
The first run reads the model from disk. Later runs start warm. An SSD shortens the first one; with 16 GB of RAM it can stay slow, since the model gets pushed out between runs.

Questions.

Why does my RTX 5090 say no kernel image is available?

The ComfyUI install has a PyTorch from before the RTX 50 series. Install the current ComfyUI Desktop or the portable NVIDIA build, which ship CUDA 13.0, instead of patching the old one.

How much system RAM does a 5090 need?

32 GB covers everything that fits the card. For Flux.2 Dev in fp8, bf16 Qwen-Image or the larger MiniMax H3 files, get 64 GB.

Is nvfp4 worth it on a 5090?

For speed, yes: it’s the fastest format the card has, and a quarter the size of bf16. The 5090 has room for full precision too, so compare both on your own prompts.

What about the RTX 5090 Laptop GPU?

It has 24 GB, not 32, and far less power. It runs what the 24 GB desktop cards run, more slowly.

Sources: [1] Flux Dev fp8 GPU benchmark thread, ComfyUI discussion #9002, [2] FLUX.2 Klein in ComfyUI, Comfy blog, [3] ComfyUI README, [4] Krea 2 files, Comfy-Org on Hugging Face, [5] FLUX.2 on RTX GPUs, NVIDIA blog, [6] MiniMax H3 in ComfyUI, Comfy blog, [7] Dynamic VRAM in ComfyUI, Comfy blog, [8] Steam Hardware Survey, August 2026.

HEISS UI

Full size, marked.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. On a 5090 it points you at the full-size models.

  • The biggest versions are marked. It checks your card and your RAM, and marks the largest version of each first model that runs well.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.