Hardware guide

What runs on a Mac with Apple Silicon.

On a Mac, count the unified memory and take about 70%: that’s what the GPU can use. 16 GB runs SDXL, 18 to 24 GB adds Flux.2 Klein and Z-Image, 48 GB runs Krea 2 at full precision. This page covers the memory math, which files to pick, and how long it takes.

Updated 29 Sep 20264 min read

GPU share
about 70%Metal’s own limit is 74 to 78%
fp8 files
no savingthey load at full precision
Smaller files
GGUFQ4_K_M to Q8_0
Install
ComfyUI DesktopM1 or later, macOS 13+

Short answer

Count the unified memory, then take about 70%: that’s what the GPU can use comfortably. 16 GB runs SDXL, 18 to 24 GB adds Flux.2 Klein and Z-Image, 48 GB runs Krea 2 at full precision, and 128 GB runs Flux.2 Dev.

It works, and it’s slower than an NVIDIA card. The same SDXL image took about 40 seconds on an M1 Max and 16 on an RTX 4060 Laptop.

How much memory the GPU gets.

A Mac shares one memory between CPU and GPU, and macOS lets the GPU use part of it. Metal calls that part the recommended working set: on a 24 GB M4 Pro it’s 19.07 GB, about 74%, and reports elsewhere put it at 75 to 78% [1]. ComfyUI itself reports the whole memory, so it can’t tell you. Plan with about 70%, which leaves macOS and your other apps some room.

MemoryAbout 70% for the GPURuns wellHEISS UI marks
8 GB5.6 GBSD 1.5, SDXL Lightning at a squeezeNothing
16 GB11.2 GBSDXL, Z-Image as a GGUF, Flux.2 Klein 4B tightlySDXL
18 to 24 GB12.6 to 16.8 GBFlux.2 Klein 4B, Z-Image, Flux.1 Dev as a GGUF Q4 to Q6Flux.2 Klein 4B, SDXL
32 to 36 GB22 to 25 GBQwen-Image 2.1, Flux.1 Dev GGUF Q8, Wan 2.2 5B, HunyuanVideo 1.5Flux.2 Klein 4B, SDXL
48 to 64 GB34 to 45 GBKrea 2 at full precision, Qwen-Image, Wan 2.2 14B in GGUF halvesKrea 2 Turbo, Flux.2 Klein 4B, SDXL
96 to 128 GB67 to 90 GBFlux.2 Dev, MiniMax H3, Wan 2.2 14B at full precisionKrea 2 Turbo; Flux.2 Dev at 128 GB

Raising the limit.

On macOS 14 and later, iogpu.wired_limit_mb sets how much the GPU may use. It resets at restart. Leave several GB for macOS: the example gives a 32 GB Mac 26 GB for the GPU.

Terminal
sudo sysctl iogpu.wired_limit_mb=26624

HEISS UI reads this value when ComfyUI runs on the same Mac, and counts it instead of the 70%.

Skip fp8 files on a Mac.

Apple’s GPU can’t compute in fp8, and ComfyUI loads fp8 weights as fp16 or bf16 on a Mac. An fp8 download saves disk space, then takes as much memory as the full-precision file [1]. Some fp8 files don’t load at all and stop with Trying to convert Float8_e4m3fn to the MPS backend [5].

  • To make a model smaller on a Mac, use a GGUF (Q4_K_M to Q8_0) through the ComfyUI-GGUF nodes.
  • Otherwise take the bf16 or fp16 file.
  • The int8 and nvfp4 files are tuned for NVIDIA cards; stick to GGUF and bf16 on a Mac.

Set up ComfyUI on a Mac.

  1. Install ComfyUI Desktop

    It needs Apple Silicon (M1 or later) and macOS 13 or newer [8]. For a manual install, ComfyUI’s README asks for the latest PyTorch nightly on Apple Silicon.

  2. Close what uses memory

    Browsers with many tabs, video apps and other AI tools share the same memory as the GPU. On 16 and 24 GB Macs, quitting them is the cheapest speed-up there is.

  3. Download GGUF or bf16 files

    Not fp8, for the reasons above. For Flux.1 Dev on 24 GB, a GGUF Q5_K_S (8.3 GB) or Q6_K (9.9 GB) is a good balance.

  4. Start small

    Generate at 1024 px, one image at a time. For video, start at 320 × 320 to check it runs at all, then go up [7].

How long it takes.

MacModel and settingsTime
Mac mini M4 Pro, 24 GBSDXL, 1024 px, 25 steps20 to 40 s[3]
M1 MaxSDXL, 1024 px, 20 stepsThe same image took 15.8 s on an RTX 4060 Laptop.about 40 s[2]
M1 Max, 64 GBFlux.2 Klein 4B, 1024 px, 4 stepsIn mflux (MLX), not ComfyUI.31.7 s[4]

There are no reliable ComfyUI numbers for Krea 2, Qwen-Image or video on current Macs yet.

When it goes wrong.

MPS backend out of memory (MPS allocated: … max allowed: …)
The model and the image didn’t fit the GPU’s share. A 16 GB Mac mini hit this with SD 3.5 Large [6]. Use a smaller file or size, and quit other apps. The message suggests PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.0; that removes the limit, and the Mac can freeze instead.
Trying to convert Float8_e4m3fn to the MPS backend
An fp8 file. Download the bf16 or a GGUF version instead.
A Flux image takes many minutes
Usually an fp8 file expanded to full size, or the Mac swapping to disk. Check Activity Monitor › Memory for swap. Use a GGUF that fits your memory.
Video comes out black, or stalls
Memory pressure. Lower the size and the frame count.

Where other apps are faster.

ComfyUI runs everything on this page, and it isn’t the fastest option on a Mac. Draw Things uses Metal-specific attention and was about 20% faster than ComfyUI on the same Mac mini [3]. mflux runs Flux models through Apple’s MLX. If you only need one model they support, they’re worth a look.

A Mac and a PC.

If there’s an NVIDIA PC in the house, run ComfyUI there and work from the Mac. Start ComfyUI on the PC with --listen, and point the Mac’s front end at the PC’s address.

Questions.

How much memory do I need on a Mac for AI images?

16 GB runs SDXL. 24 GB is comfortable for Flux.2 Klein, Z-Image and Flux.1 Dev as a GGUF. 48 GB runs Krea 2 at full precision, and 128 GB runs Flux.2 Dev.

Why is my Mac so much slower than a PC?

Apple’s GPU does less work per second than a desktop NVIDIA card, and ComfyUI is tuned for NVIDIA first. An SDXL image took about 40 seconds on an M1 Max and 16 on an RTX 4060 Laptop.

Should I use fp8 files on a Mac?

No. They load at full precision, so they save disk space but not memory, and some don’t load at all. Use a GGUF or the bf16 file.

What is iogpu.wired_limit_mb?

A macOS 14+ setting for how much memory the GPU may use. Raising it with sudo sysctl lets bigger models run; it resets at restart. Leave several GB for macOS.

M4 Pro with 24 GB or M4 Max with 36 GB?

For image models, the extra memory matters more than the chip: 36 GB runs Qwen-Image 2.1 and larger GGUFs of Flux.1 Dev comfortably. The Max also has a bigger GPU.

Can a Mac make video?

Yes, slowly. Wan 2.2 5B and HunyuanVideo 1.5 run on 32 GB and up, Wan 2.2 14B as GGUF halves on 64 GB. Start at a small size to check it runs.

Sources: [1] HEISS UI hardware notes, [2] SDXL on an RTX 4060 Laptop, lilting channel, [3] Local image generation on a Mac mini M4 Pro, heyuan110, [4] Flux.2 Klein 4B on an M1 Max, lilting channel, [5] fp8 files on Apple Silicon, ComfyUI discussion #13273, [6] MPS backend out of memory on a 16 GB Mac mini, ComfyUI issue #7171, [7] Wan 2.2 on Apple Silicon, Papaya Bytes, [8] ComfyUI Desktop for macOS, Comfy docs.

HEISS UI

Counted the Mac way.

HEISS UI is a prompt box and a gallery on top of your ComfyUI. It knows how much of your Mac’s memory the GPU can really use.

  • The model that fits your Mac is marked. It counts the GPU’s share of the memory, not the whole number, and never points a Mac at a file that saves no memory there.
  • Drop in any model and it runs. It recognises the file, even renamed, and uses settings that fit it. No graph to wire.
  • Missing parts, shown first. Anything a model still needs is listed with its size and a button. Nothing downloads behind your back.
  • Out of memory? One tap. A run that runs out offers to free memory and try again, or to try again smaller with the same seed.
  • Runs on the ComfyUI you have. No second install. Your models, your output folder.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI. Starter models download in one tap; other models and GGUF files you bring.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.