Image model

How to run Flux.2 Dev locally.

Flux.2 Dev runs in ComfyUI with three files: the 32B model, a 24B Mistral text encoder and the Flux.2 VAE. In fp8 they come to 53.8 GB, and they run on a 16 to 24 GB graphics card when the computer has 64 GB of system RAM. It’s non-commercial: client work needs a licence from Black Forest Labs.

Updated 29 Sep 202610 min read

Maker
Black Forest Labs
Released
Nov 202525 November 2025
Licence
FLUX Non-Commercialv2.1. Paid work needs a BFL licence
Memory
16 GB and upplus 64 GB of system RAM

What Flux.2 Dev is.

Flux.2 Dev is the large open model in Black Forest Labs’ Flux.2 line, released on 25 November 2025. It’s a 32B image model paired with a whole language model as its text encoder: Mistral Small 3.2, 24B. The encoder alone is 18.0 GB in fp8, which is why this model needs so much memory.

  • One model for making and editing. It draws from a prompt, and it edits from reference images: up to 10 of them, according to BFL and Comfy.
  • Up to 4 megapixels, in any aspect ratio.
  • No negative prompt. It’s guidance-distilled, so it runs at a fixed guidance value instead of a real CFG.
  • A new architecture. Flux.1 LoRAs don’t load on it.

If your card has 12 GB or less, or you need a model you can use commercially, Flux.2 Klein is the better pick. Klein 4B is Apache 2.0 and fits 8 GB.

Files you need.

Get them from Comfy-Org/flux2-dev, not from BFL’s own repository. BFL’s repo is laid out for Python code, with transformer and text_encoder folders, and its model file is the 64.5 GB full-precision one.

  • Model flux2_dev_fp8mixed.safetensors ComfyUI/models/diffusion_models/ full precision: flux2-dev.safetensors, 64.5 GB, gated at BFL
    35.5 GB Download
  • Text encoder mistral_3_small_flux2_fp8.safetensors ComfyUI/models/text_encoders/ or _bf16, 35.6 GB, or _fp4_mixed, 12.3 GB
    18.0 GB Download
  • VAE flux2-vae.safetensors ComfyUI/models/vae/
    0.3 GB Download
  • Turbo LoRA Flux2TurboComfyv2.safetensors ComfyUI/models/loras/ optional: 8 steps instead of 28 to 50
    2.8 GB Download

Comfy’s newer templates use BFL’s small decoder as the VAE (full_encoder_small_decoder.safetensors, 0.25 GB). It decodes about 1.4 times faster with little visible difference. Either VAE works with every Flux.2 model.

ComfyUI/models
models/
├── diffusion_models/
│   └── flux2_dev_fp8mixed.safetensors
├── text_encoders/
│   └── mistral_3_small_flux2_fp8.safetensors
├── loras/
│   └── Flux2TurboComfyv2.safetensors
└── vae/
    └── flux2-vae.safetensors

Smaller files: GGUF

city96’s GGUF builds load through the ComfyUI-GGUF nodes and go in diffusion_models. For a GGUF text encoder, use a mainline llama.cpp conversion of Mistral Small 3.2 24B, such as unsloth’s. The gguf-org encoder files don’t load in ComfyUI-GGUF.

QuantSizeQuantSize
Q8_035.0 GBQ4_K_M20.1 GB
Q6_K27.4 GBQ3_K_M16.0 GB
Q5_K_M24.1 GBQ2_K12.9 GB

Q4_K_M is the usual pick for a 16 GB card; Q4_K_S (19.3 GB) is a little smaller. The text encoder comes on top of these sizes.

RTX 50 cards: NVFP4

BFL publishes NVFP4 builds (flux2-dev-nvfp4, 21.0 GB, and flux2-dev-nvfp4-mixed, 22.8 GB). They’re fast on Blackwell cards, the RTX 50 series, and save a lot of memory. On older cards they bring no speed-up.

What fits your computer.

Flux.2 Dev in fp8 at 1024 × 1024, with the fp8 encoder. ComfyUI runs the text encoder first and then moves it aside, so model and encoder don’t need to sit on the card together, but they do both need room in graphics memory plus system RAM. That’s why system RAM matters here as much as the card.

  • 6 GBNo

    Use Flux.2 Klein instead.

  • 8 GBSlow

    With 32 GB of RAM it runs out of memory. With 64 GB, the Q4_K_M GGUF runs, and it’s rather slow.

  • 12 GBNo

    One user couldn’t run even the Q2 GGUF. Klein is the answer here.

  • 16 GBTight

    Q4_K_M GGUF used 14 to 15 GB with reference images and went over 16 GB without them. --reserve-vram 2 helped. fp8 runs with most of the model in system RAM: plan on 64 GB.

  • 24 GBOffloads

    fp8 with part of the model in system RAM, 64 GB of it. A 3090 ran out of memory on the first step until --reserve-vram was set to 1 to 3.

  • 32 GBFits

    RTX 5090: NVFP4 (21.0 GB) fits on the card, though some users hit out-of-memory with it where fp8 worked. fp8 needs a little system RAM. BFL names the 4090 and 5090 as the target cards for its fp8 and 4-bit options.

  • Mac, fp8No

    A Mac can’t compute in fp8, so ComfyUI loads the files at full precision: 64.5 GB of model and 35.6 GB of encoder.

  • Mac, GGUFUntested

    Q4_K_M (20.1 GB) with a GGUF Mistral encoder is the realistic route. We found no timings from Mac users.

Set it up.

  1. Update ComfyUI

    Flux.2 needs a ComfyUI from late November 2025 or newer, and current requirements. ComfyUI Desktop updates itself; the portable build has an update script. In a manual install:

    Terminal, in the ComfyUI folder
    git pull
    pip install -r requirements.txt
  2. Download the files

    Model, text encoder and VAE from the list above, about 54 GB together. Add the Turbo LoRA if you want 8-step images.

  3. Put them in their folders

    Model in diffusion_models, text encoder in text_encoders, VAE in vae, LoRA in loras. Then restart ComfyUI so the loaders list them.

  4. Open the Flux.2 Dev template

    In ComfyUI’s template browser, pick Flux.2 Dev Text to Image. The Flux.2 Dev template adds reference images for editing. Each has an Enable Turbo LoRA switch, off by default.

  5. Check the three loaders

    Load Diffusion Model gets the Flux.2 Dev file. Load CLIP gets the Mistral encoder with type flux2. Load VAE gets the Flux.2 VAE or the small decoder. The template may name the bf16 encoder; point it at the fp8 one if that’s what you downloaded.

  6. Run, and give it room

    The first run loads everything and takes a while. If it stops on the first step with an out-of-memory error, start ComfyUI with a little memory held back:

    Terminal, in the ComfyUI folder
    python main.py --reserve-vram 2

Settings that work.

Flux.2 Dev samples through SamplerCustomAdvanced with Flux2Scheduler, which sets the noise schedule from the step count and the image size. There is no shift to set, and no CFG: guidance comes from a FluxGuidance node. The latent is EmptyFlux2LatentImage.

Base model

Steps
28 to 50
Guidance
4
Sampler
euler
Scheduler
Flux2Scheduler
Size
1024 × 1024
Negative
none
Guider
BasicGuider
Text encoder
Load CLIP, flux2

BFL’s model card uses 50 steps and calls 28 a good trade-off. Comfy’s template uses 20, which is faster and a little rougher. Without a FluxGuidance node, ComfyUI falls back to guidance 3.5; BFL and the template both use 4. Any aspect ratio works up to 4 megapixels.

With the Turbo LoRA

Steps
8
Guidance
4
LoRA
Flux2TurboComfyv2
Strength
1

The Turbo LoRA is fal’s, converted for ComfyUI. Comfy-Org hosts two versions of the same size; the v2 file has corrected LoRA keys, so take that one.

How fast.

GPUFilesJobTime
RTX 3090 24 GBfp82 images per batch9.5 s per step[1]
RTX 5070 Ti 16 GBfull precision, 96 GB RAM1280 × 7202 to 4 min[2]

User reports, not benchmarks. The 5070 Ti user found the fp8 file faster and much lighter on RAM than the full-precision one.

When it goes wrong.

torch.OutOfMemoryError: Allocation on device on the first step
ComfyUI under-estimated what the weights need. Start it with --reserve-vram 2 (1 to 3 works for most people), and update ComfyUI.
mat1 and mat2 shapes cannot be multiplied (402x15360 and 4096x3072)
The text encoder and the model don’t belong together. Flux.2 Dev takes the Mistral encoder; Klein takes Qwen3.
LoraLoaderModelOnly: Required input is missing: model
The template shipped with the Turbo LoRA loader unconnected. Connect the model output of Load Diffusion Model to it.
'NoneType' object has no attribute 'Params'
An old comfy-kitchen package. Update ComfyUI’s requirements (pip install -r requirements.txt).
No module named 'transformers.models.pixtral.convert_pixtral_weights_to_hf'
The transformers package doesn’t match your ComfyUI. Update ComfyUI and its requirements; the report has no confirmed fix beyond that.
No live preview while sampling, or a preview crash
Update ComfyUI. If previews still crash, set the preview method to None in the settings.
Bad hands or crowded scenes, and a negative prompt doesn’t help
Flux.2 Dev has no negative prompt. Describe what you want instead, or try a different scheduler; several users find Flux.1 or Krea better at hands.

Which files do I download for ComfyUI? The BFL repo has transformer and text_encoder folders and I don’t know what goes where.

Hugging Face, FLUX.2-dev

Why does bf16 run out of memory even on a 96 GB card? The text encoder is a whole language model, and that’s what fills it.

Hugging Face, FLUX.2-dev

Questions.

How much VRAM does Flux.2 Dev need?

It runs on a 16 GB card as GGUF Q4_K_M, or in fp8 on 16 to 24 GB cards with 64 GB of system RAM. An RTX 5090 with 32 GB runs fp8 with 64 GB of RAM, or the smaller NVFP4 build. On 12 GB or less it doesn’t run well; use Flux.2 Klein there.

Which text encoder does Flux.2 Dev use?

Mistral Small 3.2, a 24B language model. The ComfyUI file is mistral_3_small_flux2_fp8.safetensors (18.0 GB) or the bf16 one (35.6 GB), loaded with a single Load CLIP node set to type flux2. It doesn’t use T5 or CLIP-L like Flux.1, or Qwen3 like Klein.

Can I use Flux.2 Dev commercially?

Not without a licence. It’s under the FLUX Non-Commercial License v2.1, and Black Forest Labs asks for a paid licence for client work and for products you charge for. Flux.2 Klein 4B is Apache 2.0 if you need a Flux.2 model for commercial work.

Does Flux.2 Dev have a negative prompt?

No. It’s guidance-distilled and runs with a fixed guidance of 4 through BasicGuider, so there is no negative prompt to fill in. The NAG node in ComfyUI doesn’t change that: users found identical output at twice the time.

Do Flux.1 LoRAs work on Flux.2 Dev?

No. Flux.2 is a new model trained from scratch, so Flux.1 LoRAs don’t load. You need LoRAs made for Flux.2 Dev.

How many reference images can Flux.2 Dev use?

Up to 10, according to Black Forest Labs and Comfy. Each one adds memory and time, so start with one or two on a 16 or 24 GB card.

Does Flux.2 Dev run on a Mac?

It can, with a lot of memory. A Mac can’t compute in fp8, so ComfyUI loads fp8 files at full precision: 64.5 GB for the model and 35.6 GB for the encoder. GGUF, Q4_K_M at 20.1 GB with a GGUF Mistral encoder, is the realistic route. We found no Mac timings.

Sources: FLUX.2 Dev model card, BFL flux2 repository, ComfyUI Flux.2 Dev tutorial, Comfy blog on Flux.2, ComfyUI issue #10891 [1], FLUX.2-dev memory thread [2], ComfyUI issue #11552, ComfyUI issue #12259, ComfyUI-GGUF issue #367.

HEISS UI

The big one, made easy.

HEISS UI runs Flux.2 Dev on the ComfyUI you already have. It sets up the model, its large text encoder and the VAE, and shows what’s missing before anything downloads.

  • Big downloads that don’t start over. The parts add up to more than 50 GB. Free space is checked first, and downloads resume and are verified.
  • A first model that fits. On an empty studio it’s one tap away, with the version for your GPU or Mac marked.
  • Edit from your own pictures. Add reference images in the composer and describe the change.
  • Failures that explain themselves. When a run runs out of memory, it says so and offers the fix.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.