What Flux.2 Dev is.
Flux.2 Dev is the large open model in Black Forest Labs’ Flux.2 line, released on 25 November 2025. It’s a 32B image model paired with a whole language model as its text encoder: Mistral Small 3.2, 24B. The encoder alone is 18.0 GB in fp8, which is why this model needs so much memory.
- One model for making and editing. It draws from a prompt, and it edits from reference images: up to 10 of them, according to BFL and Comfy.
- Up to 4 megapixels, in any aspect ratio.
- No negative prompt. It’s guidance-distilled, so it runs at a fixed guidance value instead of a real CFG.
- A new architecture. Flux.1 LoRAs don’t load on it.
If your card has 12 GB or less, or you need a model you can use commercially, Flux.2 Klein is the better pick. Klein 4B is Apache 2.0 and fits 8 GB.
Files you need.
Get them from Comfy-Org/flux2-dev, not from BFL’s own repository. BFL’s repo is laid out for Python code, with transformer and text_encoder folders, and its model file is the 64.5 GB full-precision one.
-
Model35.5 GB Download
flux2_dev_fp8mixed.safetensorsComfyUI/models/diffusion_models/ full precision: flux2-dev.safetensors, 64.5 GB, gated at BFL -
Text encoder18.0 GB Download
mistral_3_small_flux2_fp8.safetensorsComfyUI/models/text_encoders/ or _bf16, 35.6 GB, or _fp4_mixed, 12.3 GB -
VAE0.3 GB Download
flux2-vae.safetensorsComfyUI/models/vae/ -
Turbo LoRA2.8 GB Download
Flux2TurboComfyv2.safetensorsComfyUI/models/loras/ optional: 8 steps instead of 28 to 50
Comfy’s newer templates use BFL’s small decoder as the VAE (full_encoder_small_decoder.safetensors, 0.25 GB). It decodes about 1.4 times faster with little visible difference. Either VAE works with every Flux.2 model.
models/
├── diffusion_models/
│ └── flux2_dev_fp8mixed.safetensors
├── text_encoders/
│ └── mistral_3_small_flux2_fp8.safetensors
├── loras/
│ └── Flux2TurboComfyv2.safetensors
└── vae/
└── flux2-vae.safetensors
Smaller files: GGUF
city96’s GGUF builds load through the ComfyUI-GGUF nodes and go in diffusion_models. For a GGUF text encoder, use a mainline llama.cpp conversion of Mistral Small 3.2 24B, such as unsloth’s. The gguf-org encoder files don’t load in ComfyUI-GGUF.
| Quant | Size | Quant | Size |
|---|---|---|---|
| Q8_0 | 35.0 GB | Q4_K_M | 20.1 GB |
| Q6_K | 27.4 GB | Q3_K_M | 16.0 GB |
| Q5_K_M | 24.1 GB | Q2_K | 12.9 GB |
Q4_K_M is the usual pick for a 16 GB card; Q4_K_S (19.3 GB) is a little smaller. The text encoder comes on top of these sizes.
RTX 50 cards: NVFP4
BFL publishes NVFP4 builds (flux2-dev-nvfp4, 21.0 GB, and flux2-dev-nvfp4-mixed, 22.8 GB). They’re fast on Blackwell cards, the RTX 50 series, and save a lot of memory. On older cards they bring no speed-up.
What fits your computer.
Flux.2 Dev in fp8 at 1024 × 1024, with the fp8 encoder. ComfyUI runs the text encoder first and then moves it aside, so model and encoder don’t need to sit on the card together, but they do both need room in graphics memory plus system RAM. That’s why system RAM matters here as much as the card.
- 6 GBNo
Use Flux.2 Klein instead.
- 8 GBSlow
With 32 GB of RAM it runs out of memory. With 64 GB, the Q4_K_M GGUF runs, and it’s rather slow.
- 12 GBNo
One user couldn’t run even the Q2 GGUF. Klein is the answer here.
- 16 GBTight
Q4_K_M GGUF used 14 to 15 GB with reference images and went over 16 GB without them.
--reserve-vram 2helped. fp8 runs with most of the model in system RAM: plan on 64 GB. - 24 GBOffloads
fp8 with part of the model in system RAM, 64 GB of it. A 3090 ran out of memory on the first step until
--reserve-vramwas set to 1 to 3. - 32 GBFits
RTX 5090: NVFP4 (21.0 GB) fits on the card, though some users hit out-of-memory with it where fp8 worked. fp8 needs a little system RAM. BFL names the 4090 and 5090 as the target cards for its fp8 and 4-bit options.
- Mac, fp8No
A Mac can’t compute in fp8, so ComfyUI loads the files at full precision: 64.5 GB of model and 35.6 GB of encoder.
- Mac, GGUFUntested
Q4_K_M (20.1 GB) with a GGUF Mistral encoder is the realistic route. We found no timings from Mac users.
Set it up.
-
Update ComfyUI
Flux.2 needs a ComfyUI from late November 2025 or newer, and current requirements. ComfyUI Desktop updates itself; the portable build has an update script. In a manual install:
Terminal, in the ComfyUI foldergit pull pip install -r requirements.txt
-
Download the files
Model, text encoder and VAE from the list above, about 54 GB together. Add the Turbo LoRA if you want 8-step images.
-
Put them in their folders
Model in
diffusion_models, text encoder intext_encoders, VAE invae, LoRA inloras. Then restart ComfyUI so the loaders list them. -
Open the Flux.2 Dev template
In ComfyUI’s template browser, pick Flux.2 Dev Text to Image. The Flux.2 Dev template adds reference images for editing. Each has an Enable Turbo LoRA switch, off by default.
-
Check the three loaders
Load Diffusion Model gets the Flux.2 Dev file. Load CLIP gets the Mistral encoder with type
flux2. Load VAE gets the Flux.2 VAE or the small decoder. The template may name the bf16 encoder; point it at the fp8 one if that’s what you downloaded. -
Run, and give it room
The first run loads everything and takes a while. If it stops on the first step with an out-of-memory error, start ComfyUI with a little memory held back:
Terminal, in the ComfyUI folderpython main.py --reserve-vram 2
Settings that work.
Flux.2 Dev samples through SamplerCustomAdvanced with Flux2Scheduler, which sets the noise schedule from the step count and the image size. There is no shift to set, and no CFG: guidance comes from a FluxGuidance node. The latent is EmptyFlux2LatentImage.
Base model
- Steps
- 28 to 50
- Guidance
- 4
- Sampler
- euler
- Scheduler
- Flux2Scheduler
- Size
- 1024 × 1024
- Negative
- none
- Guider
- BasicGuider
- Text encoder
- Load CLIP, flux2
BFL’s model card uses 50 steps and calls 28 a good trade-off. Comfy’s template uses 20, which is faster and a little rougher. Without a FluxGuidance node, ComfyUI falls back to guidance 3.5; BFL and the template both use 4. Any aspect ratio works up to 4 megapixels.
With the Turbo LoRA
- Steps
- 8
- Guidance
- 4
- LoRA
- Flux2TurboComfyv2
- Strength
- 1
The Turbo LoRA is fal’s, converted for ComfyUI. Comfy-Org hosts two versions of the same size; the v2 file has corrected LoRA keys, so take that one.
How fast.
| GPU | Files | Job | Time |
|---|---|---|---|
| RTX 3090 24 GB | fp8 | 2 images per batch | 9.5 s per step[1] |
| RTX 5070 Ti 16 GB | full precision, 96 GB RAM | 1280 × 720 | 2 to 4 min[2] |
User reports, not benchmarks. The 5070 Ti user found the fp8 file faster and much lighter on RAM than the full-precision one.
When it goes wrong.
torch.OutOfMemoryError: Allocation on deviceon the first step- ComfyUI under-estimated what the weights need. Start it with
--reserve-vram 2(1 to 3 works for most people), and update ComfyUI. mat1 and mat2 shapes cannot be multiplied (402x15360 and 4096x3072)- The text encoder and the model don’t belong together. Flux.2 Dev takes the Mistral encoder; Klein takes Qwen3.
LoraLoaderModelOnly: Required input is missing: model- The template shipped with the Turbo LoRA loader unconnected. Connect the model output of Load Diffusion Model to it.
'NoneType' object has no attribute 'Params'- An old
comfy-kitchenpackage. Update ComfyUI’s requirements (pip install -r requirements.txt). No module named 'transformers.models.pixtral.convert_pixtral_weights_to_hf'- The transformers package doesn’t match your ComfyUI. Update ComfyUI and its requirements; the report has no confirmed fix beyond that.
- No live preview while sampling, or a preview crash
- Update ComfyUI. If previews still crash, set the preview method to None in the settings.
- Bad hands or crowded scenes, and a negative prompt doesn’t help
- Flux.2 Dev has no negative prompt. Describe what you want instead, or try a different scheduler; several users find Flux.1 or Krea better at hands.
Which files do I download for ComfyUI? The BFL repo has transformer and text_encoder folders and I don’t know what goes where.
Why does bf16 run out of memory even on a 96 GB card? The text encoder is a whole language model, and that’s what fills it.
Questions.
How much VRAM does Flux.2 Dev need?
It runs on a 16 GB card as GGUF Q4_K_M, or in fp8 on 16 to 24 GB cards with 64 GB of system RAM. An RTX 5090 with 32 GB runs fp8 with 64 GB of RAM, or the smaller NVFP4 build. On 12 GB or less it doesn’t run well; use Flux.2 Klein there.
Which text encoder does Flux.2 Dev use?
Mistral Small 3.2, a 24B language model. The ComfyUI file is mistral_3_small_flux2_fp8.safetensors (18.0 GB) or the bf16 one (35.6 GB), loaded with a single Load CLIP node set to type flux2. It doesn’t use T5 or CLIP-L like Flux.1, or Qwen3 like Klein.
Can I use Flux.2 Dev commercially?
Not without a licence. It’s under the FLUX Non-Commercial License v2.1, and Black Forest Labs asks for a paid licence for client work and for products you charge for. Flux.2 Klein 4B is Apache 2.0 if you need a Flux.2 model for commercial work.
Does Flux.2 Dev have a negative prompt?
No. It’s guidance-distilled and runs with a fixed guidance of 4 through BasicGuider, so there is no negative prompt to fill in. The NAG node in ComfyUI doesn’t change that: users found identical output at twice the time.
Do Flux.1 LoRAs work on Flux.2 Dev?
No. Flux.2 is a new model trained from scratch, so Flux.1 LoRAs don’t load. You need LoRAs made for Flux.2 Dev.
How many reference images can Flux.2 Dev use?
Up to 10, according to Black Forest Labs and Comfy. Each one adds memory and time, so start with one or two on a 16 or 24 GB card.
Does Flux.2 Dev run on a Mac?
It can, with a lot of memory. A Mac can’t compute in fp8, so ComfyUI loads fp8 files at full precision: 64.5 GB for the model and 35.6 GB for the encoder. GGUF, Q4_K_M at 20.1 GB with a GGUF Mistral encoder, is the realistic route. We found no Mac timings.
Sources: FLUX.2 Dev model card, BFL flux2 repository, ComfyUI Flux.2 Dev tutorial, Comfy blog on Flux.2, ComfyUI issue #10891 [1], FLUX.2-dev memory thread [2], ComfyUI issue #11552, ComfyUI issue #12259, ComfyUI-GGUF issue #367.