Pony V7 and AuraFlow.
AuraFlow is a 6.8B flow model from fal, released in 2024 under Apache 2.0. It reads prompts through a Pile T5-XL encoder instead of CLIP. PurpleSmartAI trained Pony V7 on top of it, from about 10 million images picked out of 30 million, split evenly between anime, cartoon, furry and pony styles.
Compared with Pony V6, which is an SDXL model:
- It understands sentences. V7 takes natural language as well as tags. V6 needed tag lists.
- Special tags are weaker.
score_andsource_tags still exist but matter less, and a newstyle_cluster_tag picks one of 2048 styles, which aren’t documented. - It’s heavier. About twice the parameters of SDXL and slower per step.
- It needs new LoRAs. V6 LoRAs don’t load on V7.
Reception is split. Civitai reviews are very positive; in the Hugging Face discussions, early users found attractive images harder to get than with V6. The card announces a V7.1 to fix weak special tags and faces; as of September 2026 it hasn’t appeared.
Files you need.
Three files, all from purplesmartai/pony-v7-base. The model file holds only the transformer, so the text encoder and VAE come separately.
-
Model13.7 GB Download
pony-v7-base.safetensorsComfyUI/models/checkpoints/ fp16, transformer only -
Text encoder3.0 GB Download
model.fp16.safetensorsComfyUI/models/text_encoders/ Pile T5-XL. Rename it, for example to pony-v7-t5xl.fp16.safetensors -
VAE0.2 GB Download
diffusion_pytorch_model.fp16.safetensorsComfyUI/models/vae/ the SDXL VAE; sdxl_vae.safetensors works too
Both extra files are shared with other models. The text encoder is byte for byte the one from AuraFlow v0.3, and the full-precision VAE in the repo is identical to Stability’s SDXL VAE. If you already have either, you don’t need to download it again.
models/
├── checkpoints/
│ └── pony-v7-base.safetensors
├── text_encoders/
│ └── pony-v7-t5xl.fp16.safetensors
└── vae/
└── pony-v7-vae.fp16.safetensors
Smaller files: GGUF
The repo has two official GGUFs in its gguf folder; a community set adds K-quants. They load through ComfyUI-GGUF and go in models/unet. Memory figures are the official ones.
| Quant | File | Memory |
|---|---|---|
| Q8_0 | 7.3 GB | about 10 GB |
| Q6_K | 5.7 GB | about 8 GB |
| Q5_K_M | 4.8 GB | about 7 GB |
| Q4_0 | 4.0 GB | about 6.5 GB |
| Q3_K_S | 3.0 GB | about 6 GB |
| Q2_K | 2.4 GB | about 5 GB |
Q8_0 and Q4_0 are official; the K-quants are from qpqpqpqpqpqp/pony_v7_base_GGUF. The makers recommend Q8_0.
A community fp8 version (6.9 GB) also exists. It needs a hybrid fp8 loader node; without it you get black images.
What fits your computer.
At 1024 × 1024 and above. The T5-XL encoder is small and runs before sampling, so the model file sets the limit.
- 6 GBTight
GGUF Q3 or Q4_0. Expect long waits.
- 8 GBFits
GGUF Q5_K_M (about 7 GB).
- 12 GBFits
GGUF Q8_0 (about 10 GB), the recommended one.
- 16 GBTight
The fp16 model, about 16 GB, with a little offloading.
- 24 GBFits
fp16 comfortably, at the full 1280 × 1536.
- MacUntested
No reports yet. By the numbers, GGUF Q8_0 should fit a 24 GB Mac. fp8 files don’t save memory on a Mac.
Set it up.
-
Load the official workflow
The repo’s workflows folder has PNG images with the workflow inside:
pony-v7-simple.png, a GGUF version and a LoRA version. Drag one onto ComfyUI to open it. -
Download the three files
Model, text encoder and VAE from the lists above, into their folders. Rename the encoder and VAE so you can tell them apart from other models’ files. Press R or restart ComfyUI.
-
Check the loaders
Load Checkpoint gets the model; only its MODEL output is used. Load CLIP gets the T5-XL; ComfyUI recognises it by its weights whatever type is set. Load VAE gets the VAE. For GGUF, use Unet Loader (GGUF) instead of Load Checkpoint.
-
Keep the T5 padding
The official workflow puts T5TokenizerOptions after Load CLIP with
min_padding768 andmin_length768. ComfyUI’s default for this encoder is 256, so leave the node in. -
Write the prompt in order
The card suggests: special tags, a factual description of the image, a stylistic description, then extra content tags. Name characters as species, gender, name and source, as in “pony female Twilight Sparkle from My Little Pony”.
Settings that work.
- Steps
- 20 to 30
- CFG
- 3.5
- Sampler
- euler
- Scheduler
- simple
- Size
- 1280 × 1536
- T5 padding
- 768 / 768
- Shift
- 1.73 (default)
- Negative
- yes
The official workflow uses 20 steps, CFG 3.48, euler and simple at 1280 × 1536. The card asks for “at least 30 steps” and sizes from 768 to 1536 pixels, and says to go higher rather than lower. 1536 pixels is the ceiling: the model has no position data beyond it. ComfyUI’s AuraFlow shift of 1.73 matches the model’s own scheduler, so there’s nothing to set.
In the LoRA workflow, the LoRA loads with LoraLoaderModelOnly at 0.5 strength. The repo’s custom PonyNoise node only switches between CPU and GPU noise to match diffusers; you don’t need it.
How fast.
Per step, AuraFlow and so Pony V7 is roughly six times slower than SDXL on the same card. One user found the community fp8 about 15% faster than fp16, and GGUF slower than both.
When it goes wrong.
- Load Checkpoint gives no CLIP or VAE
- Expected: the model file holds only the transformer. Add Load CLIP with the T5-XL and Load VAE.
- Black images with the fp8 file
- The community fp8 needs its hybrid fp8 loader node. Use it, or the fp16 file or a GGUF.
- Pony V6 LoRAs do nothing or error
- V6 is SDXL, V7 is AuraFlow. Only LoRAs trained on V7 work. The repo has a converter for SimpleTuner LoRAs.
- Faces and text look weak
- Known limits of this version, named on the card. Text rendering is worse than in plain AuraFlow.
- Positional embedding index out of bounds when training
- Training images above 1536 pixels. Keep the training resolution at or below 1536.
After a few days with V7, very few images come out pretty. What does style_cluster actually do?
Is there something special in the custom node, or can I use a normal workflow?
Questions.
Do Pony V6 LoRAs work on Pony V7?
No. Pony V6 is an SDXL model and V7 is built on AuraFlow, a different architecture. You need LoRAs trained on V7.
Which text encoder does Pony V7 use?
Pile T5-XL, the same file AuraFlow v0.3 uses: text_encoder/model.fp16.safetensors in the Pony V7 repo, 3.0 GB. It loads with a single Load CLIP node.
How much VRAM does Pony V7 need?
About 16 GB for the fp16 model. The official GGUF Q8_0 needs about 10 GB and Q4_0 about 6.5 GB, per the makers’ table.
What resolution should I use for Pony V7?
The official workflow uses 1280 × 1536. The model works from 768 to 1536 pixels and the card recommends going higher. 1536 is the maximum.
Is Pony V7 better than Pony V6?
It understands full sentences as well as tags, but many users find good-looking images harder to get, and V6 has far more LoRAs and merges. Try both on your prompts.
Can I use Pony V7 commercially?
The card’s summary of the Pony License allows commercial use of the model and outputs, except for inference services and apps, companies with over 1 million dollars in revenue, and professional video production.
Sources: Pony V7 base model card, official GGUF files, Pony V7 on Civitai, AuraFlow v0.3, feedback thread, fp8 thread, licence thread, ComfyUI issue #7324 [1], diffusers issue #12656.