Image model

How to run Pony V7 locally.

Pony V7 runs in ComfyUI with three files: the 13.7 GB model, a 3.0 GB Pile T5-XL text encoder and a 0.2 GB VAE. At full precision it wants about 16 GB of video memory; the official GGUF Q8_0 needs about 10 GB. It’s built on AuraFlow, not SDXL, so Pony V6 LoRAs don’t carry over.

Updated 29 Sep 20268 min read

Maker
PurpleSmartAI
Released
Oct 2025Built on fal’s AuraFlow v0.3
Licence
Pony LicenseNo inference services or firms over $1M revenue
Memory
16 GBGGUF Q8_0 about 10 GB, Q4 about 6.5 GB

Pony V7 and AuraFlow.

AuraFlow is a 6.8B flow model from fal, released in 2024 under Apache 2.0. It reads prompts through a Pile T5-XL encoder instead of CLIP. PurpleSmartAI trained Pony V7 on top of it, from about 10 million images picked out of 30 million, split evenly between anime, cartoon, furry and pony styles.

Compared with Pony V6, which is an SDXL model:

  • It understands sentences. V7 takes natural language as well as tags. V6 needed tag lists.
  • Special tags are weaker. score_ and source_ tags still exist but matter less, and a new style_cluster_ tag picks one of 2048 styles, which aren’t documented.
  • It’s heavier. About twice the parameters of SDXL and slower per step.
  • It needs new LoRAs. V6 LoRAs don’t load on V7.

Reception is split. Civitai reviews are very positive; in the Hugging Face discussions, early users found attractive images harder to get than with V6. The card announces a V7.1 to fix weak special tags and faces; as of September 2026 it hasn’t appeared.

Files you need.

Three files, all from purplesmartai/pony-v7-base. The model file holds only the transformer, so the text encoder and VAE come separately.

  • Model pony-v7-base.safetensors ComfyUI/models/checkpoints/ fp16, transformer only
    13.7 GB Download
  • Text encoder model.fp16.safetensors ComfyUI/models/text_encoders/ Pile T5-XL. Rename it, for example to pony-v7-t5xl.fp16.safetensors
    3.0 GB Download
  • VAE diffusion_pytorch_model.fp16.safetensors ComfyUI/models/vae/ the SDXL VAE; sdxl_vae.safetensors works too
    0.2 GB Download

Both extra files are shared with other models. The text encoder is byte for byte the one from AuraFlow v0.3, and the full-precision VAE in the repo is identical to Stability’s SDXL VAE. If you already have either, you don’t need to download it again.

ComfyUI/models
models/
├── checkpoints/
│   └── pony-v7-base.safetensors
├── text_encoders/
│   └── pony-v7-t5xl.fp16.safetensors
└── vae/
    └── pony-v7-vae.fp16.safetensors

Smaller files: GGUF

The repo has two official GGUFs in its gguf folder; a community set adds K-quants. They load through ComfyUI-GGUF and go in models/unet. Memory figures are the official ones.

QuantFileMemory
Q8_07.3 GBabout 10 GB
Q6_K5.7 GBabout 8 GB
Q5_K_M4.8 GBabout 7 GB
Q4_04.0 GBabout 6.5 GB
Q3_K_S3.0 GBabout 6 GB
Q2_K2.4 GBabout 5 GB

Q8_0 and Q4_0 are official; the K-quants are from qpqpqpqpqpqp/pony_v7_base_GGUF. The makers recommend Q8_0.

A community fp8 version (6.9 GB) also exists. It needs a hybrid fp8 loader node; without it you get black images.

What fits your computer.

At 1024 × 1024 and above. The T5-XL encoder is small and runs before sampling, so the model file sets the limit.

  • 6 GBTight

    GGUF Q3 or Q4_0. Expect long waits.

  • 8 GBFits

    GGUF Q5_K_M (about 7 GB).

  • 12 GBFits

    GGUF Q8_0 (about 10 GB), the recommended one.

  • 16 GBTight

    The fp16 model, about 16 GB, with a little offloading.

  • 24 GBFits

    fp16 comfortably, at the full 1280 × 1536.

  • MacUntested

    No reports yet. By the numbers, GGUF Q8_0 should fit a 24 GB Mac. fp8 files don’t save memory on a Mac.

Set it up.

  1. Load the official workflow

    The repo’s workflows folder has PNG images with the workflow inside: pony-v7-simple.png, a GGUF version and a LoRA version. Drag one onto ComfyUI to open it.

  2. Download the three files

    Model, text encoder and VAE from the lists above, into their folders. Rename the encoder and VAE so you can tell them apart from other models’ files. Press R or restart ComfyUI.

  3. Check the loaders

    Load Checkpoint gets the model; only its MODEL output is used. Load CLIP gets the T5-XL; ComfyUI recognises it by its weights whatever type is set. Load VAE gets the VAE. For GGUF, use Unet Loader (GGUF) instead of Load Checkpoint.

  4. Keep the T5 padding

    The official workflow puts T5TokenizerOptions after Load CLIP with min_padding 768 and min_length 768. ComfyUI’s default for this encoder is 256, so leave the node in.

  5. Write the prompt in order

    The card suggests: special tags, a factual description of the image, a stylistic description, then extra content tags. Name characters as species, gender, name and source, as in “pony female Twilight Sparkle from My Little Pony”.

Settings that work.

Steps
20 to 30
CFG
3.5
Sampler
euler
Scheduler
simple
Size
1280 × 1536
T5 padding
768 / 768
Shift
1.73 (default)
Negative
yes

The official workflow uses 20 steps, CFG 3.48, euler and simple at 1280 × 1536. The card asks for “at least 30 steps” and sizes from 768 to 1536 pixels, and says to go higher rather than lower. 1536 pixels is the ceiling: the model has no position data beyond it. ComfyUI’s AuraFlow shift of 1.73 matches the model’s own scheduler, so there’s nothing to set.

In the LoRA workflow, the LoRA loads with LoraLoaderModelOnly at 0.5 strength. The repo’s custom PonyNoise node only switches between CPU and GPU noise to match diffusers; you don’t need it.

How fast.

GPUModelSpeed
RX 7900 XTXAuraFlow1.6 s/it[1]
RX 7900 XTXSDXL, same card4 it/s[1]

Per step, AuraFlow and so Pony V7 is roughly six times slower than SDXL on the same card. One user found the community fp8 about 15% faster than fp16, and GGUF slower than both.

When it goes wrong.

Load Checkpoint gives no CLIP or VAE
Expected: the model file holds only the transformer. Add Load CLIP with the T5-XL and Load VAE.
Black images with the fp8 file
The community fp8 needs its hybrid fp8 loader node. Use it, or the fp16 file or a GGUF.
Pony V6 LoRAs do nothing or error
V6 is SDXL, V7 is AuraFlow. Only LoRAs trained on V7 work. The repo has a converter for SimpleTuner LoRAs.
Faces and text look weak
Known limits of this version, named on the card. Text rendering is worse than in plain AuraFlow.
Positional embedding index out of bounds when training
Training images above 1536 pixels. Keep the training resolution at or below 1536.

After a few days with V7, very few images come out pretty. What does style_cluster actually do?

Hugging Face, Pony V7 feedback thread

Is there something special in the custom node, or can I use a normal workflow?

Hugging Face, Pony V7

Questions.

Do Pony V6 LoRAs work on Pony V7?

No. Pony V6 is an SDXL model and V7 is built on AuraFlow, a different architecture. You need LoRAs trained on V7.

Which text encoder does Pony V7 use?

Pile T5-XL, the same file AuraFlow v0.3 uses: text_encoder/model.fp16.safetensors in the Pony V7 repo, 3.0 GB. It loads with a single Load CLIP node.

How much VRAM does Pony V7 need?

About 16 GB for the fp16 model. The official GGUF Q8_0 needs about 10 GB and Q4_0 about 6.5 GB, per the makers’ table.

What resolution should I use for Pony V7?

The official workflow uses 1280 × 1536. The model works from 768 to 1536 pixels and the card recommends going higher. 1536 is the maximum.

Is Pony V7 better than Pony V6?

It understands full sentences as well as tags, but many users find good-looking images harder to get, and V6 has far more LoRAs and merges. Try both on your prompts.

Can I use Pony V7 commercially?

The card’s summary of the Pony License allows commercial use of the model and outputs, except for inference services and apps, companies with over 1 million dollars in revenue, and professional video production.

Sources: Pony V7 base model card, official GGUF files, Pony V7 on Civitai, AuraFlow v0.3, feedback thread, fp8 thread, licence thread, ComfyUI issue #7324 [1], diffusers issue #12656.

HEISS UI

The two parts, fetched.

HEISS UI runs Pony V7 on the ComfyUI you already have. Pick the file and it offers the text encoder and VAE it can’t run without.

  • The parts it can’t run without. Both are listed with their sizes and fetched in one go. An SDXL VAE you already have counts too.
  • Drop in the file and it runs. It knows the model from the file itself, even renamed, and picks settings that work.
  • Your LoRAs, sorted. The ones made for this model come first, with their trigger words.
  • One gallery for everything. Search by prompt, model or LoRA, star the keepers, and bring old ComfyUI, AUTOMATIC1111 or Forge folders along.

Free and open source. macOS, Windows and Linux. Runs on your ComfyUI.

HEISS UI with a gallery of generated images and the prompt composer at the bottom.