This is a modular model - download components below
Type
CLIP
Stats
4110 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
Reviews
(25)
Published
Oct 2, 2025
Base Model
Training
Steps: 5,000,000
Epochs: 20
Hash
AutoV2
3987E95EF4


70K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9K
652K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9K
5.1M0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9M

License:
Illustrious LicenseBalanced CLIP (1M)
Training CLIP-G took >15KwH of energy, CLIP-L took far less <1KwH
The full negative reinforcement (Cosine Dissimilarity) is available on my huggingface, this was paired with a positive reinforcement (Contrastive Loss) using the full frozen vision model in latent space.
PONY CLIP-L has a further 10 epochs using ASGD for very fine-tuned loss.

