This is a modular model - download components below


70K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9K
652.1K0 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K
5.1M0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9M

License:
CreativeML Open RAIL++-MBalanced CLIP (1M)
Training CLIP-G took >15KwH of energy, CLIP-L took far less <1KwH
The full negative reinforcement (Cosine Dissimilarity) is available on my huggingface, this was paired with a positive reinforcement (Contrastive Loss) using the full frozen vision model in latent space.
PONY CLIP-L has a further 10 epochs using ASGD for very fine-tuned loss.

