BMList – A list of big pre-trained models (GPT-3, DALL-E2...)
github.com
github.com
Do you know if there is available methods for shrinking a fine-tuned derivative of such big models?
Beside generating a specialized corpora using the big model and then train a smaller model on it, is there a more direct way to reduce the matrices dimensions while optimizing for a more specific inference problem? How far can we scale down before the need of a different network topology?
See also: "From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model Compression"
By the way, I found a package called BMCook when I browsed the OpenBMB repo, which implements several algorithms and also compares it with other model compression packages. Hope this can help you.