Dumb question: can you train a model to predict the next byte of ANOTHER MODEL
So apply this same logic to compressing a bigger model within a smaller model
I know this is absolutely regarded, but humour me please
So apply this same logic to compressing a bigger model within a smaller model
I know this is absolutely regarded, but humour me please
Compression is such an interesting field