Is there some work on how small a model with some specific epsilon perplexity could theoretically be? Given a fixed architecture and a fixed dataset, I presume there is a minimal number of parameters required for optimal representation.
And the whole result is also conditional on the optimization/training process used. Which is an area where we have no reason to think that we are optimal... So we can do studies with practical results (given sufficient money), but we are far from being able to identify the actual maximums available.