This is true for classical ML, but unfortunately it does not hold very well for hard problems. The reason is that GPT-3 at al. are meant as multipurpose pretrained models their size is their strength in that they can model highly complex datasets.
In my opinion the right approach to downsizing is to train at full size and then either distill or prune the model to keep the parts that are relevant to your problem.