Optimazing for training could help distillation also.
Don't know if there any public technical reports by any of the big AI companies about this, as its pretty new.
Though this theory has the defect that GPT-4 is, I think, more expensive than GPT-3, but as I recall it was considered unlikely that GPT-4 is larger than 175 billion parameters. Not sure.