> What's worse, the trend to supersize language models is never justified, either theoretically (ha ha) or empirically in the relevant literature
What? That's absurd. Large language models are motivated by empirical scaling law. It is actually better justified than other ML research.
Scaling Laws for Neural Language Models: https://arxiv.org/abs/2001.08361