New AI tasks are being unlocked by (large-scale) foundation models (Liang, 2022).
Fine-tuning in low-resource (few-shot) scenarios is now possible for many new applications.
However, these new AI applications relied upon a huge pretrained model to get there. Because the old approach of training from scratch on 100 labeled examples didn't work well.
Thus, we want to distill the knowledge so that the model can be deployed in low-resource scenarios.
[edit: I see your below comment about the concern about transformer cost. Agreed. This is one of the many concerns around foundation models that must be understood. The happy path is that training the foundation model is a one-time cost that pays dividends in the many tasks it unlocks. However, you are correct that the research to get there is quite spendy. I encourage you to skim this paper. It's long but very accessible: https://arxiv.org/pdf/2108.07258.pdf
ps the reason commercial use cases all use simple models is not because simple models are ipso facto commercially valuable. It's just that industry practitioners are too overworked to do fancy bleeding edge stuff. Thus, the dirty secrets is that most fancy ML companies are just using logistic regression for everything. Foundation models allow industry practitioners rapidly to train powerful accurate models. The question is, now, how do they deploy them.]