I've always thought it was abundantly clear how to make smaller models perform as well as large models: keep labeling data and build a human-in-the-loop support process to keep it on track.
My perspective is more pessimistic. I think people opt for huge unsupervised models because they believe that tuning a few thousand more input features is easier than labeling copious amounts of data. Plus (in my experience) supervised models often require a more involved understanding of the math, whereas there's so many NN frameworks that ask very little of the users.