But all these smaller models are based on those 100b+ models! Like OP on TinyStories relies on larger models twice, to generate and then evaluate. Or the recent wave of small models which are all based on extracting data from GPT-3/4 to borrow the capbilities+RLHFing for free.
You can talk about how interesting it is and how much of an overhang neural nets have in terms of being overparameterized (which is an important AI safety/capabilities issue - the first AGI will be the largest, slowest, and worst one, by a long shot), but the one thing it doesn't tell you is that smaller models can replace big models entirely. Because they are parasitic on said big models, so you still need to train the big models to begin with.
(The smaller models also seem to lose a lot of the qualitative capabilities of larger ones, like meta-learning. Their TinyStories model does only stories. I doubt you could get it to 'dax a blick'.)