Not all of them per se, take a look at something like Mistral. It's a 7B model displaying incredible performance. IMO, we still haven't even scratched the surface of what is possible with small LLMs. Especially not with pre-filtered/classified pre-training data. (Interesting LLMs based on their data approach and relatively small size: Qwen, InternLM, Mistral, Phi)