This is wrong.
I saw so many people who think this though, even smart people. But it's just clearly not true if you think about it. It's like saying that you can't train a model to predict a trend from a scatter plot because the model can't be smarter than the average point in the scatter plot (or even the smartest one), and points in a scatter plot aren't smart at all.
I think the gap in understanding is that people aren't used to 'models' being treated themselves as 'data points'. So when they imagine a model being trained over model-like data points, they start getting confused between what is a model and what is a data point, and they start thinking that the model being trained can't be more capable than the smartest (or some even say average lol) data point (which is itself a model) in its training set.
Another reason why this kind of thinking is unintuitive is because of the raw scale of these LLMs. The good ones like the first ones that people are saying might become super-intelligent are going to be entire data centers, or data center sized supercomputers like Aurora, and they will cost billions of dollars in training. During training they will have more than a trillion parameters and trained on more than tens of trillions of tokens, so more than 10,000,000,000,000,000,000,000,000 numerical updates during their training. That number is just very large and outside the realm of human intuition for things like running through a for-loop in your mind when you are imagining the algorithm.