That's apples to oranges.
No human ingests that many tokens of speech; individually we learn from far fewer tokens.
No human ingests that many tokens of speech; individually we learn from far fewer tokens.
"We only need x tokens of ingested speech" to learn language doesn't hit the same when you have billions of years of the brain baking in an oven to get to that point.
But i agree, it's not a direct comparison.
Why can't we start out a transformer at a better state and then teach it language in few tokens? Seems like a problem with architecture.
There's nothing wrong with blank slates. It's not a problem, whatever that means. It just is.