I think the way to swallow this bitter pill is to acknowledge they can "generalize" because all human knowledge is actually a relatively "small" finite distribution that models are now big enough to pattern match on.
There's simply no way an LLM can even train on all of that because each bit of true expert knowledge necessarily comically underrepresented in any possible training set.
Though I'm not even sure about "/s", it is more than feasible to build such a bot that would gather quality information sources.
I mean, if it can reason about and process the data as it ingests it?