> language models aren’t large enough to actually know everything
I'd say they don't know anything.
An LLM base model, before it is post-trained with RL, just has access to a sliced and diced corpus of human output. Take the contents of 4chan and WikiPedia, put in blender and mix and chop into "training sample" sized bites, then learn the statistical regularities of this blended mess. It is what it is - not exactly what I'd call a knowledge base, even though there are bits of knowledge in there.
When you add RL-based post-training for reasoning, all you are doing is trying get the model to be more selective when you are sampling from it - encouraging it to suppress some statistics, and emphazise others, such that when you sample from it the output looks more like valid reasoning steps and/or conclusions, per the verified reasoning examples you train it on.
I'm well aware of how useful RL-tuned models (whatever the goal) can be, but at the end of the day all they are doing is taking a statistical babbler and saying "try to output patterns more like this". It's not exactly a recipe for factuality or rationality - we've just gone from hallucination-prone base models, to gaslighting-prone RL-tuned "reasoning" models that output stuff that sounds like coherent reasoning.
What missing from all of this - what makes it different from how animals learn - it that the model has no experience of it's own, no autonomy or motivation to explore, learn and verify, and hence no episodic memories of how it learnt something (tried it and ran controlled experiments, or just overheard it on the bus), and what that implies about it's trustworthiness.
It's amazing that LLMs work as well as they do, a reflection of how much of what we do can be accomplished by reactive pattern matching, but if you want to go beyond that to something that can learn and discern the truth for itself, this seems the wrong paradigm altogether.