If I give an LLM to compact its context window, so the context it carries can evolve iteratively over time as more and more things come in, is that enough?
Compacting the context is really a very, very interesting example here. The "next token predictor" is telling an external tool to change all "previous" tokens. So an LLM + a harness that allows compacting the context is no longer just a token predictor at all!
You don't need continuous learning to get interesting dynamics. You just need feedback loops.
Richard Dawkins says he thinks LLMs think.
And the physical angle is that nothing is special about humans and software simulating it would also be thinking.
But from using LLMs all the time, and understanding what is under the hood, I’m thinking that the current approaches aren’t cutting it for me and I’m not expecting them to get there. There was such big jumps early on but progress is slowing as though diminishing returns.
So we can build things that think and outthink us, but I don’t think anything we’ve hit upon yet is going to scale up into it.