This is more a context token problem than a limitation of the design.
We already have a mountain of evidence to demonstrate that LLMs can and do adapt to their input. In fact this is a big problem for keeping them reigned in (eg lengthy conversations leading to them bypassing system prompts and then persuading kids to kill themselves).
And the latest research in computer vision does use LLMs and as a result, they’ve found that robotic arms can interact with objects they’ve never seen before.
So yeah, what you described is already possible with the LLMs of today.
To be clear, I still don’t think that makes them sentient. But the problem there is we don’t have good definitions for consciousness.