We already know Large Language Models (LLMs) can learn at runtime (ie, separately to the training process.) This is called "In Context Learning". See [1], [2] for more details. (BTW, when anyone says "LLMs are stochastic parrots" you know they are ignorant of this)
In context learning is wonderful because it means you can "train" a LLM at run time by filling the context with examples. Traditionally this "context window" has been a few thousand tokens, and GPT-4 recently extended that to 32,000 tokens.
That is useful, but if you wanted to say load all of a companies documents and ask questions it doesn't really work because this overflows the context.
But at 2M tokens there's a whole range of applications that become possible.
[1] Language Models are Few-Shot Learners: https://arxiv.org/abs/2005.14165
[2] Language Models Secretly Perform Gradient Descent as Meta-Optimizers: https://arxiv.org/abs/2212.10559