What is the current hypothesis on if the context windows would be substantially larger, what would this enable LLMs to do that is beyond capabilities of current models (other than the obvious the
now getting forgetful/confused when you’ve exhausted the context)?