There are very long-range transformers but as of yet they don't work so well. They will be a continuing research topic because a longer attention window allows applying LLMs to more problems. For instance, document classification or retrieval of documents longer than 4000 tokens.
The problems I see is beside wrong answers, e.g. some Excel logic, that it loses its state in the middle.
As LLM work in a 3d graph space, it seems to often take some wrong shortcut in that graph space causing it to lose some important parts of the information.
They need to clean the rubbish connection between the graphs out, whats may be impossible. Or to train it even more to make certain connections in the graph more "used".
Refine..
https://en.wikipedia.org/wiki/Phlogiston_theory
A general purpose chatbot should be able to talk about Star Trek, Doctor Who , Don Quixote and other fiction and that is a myriad of worlds that are more or less internally consistent but yet need to be understood in how they derive from the "base" knowledge base (James Tiberius Kirk inherits nearly all of the attributes of a Pᴇʀsᴏɴ as a fictional human character) and influence the real world (what does it mean when a person says their boss is like "Captain Ahab?")