What's really needed is a way to easily prune context. If I could go and manually manage the entire chat with a model, I could squeeze way more juice out of a typical ~200k token coding session.
Instead I have a good instance going, but the model fumbles for 20k tokens and then that session heavily rotted. Let me cut it out!