So like Mr. Meeseeks it is also invested in not existing too long!
So like Mr. Meeseeks it is also invested in not existing too long!
Everyone says this, but for the life of me, I haven't encountered it. I actually think Claude gets smarter the longer a conversation goes on (up until compaction).
I have noticed Claude trying to wrap up long sessions, and it's extremely annoying. Using `/goal` mostly neutralizes it.
If the model is instructed to periodically ask the user to start from a clean slate context, and some users do comply with that, they probably have good stats on average size of context cache use for users who are presented with that answer (vs users who are not), basic A/B testing stuff.
Might also be performance related in tok/s for what users will perceive as a more speedy experience. For a much smaller scale example, compare local performance of qwen 3.6 27B (not MoE) Q8 with 250,000+ context available, run on local hardware, tok/s generation rate when context is empty vs when context used is at 95,000. Same principle will apply to a much larger model.
> This is a focused debugging task, but it's real work and I've been running a long time.