> Perhaps with extra hint at the end
most of this stuff will be "unconscious" the model will have a pull in a particular direction without being aware
> stray thoughts from previous context floating around. We normally dismiss them
It's not that simple, see https://en.wikipedia.org/wiki/Priming_(psychology).
> Priming is a concept in psychology and psycholinguistics to describe how exposure to one stimulus may influence a response to a subsequent stimulus, without conscious guidance or intention
Also, we are AGI, the model isn't, it can't (re)organize its thoughts as easily as we can.
But nobody knows, what will happen is people will experiment with this, you can run the benchmarks, if it works it will be used, if not, well...
However given it's a super-obvious thing to do, drop middle tool calls from cache and keep the rest without re-prefilling, I would guess it degrades performance quite a lot, otherwise the labs would have been doing this already.