I guess I'm skeptical that this actually improves performance. I'm worried that the middle man, the tool outputs, can strip useful context that the agent actually needs to diagnose.
Given our observations, the performance depends on the task and the model itself, most visible on long-running tasks
We state this directly appended into the outputs so the model knows exactly where the lines were removed from.