I've been doing a lot of testing around prompt injection and tool misuse in agent workflows recently.
One pattern that keeps showing up is that once agents start reading external content and calling tools, the security model changes quite a bit.
Traditional jailbreaks focus on overriding the system prompt. But in agent systems the bigger issue seems to be boundary confusion: untrusted content ends up influencing how tools are used.
Curious how others building agent systems are thinking about this.