We fixed jailbreaks. We didn't fix agents
twitter.com
twitter.com
One pattern that keeps showing up is that once agents start reading external content and calling tools, the security model changes quite a bit.
Traditional jailbreaks focus on overriding the system prompt. But in agent systems the bigger issue seems to be boundary confusion: untrusted content ends up influencing how tools are used.
Curious how others building agent systems are thinking about this.