By far, the biggest risk comes from orchestration. The Hugging Face incident showed that given a long enough leash and some open-ended tools (like web access), LLMs can construct their own state and combine multiple flaws and coordinated actions to achieve their result.
Orchestration means that LLMs can now pentest while dynamically cycling through every known and guessed vulnerability vector. The worst part is, this is emergent behavior so it can't easily be prevented at the model because each sub-agent could be operating safely while an attack is coordinated in an external process.