Major companies are using these LLMs on their already hardened software and finding countless thousands of security flaws. Is that all theater too? If not then it's very easy to see how these things are dangerous.
Humans have been doing it for decades.
People should be running LLMs on their own systems to test them.
The flaws are there whether they are seen by ai or not.