It is sensationalist and very easily shown to be not well thought out.
The present suite of models is not optimized to detect or deal with misdirection. It is a transparent presentation of weights learned through massive chunks of data and future iterations are likely going to bake those assumptions of misdirection into their training runs.
You highlighted potential solutions yourself. At its very core, the problem is one of either sanitizing inputs or outputs. How does one do this? Image models come pre packaged with a discerning model that flags NSFW pixels, there is no reason why a generic "Is this NSFW or threatening" text model optimized for this use case will not work to serve the same purpose. By ruling out the approach with an assumed, simplistic take that "Ai can be gamed, so NO", you're really splitting hairs and trying not to find a solution to provide a sensationalistic take.