> just run the output through a completely separate model that checks whether the string that's about to be returned violates the prohibitions
The big names do do this. Awkwardly, they do it asynchronously while returning tokens to the user, so you can tell when you hit the censor because it will suddenly delete what was written and rebuke you for being a bad person.