This just means the user generated content gets sent to the API once with different framing (risking a ban or strike or whatever) and if it doesn't trigger your detection gets sent again with the normal framing (giving another chance at a provider ban, strike, etc.)
Seems like that would just accelerate your ban by having you send each potentially-violating interaction twice, with slightly different context, giving more chances of a violation and possibly doubling violations for some content.
You can probably do better at reducing your risk by running a local classifier (or a comparatively small local LLM) as your trouble detector, before deciding to send a request to the backend, though validating the trouble detector setup may be problematic.