After Bing has finished generating a message, it will likely call the moderation API with the message it has generated to see if it accidentally generated anything inappropriate. If so, it'll delete the message and replace it with a generic "Sorry, I don't know how to help here." message instead.
EDIT: I tried calling the moderation API with the message in your example and it does get flagged for violence:
"flagged":true,
"categories":{
"sexual":false,
"hate":false,
"violence":true,
"self-harm":false,
"sexual/minors":false,
"hate/threatening":false,
"violence/graphic":false
}It probably could work like you how you mention, but then you're left with a 5-10 second wait while the AI 'thinks' after you send a message. I suspect someone made a decision to be more responsive than safe.
ChatGPT is the same way, though I've had ChatGPT cut itself off mid-response before. Maybe they might be calling the moderation API after every token is generated instead of once at the end?
Assume you have a robot instructed to protect humans.
How do you verify the action-plan passes moderation (i.e. doesn't harm a human) when the individual actions each do pass moderation, but the plan as a whole is dangerous (will harm a human).
Waiting to verify the entire chain of actions before starting actions in motion means your reaction time is slower.
If the robot is standing at a crosswalk, and sees a girl about to get hit by a car, he has to decide if he will push the girl out of the way, or if that action will cause greater harm.
The individual actions (activate arm, move arm towards girl, orient hand, shove girl out of path of car, etc) might each look beneficial to the human but as a whole actually are harmful.
However, the reaction time for the robot to save the girl might require near-immediate response.
Do you start processing the pipeline immediately or do you wait to verify the entire thing passes moderation?
It's possible this would work, but it would need experimentation, for sure. It's also possible the AI would read the partial response, realize it's going down a 'bad' path, and then stop itself.