They should probably have been removed. This gives me the overall impression that the testers treat GPT-3 a bit too much as something like an artificial human, and not enough like an algorithm (which will work better with sanitized input). This is not a major criticism, the experiment is still interesting.
Could it be that the marketing from OpenAI it to blame? From the OpenAI front page:
> Discovering and enacting the path to safe artificial general intelligence.
> Our first-of-its-kind API can be applied to any language task, and currently serves millions of production requests each day.
Does that seem misleading?