There's a post training.step where they get humans to interact with it.
As I understand it, there is (or was) a step where they ask people what's the 'better' response.
These linguistic forms sound good the first time you hear themz even if they are rare in real speech, so rapidly got trained in.
Now they distill off previous models, I imagine these weird linguistic forms are quite hard to get rid of.