I'm flabbergasted. I translated the tweets to Hebrew and reran the example - it returned the correct results. I then changed the input to a negative, and it again returned the correct results. So it's not only in English, and I'm sure that the Hebrew dataset was much smaller. Perhaps it is translating behind the scenes.
Thanks! Perhaps I'm not good at prompt engineering, but I could barely get anything useful out of it.
It's mainly for text classification, which explains why it's not really giving comparable outputs to GPT3