Short answer: Yes, they tested the model with racist speech, and it failed.
> I entered a selection of racist tweets I’d received in the past year into Perspective’s API, of varying lengths and sentence complexity. Just to make it easy for the machine, I deliberately chose one which included a common racial slur against South Asians.
> None of these were registered as potentially toxic at all by the AI – but, “You’re a fucking G”, a compliment, popped up with a 90.29% likelihood of being toxic.
Your hypotheses that white MPs experience more racist speech than non-white MPs, is testable. I have a pretty strong feeling that you are wrong, but I haven’t seen any data to support either (only anecdotes; and this deeply flawed analysis). What you can do to test your hypotheses is find all British MPs on twitter, scrape all tweets directed at them, and measure the proportion of racist speech against them. However, I advice you to do a more traditional cluster analysis, rather than a training model, the latter is very likely to yield a flawed model just like the journalist’s “8 month labor of love”.