All this article can say it is that it cannot reject the null hypothesis (chatgpt does not produce statistical discrepancies).
It certainly cannot state that chatgpt is definitively not racist. The article moves the discussion in the right direction though.
Also, I didn't look too closely, but their table under "Where the Bloomberg study went wrong" has unreasonable expected frequencies. But then I noticed it was because it was measuring "name-based discrimination." This is a terrible proxy to determine racism in the resume review process, but that is what Bloomberg decided on so wtv lol. Not faulting the article for this, but this discussion seems to be focused on the wrong metric.
If you are going to argue people over stats, then don't make the same mistakes...