They don't specify GPT-4. The fact they generically refer to ChatGPT leads me to believe they're assessing 3.5. GPT-4 is significantly more competent of a model and while I won't speculate on it's math ability I'd want to see a follow up, specifically for GPT-4.