They don't specify GPT-4. The fact they generically refer to ChatGPT leads me to believe they're assessing 3.5. GPT-4 is significantly more competent of a model and while I won't speculate on it's math ability I'd want to see a follow up, specifically for GPT-4.
Yup. Page 19: "we focused on the 9th-January-2023 version of ChatGPT"
Unfortunately, negative but robust results like this won't get a fraction of the attention this MIT paper initially got.