There really is no arguably about it. There isn't a verifiable definition/criteria for general intelligence that GPT-4 fails that a significant chunk of the human population doesn't also fail.
GPT-4 performs nearly all tasks it's given at at least the average human level, usually well above the average baseline. In some cases (nlp, analogical reasoning - https://arxiv.org/abs/2212.09196, and some others it's top %)
You basically have to make up your own definition that can't be substantiated (don't worry, plenty seem to do this). It almost always boils down "Well it's not "real" understanding, the difference is huge, i can't show you this supposed huge difference but trust me bro, it's not real"