First they have a very small sample size of around n = 150. Second, they don't represent average husband, they use people working on mechanical Turk for $9 an hour.
Finally, they don't have any standard measures of divergent thought - they just use 4 different types of question answering problems, and those are supposed to measure creativity. Whether they do or don't would require us to evaluate the questions.
It's important to be clear that this is saying that LLMs might be more creative than low paid workers on mechanical Turk, given no training time. That's not a very good representation of the average human, or the fact that humans need training.
So for example, an average doctor or scientist might still be more creative than LLMs.
Tldr; this is not a good paper to throw around if you want to convince people. It's truly awful in its design, and the way it's presented is not always honest.
So, yeah, between that, and that both your average person and someone arguing from Turing tapes can show AI doing things not in the training data, I'm going with actual data thats being "thrown around" rather than the rando blog post that argues from "by definition it can't be creative because it isn't original because it only repeats training data" Call me when they're at N=152 on a blind test.