Oops, sorry dang. Being a CS grad of a similar school I've taken that blow enough to be numb to the whole thing.
The paper explains that they used a panel of evaluators to judge the "humanness" of the interactions : )
Said panel of evaluators found that AI agents pretending to be humans had more "believable" responses than humans pretending to be AI agents pretending to be humans. So that's... a result.