Having said that it is also possible that the mistakes and claims were a human error, sure a lot gets ai generated these days but there is a chance in which case the accusation does not look so severe anymore.
Having said that it is also possible that the mistakes and claims were a human error, sure a lot gets ai generated these days but there is a chance in which case the accusation does not look so severe anymore.
I've noticed this over and over again with "professionals actually prefer LLM responses" studies. Typically the human generated responses seem better to me on a quick sample, but if I had to review 50 of them I'd probably start taking lazy shortcuts; using superficial language aptitude or factual comprehensiveness instead of critically reading.
It does seem like the human judges here might have given credit for e.g. a 20pg arXiv paper without actually reading it. I can blame them professionally but emotionally I have nothing but sympathy. I truly hate LLMs.