"Lots of people in CS are (almost surely) GPT-ing their peer reviews"
twitter.com
twitter.com
Maybe we're going to need a way for people to signal "If you send LLM-generated text to me under the pretext that you wrote it, I will (depending on context) urge that you be sanctioned by the academic/professional organization, refer you to academic discipline committee for suspension or expulsion, have you put on a PIP or fired, or downrate your friendship level."
I doubt this response will ever become dominant, in part because the tool is already really useful despite its limitations.
I'm not sure how it's going to shake out, given the tools are not likely to be static.
What if people used a disclaimer to actively signal, ChatGPT was used for the English here, at the top/bottom of emails/whatever when it was used. would it be as much of a problem for you then?
Nonetheless I've been tempted to add an "AI-free" or "written by a human" logo on my website. There just isn't a place where it would make any sense.
Okay, and now I've written enough that it turns out I feel like adding a disclaimer: this comment written by a human, with no help from an AI.
We know the exploitation is quite prevalent and we should keep saying so out loud every time there is an opportunity to do so.
I know some professors definitely can be exploitative, but learning how to do peer review well is a critical skill that can really only be developed by doing it - with good mentoring.
Good for you if you got lucky. Hurrah. You should be calling the exploitation out the loudest and most often on that basis, right? Otherwise it might erroneously appear that you're ok with your competition to secure an academic career being abused like this because it kinda works in your favor. And I'm sure you're not ok with it on that basis.
> Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences [ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023] could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates.
I will gladly utilize AI to feed these corporate drones.
Depth, sure. ChatGPT is neat, but much to lazy for a full review — it tends to give me a handful of suggestions and then stop. I presume this is because the training set is humans on the internet and that in turn means "make 1-3 points in response to the previous commenter, then stop", not "mark this exhaustively according to a well-defined rubric, including for spelling and tone".
Now, that said, I've just re-read your two comments, and they absolutely pattern match to the kind of thing ChatGPT says. I'm curious, were either or both actually from ChatGPT, or has that style just become so prevalent that humans are mimicking it?
See, the beauty in what we're now calling the joke, is that it's both a joke for those who know, and it's actual valid commentary that meets your expectations of this forum.
I guess it does make it easier for people to be lazy and get away with it. That being said, I’d be surprised if ChatGPT/Claude/Gemini actually provided _zero_ useful feedback.
Not exactly the same, but instead of my co-workers just pasting “LGTM” on everything, I’d probably prefer an LLM review.
Most people can't even be bothered to read a function they copy off StackOverflow to understand what it does, it's not a stretch to believe they aren't going to bother reading a whole paper.
>Not exactly the same, but instead of my co-workers just pasting “LGTM” on everything, I’d probably prefer an LLM review.
You could do your own LLM review before sending it to your co-workers, if you'd like that kind of review. It may catch basic syntax stuff or suggest ways to write something differently, but it won't catch bigger issues of logic or "is this a good idea". I'll admit to a lot of LGTM reviews going around in my team, but last week I saw someone make something extreme dangerous and the review stopped it from rolling out to prod. I am almost 100% certain an LLM would have said LGTM, as it was written fine and did exactly what it said it would do... it was just a horrible idea to do that thing.
Writing detailed, respectful letters to idiot officials is not a skill that I would enjoy getting or practicing.
Same reason as all the other tools when they were new: we're still all figuring out the best practices, the limitations, and the strengths.
Also: If all you have is a hammer, every problem looks like a nail — Outside my speciality, ChatGPT is my hammer; inside my speciality, I can see all the LLM's flaws while also having good working knowledge of many other tools.
How bad is the latency in your actual brain that this is a regular option for you?
"I use calculators in all my calculations now, and I'm really transparent about it. For $20, I now have a pocket calculator, and can at any time whip it out and calculate a long, difficult problem - and get a correct calculation based on my inputs just a few seconds later. Is there a single reason not to use it?"
Are the concerns in the other replies consistent with their own outsourcing of willpower & mental effort to technology?
https://hai.stanford.edu/news/researchers-use-gpt-4-generate...
Now it will be hard to tell, if a given review has been done exclusively by this drafting tool, or with some human input/oversight.
This would also surely create an LLM black market. One kid in school sets up his own LLM and then charges kids for access to write their papers with a system that won't log to the central DB.
Also, CS likely has a higher proportion of non-native English speakers compared to other fields. The increased frequency of certain tokens could be due to non-native speakers running their draft reviews through prompts like "Please edit this draft written by a non-native speaker, clean up infelicities and improve word choices where appropriate."
I presume people with English as a second language (a) lean on GPT more, (b) don't deeply reedit the output.
It leaves the reader to decide if it's a good or bad thing.
On the one hand scientists are not always good writers. Bring good at science does not imply good communication skills. Clearly tools like spell-checker and grammar-checker have been jn use for decades, is this not just the next iteration?
On the other hand LLMs are just probability machines. So while the review might accurately reflect the authors opinion, it may come across as "bland". All the personality of the reviewer is stripped away.
It's like fast-food. It's all the same. Which has up-sides. But it's considered "low quality" in part because it's "designed by focus group". By contrast Mamas Italian Kitchen reflects mamas unique personality.
Perhaps they are both OK, in their own way.