What should I do if I suspect one of the journal reviews I got is AI-generated?
academia.stackexchange.com
academia.stackexchange.com
1. Don't underestimate how bad human reviewers can be. I've seen really bad reviews before. But the worst were for conferences, not journals.
2. The job of an associate editor is to field the reviews and make decision recommendations to the senior editor. A good associate editor will take care of this stuff, but may let a bad review through for the sake of the process. They might emphasize a particular review to help the author understand what the editor actually thinks is important, as opposed to letting the author think that all reviews are equal. That being said, it's up to the author to respond. If a reviewer is unequivocally wrong about something, the author can explain why they didn't follow the reviewer's recommendations. What the senior editor (and to some extent the associate editor) thinks is what matters, not what the reviewer thinks.
3. If the associate editor is not doing their job of fielding and reviewing the reviews, I question whether the journal is actually a top journal. My impression thus far is that top journals take their editing seriously. So far in grad school, I've met multiple editors from multiple journals, and gone to multiple journal workshops. The amount of work these people pour into doing journal work, many for free, is staggering. The burnout rate is significant accordingly, but the ones who stay keep it up because they want to be serious custodians of their discipline's research authority. It's a massive amount of work. I'm not sure I'd want to do it myself. I can't imagine such people brushing aside bad reviews and not realizing how bad they are. This is partly why good journals also have workshops to teach how to review. It's not easy to become an associate editor either. You need to become respected enough in the community to get nominated by editors and then voted in by editors. They have a standard based on how much they respect researchers.
Now... it's possible that my discipline (information systems) is unique in this manner. Is it possible that the top journals in computer science, physics, or other don't take this seriously? I doubt it?
Keep "AI" out of it. As described, the (suspected-AI) review seemed to only be based on the Abstract (didn't bother reading the rest of the submitted paper), and mentions several papers from irrelevant fields. Politely suggest to the editor that that reviewer was obviously struggling to review a paper well outside his area of expertise, and might best be replaced with a reviewer who is a better fit for the subject matter of your article.
This is only irrelevant if you place no value on the time of yourself and others.
I do get the point that LLMs make producing crap easier but that's somewhat independent of LLMs being used generally--which is going to happen in any case.
We should not be predicating our concerns about LLM-generated content solely on its "quality", because ultimately, the problem with it is that it is generic. I think it unlikely that it will have the ability to produce a genuine and thoughtful critique of a journal article until and unless there are significant breakthroughs, possibly even to the level of achieving AGI or something like it. Even using a more-advanced review-specific LLM like I describe above more widely does present serious concerns, because it runs the risk of suppressing articles that deviate from the "norm" in ways that the LLM doesn't have any way to appreciate, but which can present the findings better or even make the science better.
As the percentage of garbage that goes into peer review (or any other filter) increases, the percentage of garbage that manages to sneak through will increase.
Someone submitting AI generated reviews becomes relevant again when deciding whether to keep a reviewer around - a pattern of useful looking but useless and time wasting 'contributions' is relevant.
Basically don't over index on whether someone is using AI to be a shitter, focus on the problematic behavior.
Historically, we speak of "plagiarism" as being something you do against human text, because human text is all there was. But I would suggest that most of the issues with plagiarism are actually around misattribution, which means that it is perfectly sensible to speak of "plagiarizing" an AI. The AI may not be victimized, but victimization is not the only issue with plagiarism and most or all of the rest of them apply here. It matters over time where the text comes from. Even if the text of the review is high quality, in order to tune the editor's own tracking of reputation they need to know if it is from a human reviewer, GPT-1, GPT-7.5, or NotGPTAtAllSciAI-2026.
This is especially true in this case, because the entire point of a reviewer's review is that they are doing something the editor is not supposed to be doing! If the editor has to do a deep due diligence on all reviews, the reviewer are failing to provide any value as the editor might as well directly review the paper in question. So reputation is not something we can just wave away with "well if it was a good review it doesn't matter"; trust is a huge deal here. The editors need reviews to be properly attributed. Even if they are fine with AI reviews they need to know they are from AIs, and as I said, which AIs.
The reviewer certainly shouldn't depend on the results in either case. But it certainly seems reasonable to refer to references that aren't solely in their head.
And for the same reasons I laid out, if a reviewer thinks they need to "refer" to an AI, they should just tell the editor they're not qualified. The editor does not need a reviewer to serve as a middleman between them and a GPT. The reviewer is adding no value at that point; the editor can already "refer" to an AI themselves if they want to.
If you give busy reviewers an easy "out", where they can just run the paper through an LLM, do a bit of editing then send off the review, people are going to do exactly that.
And the resulting review, with the right editing, might seem perfectly plausible and human-like. But that review isn't going to be able to offer suggestions with insight from recently published papers. It isn't going to be able to point out issues with the data, or with the statistical analysis, or with the paper's logical conclusions.
Maybe someday AI will be capable enough to replace the role of human reviewers. But right now, encouraging this practice is just going to let a lot of bad science slip through to publication without genuine peer review. (even more than the large amount that already does, let's be honest ...)
The small issue is whether people are responsible for what they publish under their own name. Seems like a straightforward "yes", and whatever helper tools they use are irrelevant.
The much bigger issue is why scientific publishing's standard for a review is only "plausible and human-like", allowing people to to submit a LLM generated summary of an abstract without fear of responsibility.
Reviewers looking for an "easy way out" might be inevitable if they continue to remain uncompensated and uncredited for their time and efforts while journals get all the profits for other peoples' research.
I read a review of The Singularity is Near by Ray Kurzweil where it was described as seeing a table full of what appears to be very delicious food, but it is then revealed that there is absolutely some amount of dog feces mixed in with some of the dishes. You can't tell which is safe, and which is carefully crafted with dog feces.
An LLM in a peer reviewed journal currently has no place, unless it is part of an experiment where it is trained on the Journal's body of work and then tested for accuracy with future articles. As the tech progresses it may find a place but if it takes twice as long to fact check the LLM output it's saving nobody time and possibly hallucinating in hard to catch ways.
Humans usually try not to lie, and there's a particular shape to the sorts of details they tend to forget / confuse. Compulsive liars often don't even notice themselves lying. That's closer to an LLM. I trust the output of an LLM about as much as stuff George Santos says.
I get what you're saying, but this is exactly the point of peer review though. It wouldn't be worse if the original author was doing shoddy work in some parts.
This isn't one of those contexts. it's called peer review for a reason. You don't get to outsource your duty to either a machine or some random person. It's explicitly you others have vested their trust in.
>The output is either good/reasonable or it's not.
In the world of human beings this isn't the only thing that matters. Reminds me of Zizek who pointed out the end result of the "AI revolution" isn't going to be machines acting like humans, but the reverse, humans LARPing as machines. Humans as obtuse as robots, rather than the other way around.
It's also good for the editor to know about. LLMs represent a new acute threat to review quality that they may currently be underestimating. I've literally heard of people bragging about using ChatGPT instead of doing reviews themselves. People who aren't LLM experts don't necessarily understand their limitations or that using them in this way should be unacceptable. The editors should know so they can improve the communication of review expectations.
Those using AI tools in such situations should be expected to remove anything from the LLM's output that they can't verify with their own expertise. Reviewing out of your expertise doesn't necessarily inevitably lead to mistakes, but unchecked AI output will.
In practice - my advice is for the academic, who is trying to get an article published in "one of the well-reputable journals". That is a weak hand to be playing. Vs. the journal's editor is in a far stronger position, to hit back hard at whoever seems to be farming out their review job to a cut-rate bot.
Edit: 's/is farming/seems to be farming/'
Or for that matter if it had been dictated - and then unchecked?
Can happen to students too (sleep deprivation will do it), but not with the same inevitability as the LLM tool. To phrase another way, you could set up the 'human factor' in a way that you can trust unchecked output from another human (e.g. checking their expertise in academia, or if they are a commercial aircraft pilot, checking their pre-flight notes on how much sleep they got), but not for LLMs.
The author can't (and shouldn't) do anything directly about the anonymous reviewer, all the responsibility, authority and duty is up to the editor, who at least knows who that person is.
Last I heard, the one had been censured by the court, but courts generally have no power over law licenses. We might have to wait awhile to find out if there will be any more serious repercussions.
I think that in many professional settings, we might in the near future discover that some large fraction have been "faking it until they make it", but without the "making it" conclusion.
The fun part is when congressional staffers use this for gigantic 10,000 page bills too large for anyone to catch it before the vote. It might already be happening.
Also, at least some of the "increased productivity" boils down to "I spent less time thinking about the underlying problem while I was composing text and copy-editing".
That's what they are there for.
In this situation, it would be reasonable to raise your concerns with the journal editor. While you may not have definitive proof that the review was AI-generated, your observations and the results from the AI detection tools provide enough basis for a respectful inquiry. Expressing concerns about the review process is important for maintaining the integrity and quality of academic publishing. However, it's crucial to approach the matter diplomatically, focusing on seeking clarification rather than making accusations. Remember, the goal is to ensure constructive and relevant peer review for your paper, not to challenge the decision or the reviewer's credibility.