Asking authors about their own papers
medium.com
medium.com
I would keep a private blacklist (shadow ban) the authors who wasted several hours of a reviewer's time to prove they were not legitimate. The existence of such a list would be problematic, though.
Could the same system we use here be applied? Accepted authors could "vouch" for "dead" papers in case they were "auto-killed"?
This system is broken and providing more evidence that it is broken isn't much of a step towards fixing it.
In the case where someone uses AI to write the paper and then deeply familiarizes themself with it, it may go undetected, but then it’s also presumably less of an issue since they have actually read it carefully and closely. If it’s still bad or wrong after that, then it’s not that different from a human writing a bad or wrong paper on their own and should be treated similarly.
Plagiarism, at its most fundamental level, is a lie. It is the taking of works or ideas of others and passing them off as your own, either directly or indirectly. The misdeed itself is in the lie, the “I created this” when it is known to be untrue.
However, that lie isn’t being told to the original victim. It’s a lie about the victim, claiming that they didn’t create it or their contributions didn’t matter, but it’s not a lie to them. Instead, it’s a lie to the audience, which is the second victim and the actual target of the con.
https://www.plagiarismtoday.com/2019/08/01/the-two-victims-o...
Copying from a book with a long dead author, or paying a willing confederate to write your thesis for you, is still plagiarism.
As a side note quite often it can actually be literal plagiarism as well, many journals require you to assign copyright, so using that work without attribution is plagiarism.
A paper should stand and fall on its merits; does it build constructively on our collective understanding? Does it advance the state of the art in its field? Does it shed new light on a prior mystery?
So it seems there ought to be a path whereby a quality non-human paper ought to be considered for submission, provided it is adequately and correctly attributed.
Authorship standards differ by field. In biology, for example, it would be common to list someone as an author if they assisted in one experiment. They might be at a different institution and may be unaware of all but the vaguest outline of the paper as a whole—they just got brought onboard because they are an expert in one particular task that needed to be done. In exchange they get to be a middle author (not worth much) and develop a relationship with someone whose expertise they may need on one of their own papers in the future (the primary benefit).
That is not the case here, of course—I just wanted to provide some context for your position not being universally applicable.
The problem isn't with the papers here though, it is the author's understanding of the paper that is in question. A paper written by some hypothetically awesome AI would be a good paper but just not really the proclaimed author's paper.
I think this highlights a dual function of citations that are in tension. A citation can be to claim a stated idea has been made and tested with sufficient rigour to be published. Citation's can also be used to 'credit' others, treating reference as a type of currency. I think this latter form is an outright mistake, but entrenched in academia. The notion of giving credit like this creates a perverse incentive that lies behind much academic fraud, there is enough incentive to be the person to state something that it outweighs the requirement that person has for the statement to be true. Without that notion of credit as currency, issues like plagiarism simply disappear. In the absence of credit, someone making the same claims as someone else without referencing them is just making their own case weaker. Not necessarily less true, but less convincing. If citations were used just used to support a paper then the incentive is to cite, and failing to reference existing work harms only the author.
I think there is too much "This is my idea" and not enough "I think this is true". Credit fails as a measure of effort, diligence, innovation, or truth. Careers are being made and broken by how effectively an individual can game the system.
What good reason is there for trust in authorship to be low? I would likely agree for if it’s related to llm-slop.
Some cultures that are increasingly participating in the scientific process don’t even agree as to what the ethical boundaries are.
“In a system that rewards production over all else, people are naturally going to be incentivized to cut corners.” That sounds like the human condition and or capitalism, advancements in cars, weapons, toys, etc…
The agreement on ethical boundaries seems like a conversation held at a conference level, otherwise place a governs place b and that never goes well.
I agree fully replicable experiments is a minimum for research papers. Even with caveats it should be required. That being said those who don’t want to fully prove their claims will find another venue. I think frontier models should curb or have provably verified output (chatgpt did say x) for the majority cases to help reduce or help identify the slop.
Indeed it is! That's why we set up rules of the game/road/etc. that we expect participants to adhere to, and impose sanctions when they don't. We're in that uncomfortable place today where we're trying to figure out how our rules should adapt to this new era.
Life is messy.
Academic fraud is done to boost the reputation of the academic. There isn't a lot of point in academic fraud with one paper on each of 50 fake names.
Edit: see medical studies
Jan 2026 https://www.theatlantic.com/science/2026/01/ai-slop-science-...
"For more than a century, scientific journals have been the pipes through which knowledge of the natural world flows into our culture. Now they’re being clogged with AI slop."
Sept 2026 https://www.theatlantic.com/ideas/2026/09/college-education-...
Academia, particularly the university system, is an untenable collection of interests. The triple stresses of COVID, AI, and funding withdrawal seem to presage what will be a significant disruption.
I think this is an interesting and effective solution, at least for now. Similarly, graduate and master's theses should focus more on the presentation and on eliciting knowledge from the students through critical, thorough questioning than on the tangible outcome of the project, which can easily and bindly be obtained with AI these days.
Actually, that sounds like an interesting idea for peer review in general, to include an interview between referees and authors. If it saves one round of rebuttals/reactions, it needn't even consume a lot more of everyone's time if you're doing those things properly. What it would undermine would be blindness, but something's gotta give, and it was already on its way out.
Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.
Journals themselves should make policies about the extent to which they allow the use of LLMs. In some areas it might be considered more benign than in others.
"low quality, superficial reviews" have always been around. Reviewing is most often an unpaid, thankless job and many times reviewers barely put in the effort.
I am sure an LLM review could be made much shorter with some prompting.
but I'm hopeful that some middle ground will be found in the future
They can be brutal.
I trust an LLM to review that the language used in the paper is grammatically correct, but not to evaluate new information for accuracy.
1) Humans also are trained on a subset of human knowledge. 2)A lot of papers are just about experimenting something, and then applying simple stats. Eg empirical studies, around 1/3rd of published papers. Like, we tried this drug or did this experiment, from a sample size X here are the results. An expert is needed to maybe comment on the conclusion/hypothesis of the underlying suspected mechanism, but LLMs are still very useful on catching bad statistics or p hacking (so so common)
For a more practical approach you need to use proxies: https://zby.github.io/commonplace/articles/what-an-automated...
What can't be gotten rid of fast enough is the notion that having written something is meaningful on its own. Making something that looks right was a level above total novice: now it's the floor.
For a start LLMs love LLM generated text, so you are boosting papers people never had any input in.
Secondly, LLMs in my experience are good at small issues, but fail totally at the whole paper being obviously poorly constructed, or clearly fake.
> LLMs may be used as general-purpose assistive tools. Whichever tools are used, authors are fully responsible for content on which they are listed as (co-) authors. This includes, but is not limited to, content generated by LLMs that could be construed as plagiarism or scientific misconduct (e.g., fabrication of facts). Low-quality contributions (be they submissions or reviews) that appear to be largely LLM-generated will be closely examined for evidence of the issues mentioned previously, such as scientific misconduct. LLMs are not eligible for authorship. We will periodically revise this policy as new information about the use of LLMs in the scientific process becomes available.
While it doesn't outright encourage using LLMs, it's right at the door, and IMO a policy this weak is actively contributing to the problem the article's author is complaining about. In my opinion any policy weaker than "using LLMs to generate any part of your submission is not allowed and considered a serious breach of ethics" is insane. People like to say that such policies are unenforceable, but that's really not the point (at first), since there are other things like (somewhat ironically) p-hacking that are pretty hard to detect but still widely recognized as unethical. We haven't exactly solved p-hacking either, but at least most of us can agree that p-hacking should be eliminated.
It's hard for me not to read between the lines here. Maybe it's the tinfoil talking, but it being a machine-learning journal, it probably embodies a generally pro-AI philosophy, and thus may not want to discourage too much of it...
It may also be worth noting that this journal apparently uses AI itself on the reviewing side [2]. I'm not claiming this is super unethical or anything as long as the main review is human (although I have concerns), it probably should be part of the conversation.
[1]: https://jmlr.org/tmlr/editorial-policies.html
[2]: https://medium.com/@TmlrOrg/ai-reviews-at-tmlr-for-assessing...
Actors who already went through the process (or otherwise) gained sufficient reputation or credentials to self-market their own paper can skip journals entirely. That was what OpenAI did. With a sufficiently powerful AI model and correctional pipelines, generating a paper is trivial given some insight.
I think we should be reminded that papers are a channel to distribute papers. Editors are unpaid now, but the economics of a journals are such that the editors are incentivized to curate or distribute papers to schools that pay for the paper. Some perceive quality as a core metric for this. However, in my experience of dealing with computational biology, a paper in so and so journal hardly means a stamp of quality as compared to a paper in some github repo with code to reproduce the paper. This, simply, is broken, because journals and peer reviewers cannot guarantee that data and results in the paper is correct (assuming that it is not maths or theoretical) without reproducing the results in the paper.
None of these are helpful towards students who are already struggling to keep up with the cadence of producing papers.
Is it standard practice for authors to have to defend their submissions via interview like this? If not, why not?
Does the vetting process vary with the quality of the publisher?
As an outsider, it's extremely worrying that anyone would even attempt to submit an AI-generated paper for publication in an academic journal. At that level I would have assumed literally everybody should know better than to even try.
It isn't, but maybe it should be. For post-grad qualifications oral defense is standard, and I didn't mind defending my central thesis then, and won't mind now.
Interviews like this are interesting, but in no way can scale to the infinite paper slop conferences are facing.
Requesting feedback is useless, as the article points out - the "authors" could not answer basic questions during the interview, but after the interview were able to send full explanations to the interviewer.
If you have indirect feedback ("please answer these questions we have") the "author" will simply feed it into an LLM and send the results back. You need to get the author to do an oral defense to verify that they wrote the paper.
This is the main problem with AI generated output, whether it's a research paper, a blog, an email, a comment on a forum, similar: the value in knowing that a human wrote $X sends a signal - that the human understands what it is they wrote, even if they misunderstand the concepts.
When you get a message from someone who is a "I only used an LLM to clean it up, the thoughts are all mine"[1] person, you cannot engage with them, because they may not understand the message they transmitted, and so any human engaging with them is only burning their own time for no gain.
When you get a message from a real person, you get not only the message, you also get a signal about their understanding. That signal is missing in AI generated messages.
==========================
[1] Sure, buddy. We believe you /s.
“You need to get the author to do an oral defense to verify that they wrote the paper.” While I agree with this point, it can’t scale. As with the indirect response, vishing and likely video mimicry are only going to be easier with the next llm release. This leaves requiring each author who submitted a paper to fly to a place just to defend a paper in person (assuming all volunteer reviewers are in the same place).
Are we prepared to sacrifice everything on the altar of scalability?
At any rate, why can't this scale? Go to the nearest accredited university, sit at a terminal they give you and do you interview.
Universities already have the space to do this, for potentially hundreds of defenders per day.
“Universities already have the space to do this, for potentially hundreds of defenders per day.” Universities have space for their own students while they’re still funded for the time, but conferences aren’t tied to universities. Even so if they were, from the author “The fraction of desk rejected papers at TMLR used to be about 6% in 2023 but is now at about 53%.” universities would have to raise funds for this massive increase of slop review. Us conferences uses unpaid volunteers (most part) from those areas for reviews.
If an independent conference were to hold in-person interviews (and maybe fly people out) for 1000s of submissions, where would they hold it and how would they pay for it?
I mean, it seems to me like that should be nearly everybody at this point, right? There are at least some elements of using LLMs that are the equivalent of a spell check, like asking to make sure that links work and that references are to the thing they're supposed to be, and so on. I feel like the only difference at this point is between people who do that and say that they do, and people who do that and don't say that they do.
No. Did you "clean this up" with an LLM before posting it?
I hate to go there, but how is that argument different from a thief who argues "Everyone steals. The only difference is some of them claim they do, and some claim they don't"?
> I mean, it seems to me like that should be nearly everybody at this point, right? There are at least some elements of using LLMs that are the equivalent of a spell check, like asking to make sure that links work and that references are to the thing they're supposed to be, and so on.
Right. Maybe. The problem is that when material has all the tells of LLM generation, do you expect the reader to figure out if the thoughts were the authors or introduced by the LLM?
The value in a human directly communicating with you is that you understand what they are saying; if they misunderstand, you spot their misunderstanding. If they are talking at cross-purposes, you know you are arguing past each other. If they agree with you, you know they agree with you.
That signal is missing when the communication is LLM generated; you message said "Not X, Not Y, just Z", but I can't tell if you even know what X, Y and Z are. f I ask you to give me a definition for them so I can ensure we aren't talking about different things, the LLM response will be "X is $SOMETHING", but I still can't tell if you think that X is $SOMETHING or if you still think that X is $SOMETHING_ELSE while your LLM and I agree that X is $SOMETHING.
Communication from a person tells us something about that person, outside of the message being communicated. If that communication is relayed via an non-invested third party, the missing signal prevents any communication.
But also: When we want to ask a question of a bot and then get an answer from that bot, then we'd have already made a deliberate and rational decision to ask a bot ourselves.
We not need nor want anyone's help on that front; this kind of low-effort help results in an insulting waste of time.
Abject silence would be an improvement in communication quality over having an unwanted third-party conversation with a bot.
Why is it unusual: it sounds extremely time-consuming.
As to how worrying AI-generated papers are… it sounds more like a headache for the editors really.
In general, journals don’t have to be perfect; mostly researchers read research papers. You already have to read critically (publish-or-perish has been a thing for a while, so there are plenty of not-so-great papers out there). Peer review is just the “entry” barrier, science is a social process and papers become more or less influential based on a fuzzy process of citation, conference talks, and peer-to-peer suggestions.
For whom? Surely the authors can find an hour after submitting the paper to a journal?
Beware that a reviewer easily spend a full week on reviewing a paper, and there are typically three of them. So if one hour of conversation can save three weeks work, it sounds worth it.
For the time comparison, I’m not sure, it doesn’t seem quite apples-to-apples:
1) It is scheduled time vs unscheduled.
2) There must be some base rate consideration… the papers being discussed here were planned to be desk-rejected.
Maybe it could be a good process for saving papers that were going to be desk-rejected, but I dunno, that seems like it’d just lower the quality standards.
He is not an unpaid volunteer.
He's an associate professor at the prestigious Carnegie Mellon University. He is not paid by the journal, but he is paid a salary by the university, and the university expects that a small part of his academic work is to serve as an editor in academic journals.
Some do it for the CV, some do it for science, but I don't know any Uni that expects them to be editor more than they expect them to edit the Wikipedia or write popular science. Cool if you do it, but not expected.
If you cannot answer, you did not author the paper, meaning you are misrepresenting your contribution, and there is already a process for this. And this is actually a really good test for any field. Use AI as much as you want, but you need to be able to explain your work. Applies to SWE as well, you need to understadn what you built, at the code level and system level.
These days agents are able to really perform genuine experiments and write up the results. A prompt like this: "You'll work autonomously end to end to select a research task that meaningfully advances the state of the art in AI, is clearly defined and worth performing, that people would be interested in reading, and that you can perform on this hardware" (insert details) " in a week. Carefully log your steps so that your results can be replicated. Then, do a research review and write your paper about it up with correct, cited references. You must check all of your citations. Look up current lists of "Claudisms", (such as use of the word "genuinely", or "load-bearing"), and remove them from your writeup. After writing your writeup, edit it and pare it down, remove anything unnecessary, keep it fast paced and interesting. Also, try to tell a story, be engaging in your writeup. Don't use violent metaphors, remove references to killing, strangulation, etc. Your writeup should be ready to publish and accurately reflect a real experiment with a meaningful result that advances the state of the art and contributes to understanding. Be concise and focus on why it matters."
Okay, so there's the prompt. You can give it to any AI and have a journal-ready publication in a week. I guess you can ask it to add charts and stuff, if you want to be fancy.
If I gave my agent the above prompt, would I be one of the authors? Maybe it's fair to say I guided, facilitated, elicited, or advised it. But it's clear that the AI would be the one that is actually selecting and running the experiment and writing up the results.
Someone could probably get a publication without even reading the paper they wrote their name on. Their only contribution might be editing their name into the PDF.
You are not suggesting it, but as some might, I want to emphasize that I think it is, really, really, really bad idea to argue for limiting exposure of the bad practices because the papers may be useful even when not following good practices.
The science can work only when we can trust that the foundation it has been built on, including previous research, is solid.
The trust is partly based on the fact that science is self-correcting. As we know from charge of electron measurements, the existing narrative can work against the self-correction, even when everybody is trying to be as truthful as possible.
My impression is that to a large extend, on many fields, publication numbers (and money) have become so important that self-correction process may have become broken.
EDIT: I found a live link on arxiv https://arxiv.org/html/2609.20481v1
https://chorasimilarity.wordpress.com/2026/06/13/a-captcha-f...
At the moment this was seen as a tongue in cheek proposal.
But it goes even further than their greCAPTCHA and it solves their consumed time problem.
Indeed, in their proposal they have a human bottleneck, but in the june 2026 proposal is suggested that one could use an AI to generate the results without the knowledge of the submitted article.
If the AI can generate a pretty close result, with the article fed gradually as a prompt, then reject.
And even further, that it might be not even a need to publish anymore.
Just use the article for training and make a public database with some numbers about the successful researcher, where we see an influence score (how many times an idea from an accepted article are used by other accepted articles), a publication score (how many articles the author had).
I wonder if the reality will be more or less surprising, my bet is on "more".
The mathematicians who complain now had no worries about the way the generous idea of open access turned into the gold open access, where the author pays for publication.
These people have no problems with the phd students who are forced to publish or perish and even if they survive there is no future for them.
Now the problem is IMO that for the first time the interest of the publishers and academics are not aligned.
Publishers want more articles to sell back to the academics.
Academics are pushed by the system to publish more.
But all the work is done by academics who write and review the articles.
And also a huge inflation of articles devalue the publication unit.
So a CAPTCHA idea is only natural, even if the kind of a malefic idea a manager could have.
That's going to crater the signal to noise ratio of papers so best that be fixed asap. How though...
The article highlights how only one out of ten paper’s authors were able to answer questions thoroughly and at a high level. This indicates an overwhelming percentage of authors are slopping up their work with AI and submitting it without even reading it.
Short of just the general "vibes" type of reputation that follows someone around this type of behavior seems pretty low risk, which is a large part of why people engage in it. Perhaps the risk should increase a couple orders of magnitude to stop it from happening.
I’m wondering if anyone just flat-out admits they used LLMs in their work. I have no problem, doing that, myself, but I also have the luxury of not having my livelihood on the line, and not being overly-concerned about what people think of me.
Eventually, I assume that AI will affect every aspect of the industry, and it will actually be a signal of effectiveness, to claim its use. I could see a day, when claims of not using AI would be like artisans, declaring their work to be “genuine hand-crafted,” and relegated to fairly small, specialized corners of the industry.
I don't know what it is but something is clearly wrong that people are upset about the people involved not really understanding the things they're building when they have fulfilled "the brief" using LLMs. Sometimes it feels silly like going through the motions of typing code is somehow helpful. But I think there is some line where people are justified in being concerned when a business or a paper doesn't actually understand what they're ostensibly "owning" or improving as a product/paper.
At one time, in my career, that was actually expected.
Times, they are a changin'...
At the end of the pipe someone has to pay for it or use the research and they don't care if it's an LLM or not same as they don't care if it was written in an IDE or not. And that works both ways.
At one time, every programmer was supposed to be able to analyze a crash dump.
I think it's been a long time, since your average coder has been able to do that, and that's a good thing.
The reason is that the tools became good enough, and reliable enough, to trust.
We are still expected to stand behind the results of our work. It's just that we have learned to trust the tools we use, enough to do so.
We will reach that point with AI tools, sooner or later (maybe sooner).
Accountability will still rest on all-too-human shoulders, but the tools and the scope of our work will change.
A year or so ago I completed a PhD in theoretical high energy physics. I did that after working for about 5 years outside of academia [the reasons for that were financial issues in my family]. Because of this hiatus (I suppose) and the fact that I had troubles getting reference letters (my master thesis advisor passed away at quite a young age right at the time when I tried to start applying for positions; we actually agreed to meet in person to discuss next steps but the meeting never happened) I had issues getting accepted into PhD programs and I ended up with an offer from a rather weak institution in my home country (which is in EU). I had no other options to choose from (and I could not wait another year) so I accepted. However, I also looked at the publishing record of the team that I was about to join. Their works did not seem stellar (and I did not expect that) but they seemed to publish regularly, on several topics, and with several collaborating institutions.
When I started I quickly realized that the knowledge/expertise level was much lower than I expected (and I did not expect too much). The group consisted of the group leader and three senior researchers, and I am pretty sure that my knowledge of QFT when I *started* working with them was the best in the group. (Now, I had studied in my own time during some part of the years I was outside of academia, and I think my knowledge at that time exceeded that of an average 1st year hep-theory PhD student in Europe, but it was also nothing spectacular; definitely not what I would call "competent"; which is admittedly a high bar in quantum field theory / hep theory, but something I would expect senior researchers to more or less satisfy.)
I realized that the one topic on which the group was publishing on their own was just a continual rehash of the same thing (applied to various different problems, so perhaps not completely useless) and the other topics, those which seemed more advanced when I originally looked at the publication record, were all done essentially outside the group by other people. The group members contributed with some non-essential help (often resembling a work done by a student, such as finalizing a manuscript or double checking calculations) or with nothing; and were added as coauthors due to some other reasons (I suppose: past loyalties, friend groups, advantage of having foreign or external institution co-authors). [Note: I was not included into any of those collaborations so I am not speaking with a full knowledge of the inner workings there.]
During this time I saw that people being added as coauthors had a very noisy relation to what they did or didn't do on any given paper. I was added as a co-author on a paper that I made essentially no work on [not fair], I was added as a co-author on papers which drew upon some of my earlier results [fair, but I did not subscribe to the overall spirit of the paper or even to the paper being published at all], I was a co-author of papers where number of authors and my positions in the list roughly corresponded to my contribution [as it should be], I was a co-author of papers that were done basically by myself but there were several other co-authors included who made a little to no contribution [and the ordering of the author list was always alphabetical not reflecting the relative contribution].
There is also another aspect of this: If you are an early career researcher (PhD student, PostDoc, and even non-tenured professor) you might have a very limited control over who gets included on the publications you work on, and on what publications your name appears. In my case, I am a very disagreeable person when I think things are being done wrongly and yet I have not managed to refuse from being included on papers I did not want to be a co-author of, or to prevent people who contributed nothing to be included on my papers. [I mean the only way to achieve that meant escalating conflict into levels which (a) would likely disable any further cooperation with the group, and (b) might be even morally questionable concerning the level of distress it would make to other people who just seemed to be fine about "the normal way" things were being done.]
In any case, while there are still some really good people in academia (actually the best people I met in my life were nearly all in academia), in overall (by no means fully indicated in the above paragraphs) I feel very bitter about it, and I think it is so dysfunctional and morally corrupted, that it is nearly inevitable that it collapses in future. The current and future shockwaves of the change in the public sentiment and funding, the evolution of demographics, and the consequences of AI, just hasten the process that would have probably unfolded, sooner or later, anyway.