Researchers cannot always differentiate AI-generated and original abstracts
nature.com
nature.com
Isn't the headline actually "Abstracts written by ChatGPT fool scientists 1/3 of the time"? Having never written one myself, wouldn't the abstract be the place where ChatGPT shines, being able to write unsubstantiated information confidently? I imagine getting into the meat of the paper would quickly reveal issues.
This is a tautology: the thing that can be validated can be validated.
I wonder if having AI models available would make it significantly easier to identify material created by that model. Seems it would be easier, but it would still be a big problem. Nets can be trained to identify ai vs not ai on a given model. And that's without needing access to the model weights, just training examples. But when there are N potential models...
Edit: Second paragraph added later.
I think you and I are basically in alignment... what this tells me is that 14% of real abstracts are so bad that other human beings call their BS. Meanwhile, this AI stuff is kinda working 32% of the time in generating legitimately interesting ideas.
So at that point - yeah, that sounds about right. The 32% is still so low that it shows AI is not anywhere near maturity, whereas 14% of human-generated is crap.
And, yeah - a short blurb like an abstract seems to be exactly the kind of text that ChatGPT is conditioned to do well generating. As others below note - once a human starts reading the rest, the alarm bells trigger.
I don't know that interesting has anything to do with whether people found them plausible.
The bar is not "this is legitimately interesting," it's "I've seen real abstracts worse than this."
This works very well to more or less stay on top of things with today's frantic publication pace in many disciplines.
Previously, we've relied on a number of heuristics to determine if something is real or not, such as if an image has any signs of a poor photoshop job, or if a written work has proper grammar. These heuristics work somewhat, but a motivated adversary can still get through.
As the quality of fakes gets better, we'll need to develop better tools to dealing with them. For science, this could, hopefully, result in better work replicating previous works.
I'm quite likely being overly optimistic, but there's a chance for positive outcomes here.
[0]: https://www.science.org/content/article/potential-fabricatio...
Even if everything is fake, the code has value for further research.
It would be nice to have that as a minimun standard at this point, as I would prefer to see much less publications that can be trusted more than the current situation.
I'd call that human-computer science partnership. If it checks out, it's not fake. Nonhuman scientists are still scientists.
If someone trains a translation model between languages they don't know, is that someone a translator?
I guess the users of said model would be "translators" as they would be doing the translation (without necessarily knowing the languages either).
Making conjectures, deriving predictions from the hypotheses as logical consequences, and then carrying out experiments or empirical observations based on those predictions.
Not sure AI is up to that, and it's debatable if it'll ever be able to make and test conjectures. There is a difference between symbol manipulation (like outputting text) and actual conjecture.
https://en.wikipedia.org/wiki/Robot_Scientist
Adam is capable of:
- hypothesizing to explain observations
That's not to say scientists shouldn't publish their data; they should.
In the limit, any claims one makes are rejected unless you also provide a way to reproduce the claim.
That's a lot of work, but perhaps the world doesn't need more monkeys and typewriters creating faux Shakespeare.
In truth, the entire weird and crufty vocabulary is simply a common set of placeholders that makes it easier to grasp research, because the in-group learns to understand them as such.
"don't write like you have a pole up your ass."
It is also true that science literature contains a lot of jargon that encodes important information. But that doesn't excuse the fact that a lot of scientific writing could be improved substantially, even if the only audience were experts in the same field.
As an educated lay person who dips into scientific papers occasionally I completely agree with this. And now I have a nice phrase I can remember next time I read a scientific paper and think "why does science hate me?".
But first you need something good to write about, and a lot of papers fall short at that point. It doesn't matter how well it is written at that point (getting the "this is well written, but not interesting" rejection is painful but usually obvious).
It contained formatting rules. The text must be double spaced size 11 font, The table of contents must be laid out in this specific way, tables and figures have to be included and referenced in this particular fashion and so on.
It also imposed rigid restrictions on language usage and writing style it basically emphasized writing short sentences. No adverbs. Avoid passive voice, Write everything in first person tense and so on.
I asked my housemate who studied creative writing to proof read my thesis. He commented on how sterile and cold the whole things felt compared to the more florid prose he was used to.
So the way scientific papers read is very much by design.
Fake scientific papers that are written with the language, vocabulary, and styling of an academic paper have been a problem for a long time. The supplement and alternative medicine industries have been producing fake studies at high volumes for years now. ChatGPT will only make it accessible to a wider audience.
My point is that this is similar to the problem of art fraud or fake news or any other thing that involves faking media: the immoral act is the problem; the use of AI just makes it easier to do and harder to catch. Yes, it is a problem, but the heart of the problem lies in the immorality of the act, not in the technology per se. Perhaps the issue is not that "many people are immoral", but that the technology enables a larger proportion of immoral people to fool us.
By "politically-popular," I don't mean primarily in a red-blue sense. If you're publishing a study for therapists to read, it should ideally use a clever methodology to re-affirm their viewpoints.
And if AI can manage that, well: https://xkcd.com/810/
Nobody takes a serious decision reading only the abstract. Look at the tables, look at the graphs, look at the strange details. Look at the list of authors, institutions, ...
Has it been reproduced? Has the last few works of the same team been reproduced? And if it's possible, reproduce it locally. People claim that nobody reproduce other teams works, but that's misleading. People reproduce other teams works unofficially, or with some tweaks. An exact reproductions is difficult to publish, but if it has a few random tweaks ^W^W improvements, it's more easy to get it published.
The only time I think people read only the abstract is to accept talks for conference. I've seen a few bad conference talks, and the problem is that sometimes the abstracts get posted on like in bulk without further check. So the conclusion is don't trust online abstracts, always read the full paper.
EDIT: Look at the journal where it's published. [How could I have forgotten that!]
How can we yield results from an industry being lead by automated derivatives of the past?
Is an AI-generated result any less valid than one created by a human with equally poor methods?
Will this issue bring new focus on the larger problems of the bloated academic research community?
Finally, how does this impact the primary functions of our academic institutions... teaching.
Even a human researcher need experiments to validate ideas. AI can generate plausible ideas, so why not run the experiment and let it learn from the outcomes? The source of learning comes from experimentation, that's how models escape the derivative trap. AlphaGo invented move 37, proving AI can be creative and smart.
> I'm quite confident that there are cliques within "science" which are admitted without as much as a glance at the body of the papers.
It's possible. I don't know every single area, but I guess it's not common in most serious branches od science.
> Some people simply cannot be bothered to get past the paywalls, others accept on grounds outside the content of the paper, like local reputation or tenure.
My friends says they can ask Alexandra for a copy. A few years ago it was also common to ask a friend in another site that has a copy.
> Others are asked to review without the needed expertise, qualification, or time to properly understand the content. Even the most honorable reviewers make mistakes and overlook critical details. Then there are the set of papers which are (rightfully so) largely about style, consistency, and honestly, fashion.
That's why you RTFP instead of thrusting the journal or the reviewer or just the abstract. I've seen "dubious" papers published in good journal.
> How can we yield results from an industry being lead by automated derivatives of the past?
It's running now in natural intelligence. Every time there is a interesting paper, other groups run to publish variants or combinations with other result to reach the anual quota, or get enough point for the graduate student. Once the GAI can read all papers and combine them, there will be very few low hanging fruits to pick.
> Is an AI-generated result any less valid than one created by a human with equally poor methods?
For now, AI is to stupid to get the details right. That's why it's important to read the full paper instead of only the abstract. One AI is intelligent enough, the result may be as bad as a cheating human. Let's hope the AI has a good PRNG to create credible noise, because that's a part humans do badly.
> Will this issue bring new focus on the larger problems of the bloated academic research community?
Nah.
> Finally, how does this impact the primary functions of our academic institutions... teaching.
I don't understand the question. You fear that professor in some areas will just send fake papers written by ChatGPT4.0 and the journal and the community will not notice? There are a lot of predatory journals, open-peer-review journals and other bad journals that are publishing a lot of crap. Usually by professor in bad universities or just as a fancy achievement for the c.v.. A good AI will increase the amount of crap, but it will be just ignored.
No, actually I'm curious if this could open up the schedule for professors who want to spend more time teaching to design and develop better curriculums for their students. But I'm probably being overly optimistic about the number of profs who actually want to teach.
Also, accumulating published papers is important to get a new position in the future, so only teaching is a risk for your future career.
For teaching some of the topics to advanced students, it's important to have people that do research and is updated with the cutting edge results. Also to babysit the graduate students so they can publish their results. For teaching the students in the first years, research is probably overrated.
It's just been a while since I was inspired by anyones research I guess.
You can explain it to a technical friend.
Start explaining about the the two possible spin of electrons, and how it cause the existence of two currents inside a conductor. This is not important in a normal conductor like cooper, but it's important inside a magnetic conductor like iron. So yo get a different resistance for each one of the two currents, the one that has a spin that is in the same direction than the magnetic field and the other with the oposite spin.
You can make a sandwich of iron-cooper-iron. If there is no external magnetic field, the two iron parts have oposite magnetization and the total resistance is higher. If there is an external magnetic field, the two iron parts have the same magnetization and the total resistance is lower. Anyway, the difference is not too much.
[Ideally, your fiend should be boring now, uninterested about the abstract currents no one cares about.]
It get's more interesting if you have many layers of iron and cooper, because the difference is higher and they call is "giant". It is used in the heads to read hard disks, like in the laptop of your friend. [Your friend will never see it coming!]
It's interesting because it mix weird abstract quantum properties and engineering to make it more efficient and you get a device that everyone has. For some weird reason, no one talks about it. And the sad part is that SSD are killing the punch line of the story :( .
I am not sure if this is sarcasm or not.
Literally the whole world besides the researches read mainly the abstract and make even life changing decisions based on it. (just look at any twitter discussion linking to a paper)
I don't consider twitter discussion "serious decisions". When a thread here reach 100 comments, the quality of the discussion usually drops. I don't want to imagine how bad it's in twitter.
The problem is people overdosing from snake oil they read in twitter. I think there were a few cases with ivermectine, hydroxychloroquine and chlorine dioxide [1]. People should not self medicate, or at least understand proportions and that the dose makes the poison. They should see a medical doctor [2], that at least understand proportions and that the dose makes the poison and that they should follow the advice from the FDA made by people that read a few of research papers and extended reports, instead of only one press release about a preprint written by a moron that somehow got a position in a university or hospital.
I guess politician don't read the full paper (except Merkel?), but they hire experts to read the articles and give advice. If you are a politician and the "expert" is reading only the abstract without checking the full paper and the journal and other related stuff, you should fire the moron.
[1] The chemistry of chlorine dioxide is too simple and it's obvious that it can't possible work, so it makes me angry. The other stuff don't work, there were no good reasons to expect them to work, but at least it's not obvious with elementary chemistry that it can't work.
[2] I have horror stories about medical doctos too. Always ask for a second opinion for important stuff.
The day is already too short, with an expansion of journals. But, there's a sort of silver lining: many folks restrict their reading to authors that they know, and whose work (and writing) they trust. Institutions come into play also, for I assume any professor caught using a bot to write text will be denied tenure or, if they have tenure, denied further research funding. Rules regarding plagiarism require only the addition of a phrase or two to cover bot-generated text, and plagiarism is the big sin in academia.
Speaking of sins, another natural consequence of bot-generate text is that students will be assessed more on examinations, and less on assignments. And those exams will be either hand-written or done in controlled environments, with invigilators watching like hawks, as they do at conventional examinations. We may return to the "old days", when grades reflected an assessment of how well students can perform, working alone, without resources and under pressure. Many will view this as a step backward, but those departments that have started to see bot-generated assignments have very little choice, because the university that gives an A+ to every student will lose its reputation and funding very quickly.
It's like we've reached a fixed point, global minima for academic ability as a system. You could almost argue it's inevitable. Any system that looks to find abstractions in everything and generalize at all costs will ultimately learn to automate itself into obscurity.
Perhaps all that's left now is to critique everything and cry ourselves to sleep at night? I jest!
But it does seem immensely tiresome and deters "real science".
Yeah well we can't tell that now either. Maybe we can finally start publishing raw data alongside these "trust us we found something" papers that people evaluate based on the reputation of the journal and the authors.
As someone else pointed out, that system has already derailed decades of Alzheimer's research. It's stupid and broken and it should have changed a long time ago.
https://www.science.org/content/article/potential-fabricatio...
IMHO, it says more about the manic habits of journal editors than anything else.
> ... When given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated. Reviewers indicated that it was surprisingly difficult to differentiate between the two, but that the generated abstracts were vaguer and had a formulaic feel to the writing.
That last part is interesting because "vague" and "formulaic" would be words I'd use to describe ChatGPT's writing style now. This is a big leap forward from the outright gibberish of just a couple of years ago. But using the "smart BSer" heuristic will probably get a lot harder in no time.
Also, it's worth noting that just four human reviewers were used in the study (and are listed as authors). That's a very small sample size to draw conclusions from. The article doesn't mention level of expertise of these reviewers, but I suspect that could also play a role.
Some are focusing on the paper-mill angle. But I think the more interesting angle is ideation.
If researchers can't reliably tell the difference between machine and human generated abstracts, what kinds of novel experiments or even research programs could ChatGPT suggest that might never have been considered?
Even if you're skeptical or cynical, you'll still fall for well-written nonsense if it remotely feels authoritative, sane, reasonable. Especially when not a deep expert on the topic.
The above effect is even stronger if you have prior trust in the author.
I guess the reason is that extreme distrust is an uphill battle. It costs a huge amount of time and cognitive load to discover objective truth, often for little reward.
https://www.theseedsofscience.org/2022-general-antiviral-pro...
(after I had written the rest of the article and long after writing the academic paper underlying it.
Although, the published abstract reads nothing like the abstracts that ChatGPT generated for me because of the subtle but important factual inaccuracies it generated. But I found it helpful to get around my curse-of-knowledge in producing a flowing structure.
My edited, manually fact-checked result flowed less fluidly but was accurate to the article body’s content. Still overall glad I did it that way. I would have otherwise fretted over format/structure for a lot longer.
It is a tool. Ultimately researchers are responsible for their use of a tool, so they should check the abstract and make sure it is good, but there’s no reason it should be seen as a bad thing.
Remember that it wasn't long ago that Facebook had to pull its Galactica AI in a few days after people got it to output all kinds of nonsense: https://mobile.twitter.com/mrgreene1977/status/1593274906707...
Abstracts don’t have to justify or prove what they state.
They'd probably be better than what comes out of university PR departments.
https://en.wikipedia.org/wiki/Sokal_affair
https://en.wikipedia.org/wiki/List_of_scholarly_publishing_s...
But I'm lost at what those scientists are trying to find.. (?)
If ChatGPT can't do that (i.e. if it's attaching abstracts disjoint from the paper body), it's not the right tool for the job. A tool for that job would be valuable.
There is however a problem when the contents of the papers costs thousands/millions of $ to be reproduced (think GPT3, DALLE, and most of the papers coming Google, OpenAI, Meta, Microsoft). More than replication, it would require fully open science where all the experiments and results of a paper are publicly available, but I doubt tech companies will agree with that.
Ultimately it could also end up with researchers only trusting papers coming from known labs/people/companies
Kinda needed to be, given the rise of computer-generated proofs starting with the 4-colour theorem in 1976.
Do you mean proof assistant like Lean ? From my limited knowledge of fundamental math research, I thought most math publications these days only provide a paper with statements and proofs, but not with a standardized format
Code from ChatGPT could very well end up processing data from each of them — I wouldn't be surprised if it already has, albeit in the form of a researcher playing around with the AI to see if it was any use.
Reviewers work from an assumption that the data is valid, and reproduction (or failed reproduction) of a paper happens as part of the scientific discourse after the paper is accepted and published.
It's a single data point. Did anyone ever claim the editorial process of Social Text caught 100% of bunk? If not, how do we determine what percent it catches based on one slipped-through paper?
I'd expect scientists to demand both more reproducibility and more data to draw conclusions from one anecdote.
Is ChatGPT located in a central repository or cloud? Is centralized? If that, probably a bad idea.
A private company having access to your abstract before you publish it could easily lead to problems like plagiarism (even worse, automatized plagiarism) or give an unfair advantage to one in two teams running to publish the same result. Science has a lot of this cases.
This would seem to me a situation that is easily exploited by an AI that generate plausible text. If you pack enough jargon into your paper you will probably make it past several layers of review until someone actually sits down and checks the math/consistency which will be, of course, off in a way that is easily detected.
It's a problem academia has in general. Especially in STEM fields they have gotten so specialized that you practically need a second PhD in paper reading to even begin to understand the cutting edge. Maybe forcing text to be written so that early undergrads can understand it (without simplifying it to the point of losing meaning) would prevent this as an AI would likely be unable to do such feat without real context and understanding of the problem. Almost like adversarial Feynman method.