Delving into ChatGPT usage in academic writing through excess vocabulary
arxiv.org
arxiv.org
I'm not quite sure this follows. At the very least, I think they should also consider the possibility of social contagion: if some of your colleagues start using a new word in work-related writing, you usually pick that up. The spread of "delve" was certainly bootstrapped by ChatGPT, but I'm not sure that the use of LLMs is the only possible explanation for its growing popularity.
Even in the pre-ChatGPT days, it was common for a new term to come out of nowhere and then spread like a wildfire in formal writing. "Utilize" for "use", etc.
It's not people picking up new words, you have to understand academia is mostly about shitting out as many papers as you can and make them as verbose as possible, it's the perfect use case for LLMs, this field was rotten before, chatgpt just makes it more obvious to the non academia crowd
For them not using chatgpt would be like sticking to sailboats while the world moved to steam engines
This is the classic case of publish-or-perish, since publication metrics are ubiquitous in all aspects of academic life unfortunately. Measuring true impact is the goal, but it is hard problem to solve for.
> and make them as verbose as possible
This is just laughably wrong. The page-limits are always too low to fit all the information one wants to include, so padding a text is just not of interest at all.
With that said, I wouldn't be surprised if people use ChatGPT a lot. If for no other reason most academics are writing in a language (English) that is not their native language and that is hard to do. Anything that makes the process of communicating ones results easier and more efficient is a good thing. Of course, it can also be used to create incomprehensible word salads, but I've seen a lot of those in the pre-LLM times as well.
Making them as verbose as possible? My experience from grad school and with friends who are now faculty is that literally everybody's first draft is above the page limit and content needs to be cut.
I'm in a similar sort of circle and I can say that this, though I've not measured rigorously, does strongly square with my own anecdotal experience, and it's especially prevalent with people whose first language is not English but have no choice but to publish (frequently) and apply for grants in English.
So many of the writing assistant tools customarily used for first-pass proofreading have gone straight into full LLM integration and no longer simply just check grammar and basic "elements of style" issues.
Also professional (human) proofreaders are fantastically expensive for the very limited amount of assistance date provide. A couple thousand dollars more (source for this number, a proofreader recommended by Oxford U Pub) for someone to fix some minor semantic redundancies in your 35 page book chapter wording is financially unreasonable for a lot of people in academia.
> We will note that they has been in consistent use as a singular pronoun since the late 1300s; that the development of singular they mirrors the development of the singular you from the plural you, yet we don’t complain that singular you is ungrammatical; and that regardless of what detractors say, nearly everyone uses the singular they in casual conversation and often in formal writing.
I assume due to the societal shift to be less focused on gender.
delves
crucial
potential
these
significant
important
They're not the first to make this observation. Others have picked up that LLM's like the word "delves".LLMs are trained on texts which contain much marketing material. So they tend to use some marketing words when generating pseudo-academic content. No surprise there. I'm surprised it's not worse.
What happens if you use a prompt containing "Write in a style that maximizes marketing impact"?"
('You can't always use "Free", but you can always use "New"' - from a book on copywriting.)
On the upside, the advent of AI has made both vendor selection and project winning simpler. For the former, I can filter within seconds. For the latter, showing examples of competitors’ AI-derived speech is enough to get them eliminated.
Probably exactly what academics are being told by their university PR departments.
But I guess I'll just get flagged for being a GPT now.
https://s3.documentcloud.org/documents/3894798/CIA-RDP78-009...
Despite being published in 1962, I find has a lot of advice that still works today.
While the word frequency stats are damning, there doesn’t seem to be any evidence presented that directly ties the changes to LLMs specifically.
It would be interesting to analyze how well a language model is able to predict each abstract. In theory if the text was largely written by a model then a similar model might be able to predict it more accurately than it would a human-written abstract. (Of course the variety of models and frequency at which they're updated makes this more difficult.)
"Delving".. Sounds like the authors might have used an LLM while writing this paper as well.
"The AI researchers delved too greedily and too deep, and awoke the BaLLMrog."
(and do I have to deliberately avoid it now, to avoid sounding like AI?)
https://trends.google.com/trends/explore?date=all&q=delving&...
An old partner edited papers for a large publisher - largely written by non-native speakers and already heavily machine translated - and would almost always use ChatGPT for the first pass when extensive changes were needed.
She was paid by the word and also had a pretty intense daily minimum quota so it was practically required to get enough done to earn a liveable wage and avoid being replaced by another remote “contractor”.
It seems strange that something that experimental became essential that quickly.
“Old” might be intended to mean “former” or “ex”, which could be as recent as merely weeks ago. (i.e. ESL / lost in translation issue).
It’s an issue I’ve noticed personally, as I’m seeing an increasing number of reviews that lack substance and are almost entirely made of filler content. Here’s an excerpt from a particularly egregious recent example I ran into, which had this to say on the subject of meaningful comparison to recent work:
> Additionally, while the bibliography appears to be comprehensive, there could be some minor improvements, such as including more recent or relevant references if applicable.
The whole review was written like this, with no specific suggestions for improvement, just vague “if applicable” filler. Infuriating.
I guess LLMs have removed some of the tedium from the process while making it more tedious for the recipient. That's annoying.
https://trends.google.com/trends/explore?date=today%205-y&q=...
The leap from how the LLMs write now to how a professional sounding scientist might write their paper is probably not that big.
Its possible now to train them on your own writing style.
Then run the whole lot through Grammarly in academic mode to get rid of flowery words and tortuous structures.
Too many academic papers try to impress with complex language rather than explaining what needs to be explained in a pithy and succinct manner.
On thing that is interesting is that they allow you to add additional prompt, because sometimes you don’t want an email to actually sound like you but it to sound like the writer you want to be (leveling up so to speak)
human and machine both should aim for brevity and clarity, and feel shame otherwise.
then we can read more and better in our lives.
We truly live in the informations age.
On a side note, do people still use ChatGPT to fill out their papers? I found Claude to be way better at spitting out more content. In my experience, ChatGPT has been in the middle while Gemini is the worst, it even cuts-off in the middle of the sentence.
So yes - AI is seen as useful for people gaming a system for personal advancement at the expense of global scientific progress and the betterment of humanity.
I see no evidence actual findings are changing.
Lots of people have good science to contribute and suck at writing English. Making findings easier to digest isn’t bad per se.
I got some help from ChatGPT while writing a recent paper. At one point while writing, I couldn't find a straightforward way to express myself. I was stuck - and then every attempt at writing the same paragraph came out worse than the last. Eventually I gave my draft to chatgpt and asked it to come up with a few suggestions on how to express my ideas more clearly. That really helped. I adapted some of its writing and ended up with a better final result.
I'm not sure if anything chatgpt wrote is in the final paper. But I really appreciated the LLM assistance to help me express myself.
I'm sure there's a lot of much more egregious examples where chatgpt wholesale wrote large sections of some papers.
But its really not the tool that's at fault. Its how the tool is being used. And we simply don't have social norms around that yet, because figuring out norms takes time.
The calculator was the same. When is a calculator acceptable in a classroom? In an exam? Our teachers at the time justified "no calculator" policies by saying "You won't have a calculator in your pocket when you're going through life!". And, well, that turned out to be very wrong.
I am assuming your paper was for a class, given your comparison to calculators in a classroom. If not, then my point does not apply.
>Our teachers at the time justified "no calculator" policies by saying "You won't have a calculator in your pocket when you're going through life!". And, well, that turned out to be very wrong.
The only time we weren't allowed to use calculators was for things like arithmetic and fraction manipulation. It would indeed be absurd to have to consult a calculator for every numerical step you have to take in math work for subsequent years, so I think your teachers were right.
And If it levels the playing field between native English speakers and foreigners that is good in my book.
I don't think I did. And no, it wasn't for a class. Please be more considered before throwing out accusations of academic misconduct. That is a very heavy term.
Anyway, I don't think what I did is any more "academic misconduct" than using Grammerly is academic misconduct. They're both uses of LLMs, sure. But the uses are different and we need to start differentiating them. If you think AI generated art is plagerism, would you therefore conclude that photoshop's content aware fill is plagerism, since it uses the same AI models? Or Apple's smart selection tool?
> The only time we weren't allowed to use calculators was for things like arithmetic and fraction manipulation. [...] so I think your teachers were right.
(Emphasis mine)
Seems pretty strange to assume my teachers were right given you have no idea what country I went to school in or what their policy was on calculators.
I don’t care if anyone uses LLMs to help write papers as long as it doesn’t falsify findings.
I’m sorry for this, but if some members of this community are ready and willing to make claims of academic misconduct over the idea of using an llm as a writing assistant, I don’t think I trust this community enough to go into more details in this thread.
Please throw all of my comment into whatever LLM you used to respond, as I pointed out and explained the academic contextual assumption I made. If you are focusing on my phrase "for a class" here, and this paper was actually for an academic conference or journal, then yes! you have committed academic misconduct.
>f you think AI generated art is plagerism, would you therefore conclude that photoshop's content aware fill is plagerism, since it uses the same AI models?
Absolutely, if you are talking about some of the new ones that use stable diffusion and other visual generative models. It might be fine if you're doing it in a commercial setting and those models' licenses authorize you to do so, but using that for a situation in which the art is assumed to be yours (an art class project or for a commission in which the contract states you are producing it) would be fraud.
>Seems pretty strange to assume my teachers were right given you have no idea what country I went to school in or what their policy was on calculators.
Given that you didn't really dispute my point and rather attacked the grounds for it, and that you invented a false premise of your own (that calculator policies are a per-country thing ???), I'm gonna assume your Australian teachers' policy was applicable to my point.
>Please throw all of my comment into whatever LLM you used to respond [...]
And there goes your credibility. At least give strangers the benefit of politeness, rather than making a fool out of yourself.
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
> When disagreeing, please reply to the argument instead of calling names.
> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith.
From the HN guidelines. I see that your account is relatively new. You're right; this isn't reddit. Please have a read: https://news.ycombinator.com/newsguidelines.html
Rude. Rude and wrong.
> If [...] this paper was actually for an academic conference or journal, then yes! you have committed academic misconduct.
This is a surprising perspective given that in many ways I had the same conversation with chatgpt about my work that I might have had with a colleague or a spouse. The biggest difference is that chatgpt's suggested improvements weren't very good, and none of its suggestions ended up in the final text. Is it the product of plagerism?
I absolutely believe that ChatGPT is being overused in academia, just like stackoverflow is in professional programming. But I don't think we have any consensus as to where the line should be between "acceptable use" and plagerism. To me, banning the tool outright would be as silly as banning google search, banning calculators outright in schools or banning artists from using photoshop. (And yes, people made the same argument about photoshop when it was invented - that art made with it wasn't really art because the computer did all the heavy lifting.)
> Given that you didn't really dispute my point and rather attacked the grounds for it, and that you invented a false premise of your own (that calculator policies are a per-country thing ???)
Of course. The grounds for your argument is where it unraveled. You claimed that my teachers had the right calculator policy, but you didn't (and still don't) have enough information to make that claim because you don't know the policy my teachers had.
I would have told you if you asked, but you didn't. You simply went on the attack assuming you had all the facts. And now in this followup message, you have doubled down on your mistake.
For reference, of course the calculator policy at my Australian school in the 90s was different from yours. We generally weren't allowed calculators on the school grounds at all. That started to change right at the end of high school, when they started allowing calculators in class and non-graphing calculators in examinations. (They checked as we walked in the door). But for most of my grade school experience, calculators were banned.
Yes, it's possible to scoop academic colleagues and steal ideas from spouses. Plagiarism applies there too.
>But I don't think we have any consensus as to where the line should be between "acceptable use" and plagerism.
Perhaps, but using it to generate ideas you claim are yours, let alone quoting it directly, are all well past that line. Unless you put ChatGPT in the author list of the paper, in which case, sure, fair play.
>For reference, of course the calculator policy at my Australian school in the 90s was different from yours. We generally weren't allowed calculators on the school grounds at all. That started to change right at the end of high school, when they started allowing calculators in class and non-graphing calculators in examinations. (They checked as we walked in the door). But for most of my grade school experience, calculators were banned.
Unsurprisingly, this doesn't change or contradict my point at all. Any math you're doing in high school should not require a calculator, or at the very least can be taught without any calculators used.
>Rude. Rude and wrong.
I'm speaking with a tone appropriate for the nature of my company here (an academic fraud and one or two people who support him)
I like the idea that lots of novel scientific ideas were really invented by someone's spouse who doesn't work in the field. One day they happened to look over their partner's shoulder on a whim and said "Oh, honey, surely you mean this?" - and thus the hard scientific problem of the day was solved! Maybe that's how software gets written, too? Perhaps we should add people's spouses to the contributor lists for software? I'm sure many engineers talk to their husbands and wives about their work over dinner. It is simply unacceptable how few spouses end up in the contributors list!
I suspect what's really going on here is that we have different ideas of what constitutes "plagiarism". What is your working definition of that term? How much support would my spouse have to give me before I'm ethically required to add them as a coauthor? Would talking about the topic over dinner be enough? What if they read over my work and pointed out a spelling or grammar mistake? Should AI based code assistants be banned outright (since they are trained on opensource code)? Should code assistants be listed as contributors?
For context, I've worked at multiple universities in multiple countries. And I've spent several years teaching. That has involved telling my students where my line for plagerism is. And getting students in trouble when they cross that line. It sounds like by your reckoning I've been doing it all wrong all this time. So please, enlighten us all. What is plagiarism? How much LLM assistance, exactly, is too much in an academic context? And, if you'll indulge my curiosity, how much experience do you have in universities?
(Sadly, I must also acknowledge that I talked to my partner about this comment before posting it. So many obviously stolen ideas - I'm a fraud down to my very bones.)