Inside the university AI cheating crisis
theguardian.com
theguardian.com
Even before the 'AI detection' tool (IMHO snake-oil) the features around copy-pasting/plagiarism were so bad I always ignored them. The tool would flag things like commonly used coverpage forms. That universities pay a fortune for those tools is an indictment of the industry.
I was surprised at the level of effort
It had various tools and heuristicsv to suggest different phrasings and better words that were pretty incredible if compared to anything else back then
Problem comes when users do not care and blindly just believe in percentage value and some threshold
Maybe this is okay when we're talking about plagiarism and you can review manually, but humans can't reliably identify LLM written text
For her literature classes - where the class grade is composed of a participation grade and a written work grade - she is strongly considering getting rid of all written homework. Instead all writing work will be hand written in class.
The job of essays isn't to produce a Word document of text that can get a grade above 70 - it's to demonstrate and improve the student's writing and thinking. The computer seems to have now completely messed this up.
When you attach a literal degree to something being done "right" and not "demonstrated considerable improvement" you're going to get people that game the system. When you then make me pay $30,000 or more for a piece of paper you will get people that seek an edge. I did my fair share of underhanded tricks in school because I didnt want to pay another $1200 because I had to maintain a certain GPA to get where I wanted. This was over 15 years ago. Focus in the classes you care about, game the classes you don't. Its a tale as old as education itself.
Blame the OP's wife and the faculty and most universities for not only supporting but encouraging this behavior. If universities are for learning and improving the weight of grades shouldn't be career-determining and the cost should be commensurate to the expectation - both from the professor and the students.
Essays are only ever used in an academic context. No-one outside of academia writes an essay. This use-case has been replaced by Powerpoint presentations.
Essays are obviously not the only method of "demonstrating and improving the student's writing and thinking".
If we were to write an essay, we would use an LLM to do the actual writing, and feed it a set of bullet points for the points we want to make. Same as our parents moved from handwriting to word processing, we're moving to LLMs as a writing tool.
We're also seeing (that the article didn't mention) that LLMs are being used by staff to grade student essays. Getting to that ridiculous point of an LLM writing content that only an LLM will ever read.
So why not just abandon the essay as a teaching tool? I realise that academia is slow to change, but this might force them to.
Replace it with student presentations, or whatever the actual industry the students are heading for uses to communicate.
As for academia moving away from publishing papers (which are also increasingly being written by LLMs) into journals... that will need to happen too. What replaces that is going to be interesting.
This is patently false, with many counter examples to count. Editorials, blogs, substacks entries, long posts on twitter are all essays.
"Essays are only ever used in an academic context. No-one outside of academia writes an essay. This use-case has been replaced by Powerpoint presentations."
ooh-a for the new way to write outrageous essays so to engage people.
The bulk of assignments are useless busywork using outdated formats, and educators and educational institutions haven't been motivated to change it.
Instead of realizing this and embracing the theories and models of newer education sciences, most institutions have become near ideologues with their inertia.
True education doesn't need busywork but credentialism does. The busywork is the filter for the credential.
Also, if a person doesn't write essays in college how are we going to have all these authors that turn 20 pages of ideas into 300 page books? It would ruin the whole book industry. It takes a ton of practice to uncompress an idea into a coherent structure using 10X more words than needed so the book is the optimal length for sale when printed.
You could, of course, pay 30x more and get better service.
And then there's the total disrespect students have for just outright knowledge and practice. Being able to recite a math formularium from memory "is useless". It kind of is. But students who can do this work 5x faster than people who can't, and surely that's a good part of an education. The second thing that really matters for math is doing A LOT of exercises (like 10 intermediate to hard problems per day, ideally) over a prolonged period of time (months, minimum). That's every last exercise in a big calculus book. The exercises "are useless". Mathematica can do it, and in almost all cases, can do it better than a 40-year experience mathematician. But ...
The difference between students who've done both and students who haven't is night and day. Including the difference in using tools like Mathematica. Frankly, you can even see the difference clearly in other subjects.
The lack of ethics, like all selfishness, propagates inefficiencies into the greater society, as exemplified here for our overworked and underpaid teachers.
Our son has become quite a good chess player over the past 6ish years, during which online chess cheating via engines has proliferated. I explained to him that he is completely responsible for being able to honestly say, "No." in answer to the question, "Have you ever cheated in online (or over-the-board) chess?".
Having a decision matrix that leads to ethical action confers a specific peace that accompanies knowing that one has done nothing wrong. In a world full of liars and cheats, being honest is its own very special and quite rare power, looks to me.
We must utilize and orient our moral compass consciously in accordance with universal compassion, or we are susceptible to the ravings of the powerful, who, as you accurately said, tend to be callously cruel liars who crave only money, power, and pleasure, at the cost of anyone and everyone.
Peace be with you.
"The Way goes in." --Rumi
I know it sounds like a farce. Because it is. But it might also be a proper solution.
All this nannyware of various sorts is just more evidence, proof even, of how few of those students should even be there.
This isn't just essays, AI will happily output any known algorithm you ask for in a few seconds. CS coursework can be almost entirely automated in many cases
Here are a couple references in the article, but I would like to see a review of this question:
> Researchers at the University of Reading recently conducted a blind test in which ChatGPT-written answers were submitted through the university’s own examination system: 94% of the AI submissions went undetected and received higher scores than those submitted by the humans.
> One study at Stanford found that a number of AI detectors have a bias towards non-English speakers, flagging their work 61% of the time, as opposed to 5% of native English speakers (Turnitin was not part of this particular study).
[1] That needs to be defined: Every word? mostly? GPT-written then edited? What about ideas from the humans and writing by GPT? First draft by human and editing by GPT - what about using it as a grammar checker? For suggestions on clarifying some tricky passages? ?
Apart some hidden secret sauce by the LLM provider, detection is always going to be an arms race.
If you mean, is it feasible to empirically test the accuracy of detection software when the LLM output is so varied? That's an interesting question.
Is there a way to create a representative sample of LLM output for testing purposes? That's an important question, not least because we'd also be able to define the distribution and range of outputs (unless there's a statistical way to create a sample without knowing those things?).
It goes to the question of how deterministic is the output? How predictable, even statistically?
Maybe we just work with a large collection of output. Or maybe we can create a sample of inputs (the prompts), and work from that.
If you know something you can prompt to insert some errors into logic or different perspective. Effectively making it somewhat your work too
I don't think most common LLMs are deterministic. Especially if you work on the prompt. That is not just ask Write a synopsis of Romeo and Juliet.
Out of curiosity I tried following
> Write a synopsis of Romeo and Juliet > Rewrite it in style of a teenager, insert some slight grammar errors > Write in as if teenager was trying to sound formal > more formal > make a different take, more edgy and judgemental
I have no idea how one would combat that.
> Maybe we just work with a large collection of output.
yes, But how do you get that large collection from a single student who uses non-default prompts?
Though I wonder, are we not fighting kids using calculators in a math class?
People don't do bad things because that's their goal but because of these issues. If you give humans an easy way to do bad things, many people under pressure will do it.
Group papers are usually made into an outline at first and then we divvy up the responsibilities. If you use AI, most kids only care if your work sucks. AI consistently scores like an 80-85% (profs sometimes submit and blindly grade the responses), but almost always misses the core teaching points of classwork.
In my program, grades don’t matter (so employers can’t stack rank us). People who make extensive use of AI are really only cheating themselves. If you’re a big AI user, other kids generally know and try to avoid forming groups with you if they know in advance. You learn better when learning alongside others and if all someone does is dump AI slop in the google doc, you’re wasting my time in addition to yours.
I use AI to flesh out points, especially on assignments I don’t care about. It can help for idea generation and “connecting the dots” between ideas, but I always edit the output because the AI makes stylistic choices I don’t like. It’s definitely an accelerator for me when writing papers. I stand behind all the papers I’ve submitted, though some have sucked (regardless of AI usage or not).
In undergrad, when my priorities were less about learning and more towards dating/partying, I definitely would have abused this tool. At the end of the day, using the tool mainly cheats your own learning. I hope they transition to talking about using LLMs like one uses gambling - a little is fine here and there but if it’s all you do that’s a problem.
I’m not sure what to do about elementary age kids, because the AI easily writes “better” essays. At least in college I could do better than an AI if I applied myself. But in sixth grade? Good luck me. Cats out of the bag now and we should be really empathetic to the younger generations. Imagine getting slammed with TikTok->Pandemic->ChatGPT in the span of like 6-7 formative years. They are growing up differently and I certainly have no clue what we need to do to help them be successful.
If I hire someone to manage my money, I don't want them to do any gambling. Although given that investing is somewhat of a gamble, I at least want to set the terms and have them to disclose to me exactly how they're gambling.
Can you expand a bit on how that works? I have limited academic experience, so I’m fascinated. Does everyone end up with a 4.0 if they pass, or..?
Grades correspond to a GPA internally that may account for scholarship or whatnot. Mean raw score for some classes is like a 90+. Most grade differentiation comes from class participation. We can pull true GPA from the school if we dig, but we are explicitly told not to share GPA with employers. I’ll share my undergrad grades, but they hardly matter. As far as I can tell, no one shares their grades and employers are expected not to ask for them. This is a top 10 MBA program in the USA, I understand that other programs are very similar.
Therefore by doing as you were instructed and referencing to the article as instructed you would fail the task automatically.
I’ve been teaching at a university for more than 20 years, and only a few times have I given students a final exam. I used to assign final papers, but I stopped doing that in the spring of 2023, when ChatGPT was becoming widely known and I couldn’t decide how to deal with its use by students.
The big issue, in my opinion, is that AI can be used productively and ethically in education; it can be used to cheat; and there’s a huge gray area between those two extremes, where there doesn’t seem to be any consensus yet.
For example, suppose a student uses Claude to brainstorm topics, then chooses one from among those topics and researches it in depth, then does some more brainstorming with ChatGPT based on what he or she found, and then jots down what he or she wants to say as bullet points. Finally, Gemini is used to write a paper that presents the information in those bullet points in a logical and well-formed manner, and the student checks the paper and makes revisions before turning it in. Is that okay?
When I’ve discussed this issue with humanities faculty, they’ve generally regarded that kind of AI-assisted writing to be cheating. Science faculty, however, have been more receptive to it, as they care more about the accuracy and originality of the ideas than about how the words are strung together. My sample was small, though, and I am sure that there are different opinions on both sides.
I retired from my full-time university post in March 2023. Now that I’m teaching only one class a semester, I have each student do a final one-on-one presentation and interview with me, which makes up a large part of their grade. I’m able to do that because I’m teaching only that one class. If I had a full teaching load with a hundred students or more total, there’s no way I could interview each one individually.
Like that joke about how you write a summary, use AI to expand it, then the recipient uses AI to get a summary. It's no longer a joke - it's really happening.
I've also seen students use AI to understand material - instead of reading the material, they feed it through AI to explain it to them. This is either a mediocre student who'll have a hard time if they ever need to do original research, or it's a bad teacher. Either way they'll do OK on the exams.
Also our universities are failing us if the only thing we get at the end is a boolean piece of paper (or even one represented by a floating point number).
This is like servers at McDonald's complaining that the customers are cheating with trash cans because they aren't eating the meals they buy, just throwing them in the trash. Why would they do that?!?!? Because someone told them they need to go to McDonald's "for the experience"?
However, as scientific writer and conference/journal reviewer I once spotted a plagiarism using the crude detector (I believe it was Turnitin) embedded in one of the ACM portals we used to use for conferences. The “bad luck” of the submitter was they had plagiarised by copy paste 100% of one of my previous papers :’)
As I understand it, it records when you're writing, including all edits and such, and verifies it's human based on that. Well, see the demo at https://inktrail.co/in-action
It will probably work well right now, but I don't know how easy that would be to fool once the hucksters build tools to circumvent it.
Do you happen to know if there is some overlap between these professors and professors who also refuse AI-detectors? Apathy could be the reason but I wonder how much of it is driven by cynicism encouraged by the inefficacy of AI-detecting tools.
Ergo, if you have some students who use AI tools (and accept that those tools produce better papers and/or produce them with less effort) then those students will be ranked higher than students who study equally hard, are equally smart, and don't use AI tools.
Which then collapses simply down to more $ = better grades and higher class rank.
That's bullshit.
There have always been tutors, study groups, etc. that offer varying degrees of academic aid, but the broad accessibility of AI writing aids is novel.
This is absolutely the problem here. The work is done, the answers are right, the report is the best anyone has seen, but oh no a fancy new tool was involved so its CHEATING! This is so friggin stupid. When i get my work done I am responsible that it comes out right. When its right, thats just it. - it's right. No matter how i got there it's fine. But wait, what if i used some freelance platform to farm out my job to another person then turn in their work? Well that's breach of NDA and my terms of employment. So what, are the student over here committing wage theft of some kind even though they are the ones paying for tuition? Make this make sense to me, you can't. School is a job you pay to do so you can learn to get good at a job you get paid to do. Somehow school ends up being a bad investment and we get mad at the students? Schools are broken and useless and at this point i would prefer my doctor was great at using chatgpt and checking the results carefully over being taught at an educational institution.
The reason is because doing arithmetic by hand teaches kids how math actually works, while using a calculator just teaches them how to find the answer. This base-level understanding is crucial for understanding more advanced math. But it's also important for making mental estimations, or just having a basic grasp on how the world works.
This is why students using ChatGPT is a problem. If you never put in the actual work in the first place, you never gain any actual understanding. A doctor "great at using ChatGPT" won't know how to check the results carefully if they've relied on ChatGPT for their entire education. Having a baseline skill set is crucial for evaluating LLM output, and for handling situations that LLMs can't handle. And there will always be situations LLMs can't handle, regardless of how advanced genAI becomes.
I think kids waste way too much time doing hard arithmetic by hand. As a particularly egregious example, I visited a Montessori primary school where they demonstrated what is essentially a complex single player board game, for calculating square roots by hand. It's a process whereby a 5th grader would take ~10 minutes to calculate the root of a 3 digit number, by moving dozens of pins around. It looked so silly, and while other Montessori equipment offers geometrical intuition, this one seemed to offer none. Squaring with trial and error would have been so much easier.
Other more typical examples are of kids spending hundreds of hours practicing long multiplication and long division of e.g. 4 digits, and then getting punished and stressed out about having made silly mistakes. While in the real world, no one in their right mind would choose to perform these calculations by hand any more, and would instead use a tool incapable of making arithmetic mistakes.
I don't have a clear solution, but would tend towards an educational system that teaches only the very basics of performing each mathematical operation (probably up to the level the average literate adult would do mentally), and then focus the rest of the time on problems where the students are the ones posing the question, using a calculator/computer to solve it, and then check whether the answer makes sense in the original context.
EDIT: For those curious about it, I just found a blog post demonstrating the logic behind this square root peg board. The author of the blog argues that it does help build understanding, but from what I saw, I would politely disagree.
https://borisreitman.medium.com/square-roots-the-montessori-...
But I still think it's important that kids are taught how to do arithmetic by hand. And part of that will involve repetition, both to help reinforce and practice skills, and so educators can more accurately gauge a student's ability. You can't just bypass the need to sit down and do the work, even if the exact amount of work necessary is debatable.
I also think this argument means that theoretically someone with a 6th grade education and a 12th grade education should have the same economic output, as they both know how to read and add/multiply numbers. In real life the reason someone drops out before/during high school is always extenuating life circumstances, so the outcomes are much worse for them. But in a world where that didn't happen, do you think the outcomes would be essentially identical?
Perhaps this sounds silly, but I think that’s actually deeply interesting. There is a profoundly deep relationship between calculation and games. This applies not just to ordinary arithmetic, but also things like game theoretic proofs to prove the soundness of logical inference rules. Exposing children to that relation, even with a toy case, sounds like it could be fruitful.
> A doctor "great at using ChatGPT" won't know how to check the results carefully if they've relied on ChatGPT for their entire education.
They didn't rely on ChatGPT for their education; they relied on ChatGPT for finishing homework. If they go over the generated essay and make sure it doesn't have any inaccuracies and they consistently do so correctly, what's the problem?
That's why pencil and paper tests were worth 80% of the grade.
Students are typically agreeing to submit original work. It's a breach of the terms they agreed to, so I'm not sure it's so different.
> i would prefer my doctor was great at using chatgpt and checking the results carefully over being taught at an educational institution.
What is checking over the results carefully going to do if your doctor has never been taught medicine?