Plagiarism detectors are a crutch, and a problem
nature.com
nature.com
When I arrived, his first words to me "I think we both know why you're here." I didn't.
He explained that he had run my paper through a plagiarism detector, and it had come back as a 100% match with a paper turned into a university in another state. I was rather dumbfounded because I had researched and written the paper entirely on my own.
We quickly came to an impasse, because he had faith in this rather damning score from the plagiarism detector, and I obviously was quite sure I had in fact written the paper myself.
As I left the office and began gearing up for a battle with the Office of Academic Integrity, I got another email from him profusely apologizing and telling me he had run the paper through the same plagiarism detector again and it had come back clean.
I'm really glad I didn't have to fight that fight, but I got awful close. My theory is that there was some problem during the upload, and it just checked an empty document against another empty document and found they matched exactly, but I'll never know.
Later I was told I was disqualified because I had cheated and copied my answer from StackOverflow. I tried to explain that the user I copied it from was indeed me and that I could prove it by logging into SO in front of them but I never heard back.
Basically, it's not necessarily OK to plagiarize yourself. Always disclose when you do it.
I'll never understand the concept of plagiarizing long-established solutions to problems.
Not even different variable names, as any plagiarism detector which is halfway-competent would have to ignore them because changing variable names is one of the things any halfway-competent plagiarist would do. The arms race has gone to the point where it's impossible to detect mechanically whether an implementation of a standard algorithm has been plagiarized. It's probably impossible for a human to detect it without other context.
And that context comes down to "Has this student shown steady gains or are they suddenly handing in competent code after demonstrating complete incompetence?" which you can't know unless the class is structured such that students show what they know to the graders in person. Which requires a certain student-to-grader ratio. Which is expensive.
It's practically guaranteed. Quicksort is quicksort. You can't get creative with it without implementing something which isn't quicksort anymore. That code may well sort a list, but if the assignment is to implement quicksort, you've failed at the task.
> I think the real fail here is a company assigning something like quicksort and expecting to get original answers.
Right. You have to go to IOCCC levels of perversity to be able to turn in something which is quicksort but isn't going to trip a well-trained plagiarism detector.
Thankfully, I went through college before these automated plagiarism detectors had caught on. They were around, but not as pervasive as today. I mean, if you're going to sort a vector in C++ for a trivial app, there's only so many ways to write:
std::vector<int> stuff{...};
std::sort(std::begin(stuff), std::end(stuff));
And, lets face it: most lab assignments requiring a student to write a program are pretty small and trivial (at least mine were 20 years ago). The projects, on the other hand, were definitely more involved and you'd likely see more divergence in solutions there.I don't want to work with someone who trusts their own quicksort implementation over one from an OSS library. It's a toxic combination of optimism and egotism that I don't have time for. What other stupid shit are they doing all day long? What does it say about my team if I'm on the asking end? Interviews are short, and information exchange is always limited.
Instead I may ask them to sort a complex object, or cook a piece of data that's organized opposite of what the UI will require. lots of people fail at these tasks. They're my version of fizzbuzz, but they actually represent the work we do.
What does this mean?
I walked to the whiteboard and just wrote the implementation top to bottom without mistakes, then explained the code. I may have been a little lucky that I didn’t have any missteps, such as briefly forgetting to handle negatives or zero, or not immediately realizing that the first idea that pops into your head for pulling digits off will be a little awkward as you want to fill the string in the reverse order. But I’d never done itoa before so I didn’t think to offer any sort of disclaimer about why this was particularly easy, I was just happy with my performance and ready to do the next question.
The interviewers excused themselves and I saw them consulting the person who recommended me down the hall. I was surprised to learn that they were asking if it was possible I’d been supplied a list of interview questions and memorized the answers, to work out whether they even wanted to proceed from there.
We trust our systems way too much, when we should first be trusting our own minds, and other people.
Because of how basic the assignment was, and how rigid assembly is, there's a baseline ~40% match with all assignments, and mine just happened to hit ~50% and that was grounds for plagiarism. The match was for some submission 6 years prior.
At the time I took backups of school projects on save and could replay the history and walk through my design which was enough to convince the professor. There were at least 3 other students roped into this and were probably not as lucky.
They also biased _heavily_ in favour of exams (practicals are essentially something you had to do, but are ungraded), where there is no (reasonable) opportunity for plagiarism.
Each lab was typically booked as 2 3-hour sessions (one session per week). I lived off campus, and the commute was about an hour, in good traffic, so I wanted to minimize the number of times I had to go to campus.
Unlike the students that lived on campus, I was methodically prepared walking into the first session. I'd probably have spent 6-8 hours on a weekend working on the lab. I was meticulous. I designed the circuit using Visio (obviously can't simulate, but I was concerned with layout), so I had precise diagrams of the circuit.
My diagrams matched the layout of the bread board, and my physical implementation was clean (wires were always just as long as they needed to be and never longer). I could take the diagram and literally check off wires as I placed them (and my wiring was clean).
I had my 68K assembly written ahead of time and was generally bug-free, though some debugging was needed from time-to-time.
I was generally in and out of lab in 15 minutes of what was supposed to be a 6 hour session. After a few rounds of this, my TA suspected I might be cheating and he walked up, ripped out a bunch of random wires from my circuit and told me to fix it. I chuckled, said ok, and I pulled out my design, replaced the wires and checked them all off, plugged it back in and it passed. He was dumbfounded. He sort of had a "Wow, a student that's actually prepared for lab for once" kind of look on his face.
He never bothered me again.
I think the TA was doubly surprised because I also had a habit of obviously sleeping during lecture, dead center in the front row.
What you described sounds like it might come in handy some time for me personally
IntelliJ has a feature called "Local History" which stores basically is storing undo history, but it stores other things too (like test execution).
(I'm lucky enough to be old enough that this wasn't an issue. OTOH, we couldn't just Google or search Stack Overflow from source code examples...)
Clearly they hadn't read the Zen of Python.
Fast forward 6 to 12 months. I found what seems to be a descendant of the paper published in another journal. While the vast majority of the plagiarism I identified was removed, they added some parts plagiarized from my review. I contacted the journal, who replied that plagiarism detection software didn't find anything. I replied that it wouldn't because a confidential review would not be in their system. Plus, I gave them a list of plagiarized parts; all they had to do was check against the paper. I told them that I did not spend much time checking for plagiarism, so for all I know there's a lot more plagiarism in the paper. The journal representative never did anything about the paper, unfortunately. I now refuse to publish in that journal and discourage others from publishing there as well.
I wonder if you could submit DMCA requests for this.
That is, pardon my French, fucked up.
Beyond that it's egregious that any academic institution would allow the success of students/researchers/whatever to hinge on scores generated by a black box algorithm that hasn't been rigorously proven as very close to 100% accurate. There should always be a human in the loop, and if we don't have enough humans for that then I think we have a problem that goes beyond automated plagiarism checkers.
First of all, if your work is submitted through these detection systems, you almost never have the opportunity to see your 'score' and then resubmit. That's just asking for people to cheat the system.
These plagarism checkers are also not really smart enough to differentiate proper citation from actual plagarism. Thier function is to alert the person grading the papers to sections that might be problematic. It will highlight "plagarized" sections, but the professor will ignore it if it's properly quoted and attributed or if it's just a short phrase, etc.
I mean, yeah, there should be a real person in the loop. But the issues are arising from when these types of tools aren't being used properly.
It's not a problem when a professor is checking mid-terms from an upper-division class of 30. It is a problem when a grad-student with too little time is grading finals for a freshman class of 500.
Sure, the purveyors of these tools can say: "Use this thing properly! If you don't, bad things will happen." But time and time again, bad things keep happening when people don't use the tool right. After some amount of incidents, the purveyors of the tools should just re-design the dang thing [0] such that it's harder to use it wrong.
[0] Provided that the people that make the tool have any sense of morality and put human learning above dollars.
> That's way more hoops than even a semi-motivated plagiarizer is willing to jump throughThey won't show you the work, so copying the answer is not sufficient to get credit in practically any course.
"Academic integrity" has ridiculous rules (eg self-plagiarism is not really a thing that anyone should care about) and I'm sure some people would be pissed about that, but if you're doing the work I personally don't see anything wrong with getting immediate feedback on whether you've done it correctly or not. Particularly given that many textbooks do like one or two simple examples and then jump to complex problems in the assignments... being able to check a couple problems to be sure you understand the process is extremely helpful as a learning aid.)
If you want a real-world assessment of whether someone has mastered a concept, schedule an in-class quiz or a test.
Sure it is. The entire methodology of schooling is to learn through exercise; if you are asked to write a paper, it isn't because the teacher wants the information in the paper. They are wanting you to learn through the act of writing the paper.
If you just turn in a paper you have written for another class, you are not doing that exercise.
Turning in a paper that you wrote for a different class and getting upset that the school doesn't approve is like thinking that sending a picture of you lifting weights 3 years ago is going to make you stronger.
Take, for example, a class that is about learning to write effectively, which is clearly a valuable skill that adults need no matter what they do in life.
Part of that skill is being able to construct, from scratch, an effective piece of writing in a length of time that is longer than a proctored exam session. You can only measure that skill by having them do that, and if they turn in work they have written prior, you aren't learning anything about their ability to write something effectively on demand.
Research classes are similar; you can't demonstrate that you know how to do research well in a 3 hour exam. Programming classes, too; you can't measure the ability to create a large project in a 3 hour exam.
Doing a task your superior asks of you and doing it well is arguably the most important thing school teaches you, with time and stress management being a close second. This is what grades at school measures.
People love to talk about the virtue of being "well rounded" but the truth is that knowing bits of Shakespeare or knowing different types of Greek collumns doesn't benefit you as a person. All that information leaves your head pretty much immediately anyway. This is all especially true of high school. "Exams demonstrating skill" are not what matters to the people who look at your grades.
And you did, only earlier.
"is like thinking"
Funny. Do you also think that writing exactly the same paper many times is making you learn something each time?
They pass it the same text, it comes back as 100% plagiarized. Cool.
They pass it text that is close, but clearly just slightly modified original text. Comes back as plagiarized. Cool.
They pass it text that isn't close at all. Comes back as not plagiarized. Cool.
It works. Right?
Nope. We forgot to check if there are any cases where two different texts will show as being plagiarized despite knowing they've been independently created.
We also didn't check to see if a text that was essentially plagiarized can return a result of not plagiarized.
And I think those checks fall into the halting problem. You only need to find one case to throw the entire system into doubt. But to make sure there are no cases, you have to check every piece against every other piece.
And if you have false positives and false negatives, you can either double check every result. In which case, you might as well not use it. Or you give up on one the case of false negatives. If a plagiarized paper passes the system, oh well. Focus on checking every paper that gets tagged as plagiarized. Double check them and make sure it truly is plagiarized.
Metaphorically, if you mean that there isn't going to be a concrete metric that really determines whether two bits of text are "similar" with bright shining line, sure.
If you mean halting problem literally, no, this isn't affected by the halting problem. String similarity evaluation won't be strong enough to invoke the properties that result in a halting problem. For any two strings, you can certainly terminate in a reasonably period of time, and have the "correct" answer as defined by whatever the local meaning of "answer" is. (Barring a particularly perverse implementation of a string similarity algorithm probably explicitly constructed just to screw with my statement here, that for some reason starts executing the string as instructions directly or something.)
The end result may be totally convincing, but nobody really knows until they actually look at what they think they are citing. The teachers should operate at least at the lowest level they impose on their students -- Check your work. Don't blindly copy someone else's results.
The author uses a straw man argument that plagiarism detection software promises to "replace the professor" when in fact, it is just a tool.
Besides, there are highly sophisticated ways of cheating that even a professor would not catch reliably. I wrote my undergrad thesis about plagiarism and it was interesting to learn that there are students who are highly motivated to "beat the system" instead of doing the task properly.
[1] https://www.researchgate.net/publication/248606848_Support_f...
That said, I do still hear worrying comments from my colleagues (“what number is too much?” is I’m afraid a common question) - but I can clearly see the change in student behaviour - so I don’t think this is such a bad thing. Yes, markers do still need to watch for plagiarism - and there are of course other types of plagiarism that these tools can’t usually deal with (ghost-written essays being a common talking point these days), but just because a tool may be misused or may be misunderstood, it certainly doesn’t necessarily mean it is all bad.
Sure, the same quotes from another source may be used here and there (and hopefully cited). And there may be some rote definition or other short text lifted from Wikipedia, etc. that maybe doesn't rise to the level of requiring a footnote. But if any sections of text longer than a line or two aren't substantially different, something is going on.
If your course needs a big final paper, you don't grade the paper, you grade a defense of the paper, that shows they really know the material behind the paper.
All the work before the final is just practice and grades don't matter.
Keeping the algorithms secret for plagiarism detectors is about as dubious. If they are not good enough to reject gaming they are not robust enough to deliver reliable scores. If you don't trust students to see that gaming a plagiarism detector is more effort than honest work, you've admitted the wrong student, and/or you are using a duff plagiarism detector.
Teachers like them because they don't have to grade papers, textbook publishers like them because the online code prevents reselling the textbook (or at least, the new customer will need to purchase a key anyway).
I thought that, despite the title, that's what the author was talking about for the majority of this article, but then they mention things like...
> Only if a text is somehow off, and online searching does not help, should software systems be consulted.
Sorry, what? The only reason text would sound weird is if it's been specifically mangled to defeat plagiarism software (which makes plagiarism software already useful). There's practically no way you, as a student marker, busy professor, or someone reviewing hundreds of academic proposals has the time to slowly wade through each paper you get by googling choice sentences manually--something the plagiarism checker software can do just fine. I'm more-so confused, by the pair of statements that 1) "Software cannot determine plagiarism [...] That decision must be taken by a person" but also 2) "Only if a text is somehow off, and online searching does not help, should software systems be consulted". Isn't the author's whole point that plagiarism software false flags all the time? Isn't this, then, just "hey this sentence sounds funny, time to fail this student on some plagiarism."
If their argument, however, is that you should use plagiarism software for curatable results, then this seems like the opposite order of what should be done. Why waste all your time fruitlessly finessing Google if the software will straight up just find the OG source for you? You're not going to have read every single piece of writing conceivably connected to the essay/etc. you're reviewing (unless it's a field you know much about and also so narrow that you'd never consider using plagiarism software in the first place), so you're bound to be missing actual plagiarism all the time.
I don’t think these systems are nearly as egregious as the 5 paragraph essay scoring system they used around 2010 when I was in 9th grade back in 2009. I thought it was fine back then, but now I wonder how that thing had even the appearance of working some of the time.
I'd rather be the victim of a computer's bogus score than the victim of a computer's bogus accusation of wrong-doing. On the other hand it's important for everyone to experience the latter (by computer or not) at least once, and at least with plagiarism even if it's true you're not going to go to jail...
I'm not sure if the plagiarism service is just not very effective or if the fact that I turn in mostly LaTeX-generated PDFs is causing a problem with their text extraction. I actually suspect the latter, because a couple of times I've turned in MS Word files it's marked one or two random phrases as suspect.
[0] https://plagiarism.bloomfieldmedia.com/software/wcopyfind/
Today he'd just be expelled!
Is this the norm?
Haven't worked there in almost a year so I'm not sure if the product took off or not.
To be fair, they seem to be doing a good job and they are good a detecting copy/paste from online sources. But, of course, results are to be taken with a grain of salt.
I think current CS is doing itself a disservice because we want students to collaborate, but discourage them from even looking at each other's code (that is literally how students have spoken to me about it). I don't know if English or History students refuse to look over each other's work, but I doubt it.
Soloway described programmers as having "recurring basic patterns". If you accept his theories, we literally have tiny little rolodexes in our minds and when we encounter a problem we go "hang on, I've done this before" and pull out some code we've worked on in the past. We literally have tiny snippets of code we try to reuse if we can.
Instead, for well known algorithms, just give them the code! My research [1] require them to at least do it once, but either way - what good is it to have a student reinvent how A* search works? What about inserting into a Linked List? There are only a finite number of ways to implement these things. Instead of discouraging StackOverflow, and letting students try to sneak that type of work into their code, control it through the use of additional examples. They are less pressed to copy and paste code from a 3rd party if you've literally given it to them.
We want students to understand how to take something like a programming language and use it to solve a problem, but then we confound the entire process by not training them on the language. Even among SIGCSE, people can't agree what CS1 should be and somehow everyone wants it to be "more than just learning to program" like its a bad thing.
Back to plagiarism detectors for code, I think our current track discourages students to be open and help each other because they fear if they describe how they solved the exact same problem, there goes college. The result is encouraging the loner stereotype instead of reducing it. Hell, paired programming has to be forced on us because we can't "work with other people". I recognize some simply don't like it, however, my point is we've trained ourselves to not work together for so long that we designed a coding paradigm to have "controlled helping someone code".
[1] https://research.csc.ncsu.edu/arglab/projects/exercises.html
[2] https://www.sciencedirect.com/science/article/pii/B978093461...
Exams were in class, 2-3 hours long, and with a coding without Internet section. This approach does become more difficult with a 75 minute course (which is what NC State has), simply due to time constraints, so I can agree with you there.
However, your point about "Something has to give" is one of the reasons I argue that we expect too much out of CS1. In other technical skills, like cooking, music, martial arts, dance, etc., the introduction is much more about accumulating the technical skills necessary. Music theory isn't as prominent in an Intro to Singing class; however, in Intro to CS, I don't think we do that. As my course structure should indicate, I think intro courses should be more drilling to build the neuromuscular memory because then at higher levels we can do more critical thinking with the trust that students "know how to execute the plan". I do think that something like Problem Solving should be given its own separate course where its skills can be trained and honed parallel, but separate, from Programming.
"Throw it all out" doesn't seem like a real solution to me, more of a flashy headline.