Not to mention the fact that being called non-human is most definitely going to offend some people.
What exactly makes anyone think that they can detect an LLM that is outputting text? The notion seems absurd yet it keeps coming up.
Not to mention the fact that being called non-human is most definitely going to offend some people.
What exactly makes anyone think that they can detect an LLM that is outputting text? The notion seems absurd yet it keeps coming up.
When you read LLM output, you can often tell. The source of the notion is that we can do it pretty well ourselves, so if the AIs are so magical then they should be able to do it too (not saying I agree, but it is a pretty clear line of logic).
At best - you can look for logical fallacies and false facts in the output, and use that to guess, but realistically - people are fucking bad at detecting it, including all these HN posters who keep claiming they can do it reliably...
Some humans may just write in that style. Why not?
Sadly many students never move beyond this introductory framework, so their professional writing (which GPT trained on) ends up with the same uninspired style. It's a way to automate the creation of content that even the "lowest common denominator" of students can manage, but it doesn't produce good writing.
Edited LLM text is much harder: using a LLM to generate some text and then edit it in shape (often by removing extraneous paragraphs, maybe rewriting a few things slightly). Those are basically impossible to detect reliably.
There’s definitely a pattern for some prompts, where the model uses an obvious format of:
- Make generalized statement answering prompt question
- Support that statement with 3 or 4 discrete paragraphs that don’t exhibit any personal experience, give examples, or cite statistics
- Finish off with a high-school essay style conclusion statement reiterating the introduction with different phrasing, generally beginning with “Overall” or “In conclusion”
It’s also comically obvious when suddenly someone who has a reply history full of low-effort juvenile, zero grammar, curse laden posts on video game and meme subs is suddenly writing mini-theses on topics ranging from Swedish forestry management to the chemistry of dyes used in Cambodian textile manufacturing.But it's not very good and, at a minimum, lacks nuance and supporting evidence.
You don't see overuse of linking words because it was generated by LLM. You see them, because every non-native spear is literally taught to link every paragraph with them. And all the texts, blog posts, wikipedia articles, stackoverflow responses and all the other stuff they wrote was then used as training data, from where LLM had learnt to do the same.
What I am saying is that there are just much more non-native English speakers and LLMs are inherently kind of non-native speakers too. So a sign that distincts majority (non-native) speakers from the minority (native speakers) is actually a bad sign (:
Especially as a teacher, you get used to seeing the, uh, quality of work that gets produced. I mean that for college teachers too. At my decent state college, the quality of essay writing was atrocious and I was given 99/100 for basically "not needing any help". In college!
When you compare the average GPT output to what many students turn in at all levels, the difference is highly apparent.
This also opens up the reality that smart cheaters will prompt engineer. I just tried on GPT4
> You are an average 10th grade student. You read and write at a 9th grade level. You don't edit your work ever and your writing has run-ons, comma splices, misuse of grammar, misspellings, etc. Say OK if you understand
> Write a 5 paragraph essay about spring bird patterns in Alabama. Use an intro paragraph and a conclusion paragraph. Keep it around 500 words.
> You are doing a good job of being a 10th grade student, however I want you to do the previous command like a 10th grade student who has been trained in basic essay writing and does not write conversationally.
And it's already producing something that is far more believable lol.
For the record: these tools are open and available and it takes about 10 seconds for you to replicate my experiment. The value you derive from my comment should be the idea I give you, not the results I generated. My results are good but imperfect, and are the result of very little effort. Someone truly trying to cheat the system could produce way more I imagine.
The results you generated are as important as the idea because ChatGPT does not output the same thing every time. Sometimes the output is very useful, other times it is complete and utter drek.
Method without results from that method are fairly meaningless if you're trying to compare.
There was a time when I had to discuss a paper that was too well written with a student. The paper was too well written in both my eyes and the eyes of others. The student told me which journals they read and explained they read it for both content and style. Were they cheating? Perhaps. On the other hand, I have never encountered a cheat who cared enough to make it look like they cared about how they approach learning.
> We did well with this skills as we apply it in corporate world plagiarising collegues works and even competitor a product. We even glorified how Samsung did it in early days of smartphone war era.
Also, this conflates factors other than cheating, such as the respect respondents have for the people administering the test, the degree to which respondents treat this whole thing seriously, etc.
The fact that the study says what we'd expect, doesn't imply it's a good study.
british guy publishes study which finds british people are the most honest -> definitely true, british people never lie, even though psychology is the most outrageously and rampantly un-reproduceable field of study
Any institution that produces this outcome has its priorities completely out of whack. And any student who willingly and energetically participates in such a system must be completely vapid, short-sighted and materialistic.
It would be a real shame for your country if this kind of behavior were common throughout. It would surely result in a situation where the most powerful and highly esteemed are without integrity.
Cheating is the easy way to get away from that situation. And if the student is behind, it's the only way.
It's not great by any means, obviously more people interested in educating themselves is better. At the same time, it shouldn't be a mystery. It, particularly, shouldn't be something one assumes is related to a particular country/culture rather than an everywhere thing. That's not to say there won't be variance between location, rather that it's a pretty common outlook everywhere.
Certain domains, such as budgeting, use it as a synonym for "allocate" but given the context it was first used in your comment it definitely implied "taking without consent".
Taking this approach, many might find school work is, largely, effort done to grade a level of understanding the individual has of the content. On the other hand, one might find business work is, largely, effort done to accomplish a customer's request. Additional less common types of each may come up as well, and there are, of course, exceptions to the main theme of each. In general though, I think you'd agree swapping the goals to say "school work is effort done to grade whether someone understands the content while business work is effort done to see if the particular person you make the request to is able to complete it on their own" seems rather unlikely in comparison. Because of this, swapping who does the work in each case results in different treatments. Not because it's a question of whether work gets done rather what the goal in doing the work is.
No, the assumption is that the purpose of education is to prepare you to be able to function in the real world. IMHO learning how to get other people to do your work for you is the single most valuable skill one could possibly acquire.
Even if we took this line of thought as the truth though, also came to an agreement this is the most valuable skill of all to teach, and also took it as being universally good to do regardless of context, what's the reasoning for assuming college only seeks to teach and grade success on a single skill in the first place?
Yes, and this is exactly the problem, because in point of actual fact most people who go to college do eventually leave. And if you think about this even for a moment, it has to be that way because someone has to do the actual work required to maintain the the civilization that makes college possible in the first place.
Liberal economics is based on several freedoms. To sell one’s labour, to buy the labour of others, to own property and to freely associate with others on economic activities. Put together, and with a fair legal system to regulate it, and that’s capitalism. Which of those freedoms do you disagree with and why?
That's true. But, with the possible exception of the bread, you did not pay the people who did the work, you paid someone who appropriated their work and sold it to you as their own.
I’m reading definitions of appropriation, and I don’t see the applicability. The bit about ‘without permission’ doesn’t seem to apply. I’m not taking anything, they offered their labour for sale.
Can you describe how a civilisation might function without appropriation as you define it?
The problem here is the incentive structure. The nature of grades is that some people will get As and some people will get Cs. If you get Cs, you might like to get As, but maybe you can't for some reason, from lack of intelligence to other time commitments to alcoholism.
Then your choice is to get a C on your own or maybe get a B+ by cheating. Given the same amount of learning, getting a better grade is better, so the only real incentive not to do it is the risk of getting caught. Some of the old methods for this were pretty unambiguous. If you submit the exact same essay as another student, or one that has been on the internet for ten years, what other explanation is there?
Rich kids would avoid this by paying someone else to do it for them. That has always been a problem. But now the cost of getting "someone else" to do it is approaching zero when the someone else is AI, so the problem spreads. But attempting to detect it with methods that have false positives is worthless.
On the other hand, testing students on their ability to do something that ChatGPT can do just as well? Maybe that's worthless too, because that's no longer a marketable skill when your future boss can get ChatGPT to do it too. So what they need to do is change the test to test for the thing the student is expected to be able to do better than the AI.
Do such expectations still exist? Are they expected to remain for very long?
Have you taken some time to play with these things? Try the biggest LLaMA model that will fit in RAM on your computer. (128GB of DDR4 is around $200 and will just fit the big one without quantizing, though it won't be super fast.)
There are things they're good at. Search engine-like tasks in particular, if you're willing verify the output. They're great at providing hints for further reading.
Now try to get it to develop a new kind of battery with a longer service lifetime or lower manufacturing cost per unit energy storage. Ask it to write code to do something complex and uncommon instead of something similar to what it was trained on a thousand examples of. Have it describe a new class of security concern, like Spectre or rainbow tables before they were known.
People can do those things, and have done, and those are some of the best things we can teach people how to do because they're incredibly useful. Maybe writing minor variants on common existing boilerplate code isn't something we need people to spend a lot of time on anymore, and so isn't the thing we should be testing if they know how to do.
The lecturer told us very clearly "DO NOT LOOK AT PREVIOUS YEARS ASSIGNMENTS". I was a little naughty, I asked a friend who did it the year before if I could have her group's assignment to read, because they'd got near-perfect marks. I didn't copy anything from it, I just read it to get a better idea of what was expected.
One guy in our group, he was really late to contribute his section of the assignment. When he finally gave it to me to read, I realised he'd just copied my friend's assignment from the previous year – either she'd given it to him also, or someone else in her group had. He tried to cover up his copying by rewording sentences, but it was very obvious – every paragraph made the same points in the same order, using the same (or very similar) word choices. Added to that, a lot of his attempts to reword it, the end result didn't even make much sense – some sentences, all he'd done was transform good grammar into bad.
I was angry and felt like reporting him. But, the only way I knew he was cheating was because I'd looked at a previous group's assignment, against explicit instructions not to – I couldn't see a way to get him in trouble without getting myself in trouble. In hindsight, I probably could have shown it to my friend and then got her to claim she'd spotted the plagiarism instead of me, but I didn't think of that at the time. Instead I just threw it out and redid it from scratch myself. He got the same marks as everyone else on the group assignment, without having contributed anything except a poor attempt at plagiarism.
Worse - The entire point of the LLM is that it's generating its output by picking the statistically likely next word... So from a "analyse text and detect forgery" side of things... any company that claims they can do it without false positives is fucking lying through their teeth.
When you can tell, you can tell. When you can't tell, you don't even know you can't tell.
Average ChatGPT output for the most part reminds me of the average British politician who can waffle on until you stop them, talking about all the aspects of something, but never really answering your question with all that much substance and depth, sometimes not even answering it whatsoever. Very "30,000 foot view".
Good writing only comes from writing lots of things that other people actually want to read. Course work is poorly structured for that, it isn’t peer review.
And whilst it's bad at being nuanced, it's even worse at being opinionated because its guide rails and human testers alike love its answers to be qualified with generic caveats like "depends on the specific situation"...
If you ask me, this is very "con-artisty".
A ChatGPT answer looks exactly like a politician that doesn't want to tell you the answer to your question.
It also has a "high schooler" style, but mostly because of the rigid form it uses. If actual high schoolers throw random content in their essays like ChatGPT does, they will get negative points for that.
Having a diary is one thing - having creative writing that is in the open and subject to criticism and encouraging refinement is quite another.
Well, that notion is wrong, otherwise, what's the explanation for something like this:
https://www.theregister.com/2023/05/17/university_chatgpt_gr...
For example, ask it to write in the style of Christopher Hitchens, Charles Bukowski, or Hunter S Thompson, let alone more extreme examples like Shakespeare or Dante
Probably the fact that if they admit the reality, they have to think about some difficult and profound questions. It's much easier just to posit an imaginary future technology and decide that will solve it.
My sense of the general idea (non-authorative): Since the sequence emitted by an LLM is probabilistic completion i.e. predict the next word, the examiner can also do the same by progressively processing the text. Given the assumption that the semantic relations extracted from training corpus should be fairly universal for a given domain at the output level (even though distinct LLMs will likely have distinct embedding spaces), then the examiner LLM should be able to assign probabilities to the predicted words. The idea is that a genuine human produced text will have idiosyncrasies that are -not- probabilistically optimal and the examiner can establish a sort of 'distant from probable mean' measure, with the expectation that LLM produced text should be 'closer' to the examiner's predictions of 'the next word'.
The problem (if above is correct) then is the missing 'prompt' and meta-instruction embedded therein. Those should ("engineering") affect the output, possibly skewing the distance measure, thus defeating the examiner. But of course, say in context of academia, the examiner can 'guess' as to some aspects of the prompt as well. For example, if you are examining papers for a specific assignment, the examiner can self-prompt as well. "An essay on Hume's position on the knowledge of the self".
Temperature above 1 often results in nonsense though
Can you come up with any type of system that this does not apply to?
https://www.lesswrong.com/posts/G5eMM3Wp3hbCuKKPE/proving-to...
The point being made is that the idea of detecting AI text is absurd. It won't work it can't work. The point about it having an increased negative impact on those who are already at a disadvantage is important.
Smart people don't need additional barriers to becoming benefits to society and promoting any tool that creates those barriers and does literally nothing else is terrible. And no, not all systems meet this criteria.
I think it's enough to say that it harms some people (ESL) more than others and leave it at that.
Are there any such systems? Don't all of them have the same problem of the very people who most need the system being the least capable of using it?
I think one solution would be a word processor that records the process of writing a paper and you need to turn in your paper along with your recording. Of course this is going to create added stress but what else do we do? There's going to be GPT scramblers that remove watermarks.
[1] https://www.businessinsider.com/professor-fails-students-aft...
Ideally though, mind probes should come into play to find out if they really produced the work.
You could probably do something similar with recorded sessions like what GP suggested. Even someone doing what you suggested could leave behind a distinct profile in the shape of the session.
> [This looks confusing Gus/]
One of my coworkers answered:
> [This looks confusing Gus/ Yes, let's rewrite it Pablo/]
We never changed it, and we forgot to remove the comment, so it got printed and distributed. :(
It was a small informal class, and the comment was not offensive or too bad, so it was not a big deal.
I'd be very scarred of people reading all my edit history.
I think your solution works but the idea of even less privacy, especially around thought processing, is worse than the problem solved. I am guessing this will be the solution and it will send the data to Google and or Microsoft. Terrible.
People always could get (paid) help on assignments. Now AIs just leveled the field a bit.
And you can always have a person defend his work. It shows quickly, if it wasn't written by him/her.
Why would I want to test someone's performance at say writing an email without the benefit of things like a spell checker, or grammarly style stuff, or chatgpt. They'll have access to this in the real world.
How about we test realistic scenarios. Instead of asking someone to calculate 96*451 by hand, ask them to workout the difference between buying a $700 phone and $5/month sim contract vs a $30/month contract for 3 years with RPI inflation. Give them a pop quiz to choose which of a dozen screaming offers in the supermarket aisle is the best thing with 10 seconds to decide, for added realism throw in a couple of kids causing some distractions.
For added credit have them consider opportunity costs of that $700 upfront payments, not just what they'd get with it sitting in a bank account. Instead of have them write a 1500 word essay by hand, give them 2 hours to write 1500 words explaining the benefits of X using normal tools like a modern word processor, google, wikipedia, ChatGPT etc. If they don't use the tools at all, you'd probably need to mark them down.
This is a false dichotomy. I feel like this well-worn talking point might be outliving it’s usefulness even faster with the invention of LLMs. We teach the mechanics of arithmetic and of spelling and grammar because that’s what education is for, it’s for teaching how things work so there’s an understanding of the fundamental mechanics, and it starts with the basics and builds on top of them to more advanced topics, in order to deliver a well-rounded and deep understanding. Note that the logical extension of your argument is to let ChatGPT, or the next AI, or the one after that, to start explaining your examples, and let humans ignore everything the computer can do, which now includes writing 1500 word essays. After all, just like spell checkers and calculators, access to AI is what people have in the real world.
We don’t need to choose between teaching multiplication by hand and how interest works, because we already teach both, one in a basic arithmetic class, and the other later in algebra & calculus. Same for spelling vs grammar vs writing. Spelling and grammar happen in elementary school, and essay writing happens in high school and college. Teachers already do allow calculators in algebra and calculus. This has no bearing on whether we should allow calculators in arithmetic class based on the vague notion that having access to calculators at all times is a ‘realistic scenario’.
I feel like this kind of thinking is what is leading people to try to cheat in the first place, it’s a lack of understanding of the value in basic skills, and the misguided assumption that learning something only has value if you can demonstrate you would need to use it all day every day for a job right now. The problem is it doesn’t ever get better or easier if you skip the basic mechanics that computers can do; in fact it only gets harder to learn the subjects that do actually matter that you are going to be using, when you have foundational gaps.
More than ever, what we are going to need from here on out of the education system is people who know at least a little what AI is on the inside just so we can use it effectively, not to mention build it, maintain it, control it, set legal policy, fix it when it’s wrong, etc. etc.
So students will white board their term papers and do other absurdities totally removed from any actual real world job requirements. Because in the real world everyone will just use gpt.
What I can see in the humanities at least is that evaluation will soon get away from subjective and superficial criteria of form and focus on the quality of the content again. (That's me being optimistic. The other alternative is that AI will soon evaluate the output of other AIs.)
Publish or perish!
Watermarking.
On each word of the output, you randomly split all possible words into two groups and only generate output using one of them. If you get a text of 1000 words that exactly follow the secret sequence of the groups, you can be sure that it's generated by this LLM with one in 2^1000 chance of error.
Watermarking only works if every LLM system available does it and it is impractical for third parties to spin up and use systems that don’t, otherwise it only detects content that opts in to being detected.
https://openai.com/blog/new-ai-classifier-for-indicating-ai-...