GPT detectors are biased against non-native English writers
arxiv.org
arxiv.org
Not to mention the fact that being called non-human is most definitely going to offend some people.
What exactly makes anyone think that they can detect an LLM that is outputting text? The notion seems absurd yet it keeps coming up.
Probably the fact that if they admit the reality, they have to think about some difficult and profound questions. It's much easier just to posit an imaginary future technology and decide that will solve it.
I think one solution would be a word processor that records the process of writing a paper and you need to turn in your paper along with your recording. Of course this is going to create added stress but what else do we do? There's going to be GPT scramblers that remove watermarks.
[1] https://www.businessinsider.com/professor-fails-students-aft...
Ideally though, mind probes should come into play to find out if they really produced the work.
You could probably do something similar with recorded sessions like what GP suggested. Even someone doing what you suggested could leave behind a distinct profile in the shape of the session.
> [This looks confusing Gus/]
One of my coworkers answered:
> [This looks confusing Gus/ Yes, let's rewrite it Pablo/]
We never changed it, and we forgot to remove the comment, so it got printed and distributed. :(
It was a small informal class, and the comment was not offensive or too bad, so it was not a big deal.
I'd be very scarred of people reading all my edit history.
I think your solution works but the idea of even less privacy, especially around thought processing, is worse than the problem solved. I am guessing this will be the solution and it will send the data to Google and or Microsoft. Terrible.
People always could get (paid) help on assignments. Now AIs just leveled the field a bit.
And you can always have a person defend his work. It shows quickly, if it wasn't written by him/her.
Why would I want to test someone's performance at say writing an email without the benefit of things like a spell checker, or grammarly style stuff, or chatgpt. They'll have access to this in the real world.
How about we test realistic scenarios. Instead of asking someone to calculate 96*451 by hand, ask them to workout the difference between buying a $700 phone and $5/month sim contract vs a $30/month contract for 3 years with RPI inflation. Give them a pop quiz to choose which of a dozen screaming offers in the supermarket aisle is the best thing with 10 seconds to decide, for added realism throw in a couple of kids causing some distractions.
For added credit have them consider opportunity costs of that $700 upfront payments, not just what they'd get with it sitting in a bank account. Instead of have them write a 1500 word essay by hand, give them 2 hours to write 1500 words explaining the benefits of X using normal tools like a modern word processor, google, wikipedia, ChatGPT etc. If they don't use the tools at all, you'd probably need to mark them down.
This is a false dichotomy. I feel like this well-worn talking point might be outliving it’s usefulness even faster with the invention of LLMs. We teach the mechanics of arithmetic and of spelling and grammar because that’s what education is for, it’s for teaching how things work so there’s an understanding of the fundamental mechanics, and it starts with the basics and builds on top of them to more advanced topics, in order to deliver a well-rounded and deep understanding. Note that the logical extension of your argument is to let ChatGPT, or the next AI, or the one after that, to start explaining your examples, and let humans ignore everything the computer can do, which now includes writing 1500 word essays. After all, just like spell checkers and calculators, access to AI is what people have in the real world.
We don’t need to choose between teaching multiplication by hand and how interest works, because we already teach both, one in a basic arithmetic class, and the other later in algebra & calculus. Same for spelling vs grammar vs writing. Spelling and grammar happen in elementary school, and essay writing happens in high school and college. Teachers already do allow calculators in algebra and calculus. This has no bearing on whether we should allow calculators in arithmetic class based on the vague notion that having access to calculators at all times is a ‘realistic scenario’.
I feel like this kind of thinking is what is leading people to try to cheat in the first place, it’s a lack of understanding of the value in basic skills, and the misguided assumption that learning something only has value if you can demonstrate you would need to use it all day every day for a job right now. The problem is it doesn’t ever get better or easier if you skip the basic mechanics that computers can do; in fact it only gets harder to learn the subjects that do actually matter that you are going to be using, when you have foundational gaps.
More than ever, what we are going to need from here on out of the education system is people who know at least a little what AI is on the inside just so we can use it effectively, not to mention build it, maintain it, control it, set legal policy, fix it when it’s wrong, etc. etc.
So students will white board their term papers and do other absurdities totally removed from any actual real world job requirements. Because in the real world everyone will just use gpt.
What I can see in the humanities at least is that evaluation will soon get away from subjective and superficial criteria of form and focus on the quality of the content again. (That's me being optimistic. The other alternative is that AI will soon evaluate the output of other AIs.)
When you read LLM output, you can often tell. The source of the notion is that we can do it pretty well ourselves, so if the AIs are so magical then they should be able to do it too (not saying I agree, but it is a pretty clear line of logic).
At best - you can look for logical fallacies and false facts in the output, and use that to guess, but realistically - people are fucking bad at detecting it, including all these HN posters who keep claiming they can do it reliably...
Some humans may just write in that style. Why not?
Sadly many students never move beyond this introductory framework, so their professional writing (which GPT trained on) ends up with the same uninspired style. It's a way to automate the creation of content that even the "lowest common denominator" of students can manage, but it doesn't produce good writing.
Edited LLM text is much harder: using a LLM to generate some text and then edit it in shape (often by removing extraneous paragraphs, maybe rewriting a few things slightly). Those are basically impossible to detect reliably.
There’s definitely a pattern for some prompts, where the model uses an obvious format of:
- Make generalized statement answering prompt question
- Support that statement with 3 or 4 discrete paragraphs that don’t exhibit any personal experience, give examples, or cite statistics
- Finish off with a high-school essay style conclusion statement reiterating the introduction with different phrasing, generally beginning with “Overall” or “In conclusion”
It’s also comically obvious when suddenly someone who has a reply history full of low-effort juvenile, zero grammar, curse laden posts on video game and meme subs is suddenly writing mini-theses on topics ranging from Swedish forestry management to the chemistry of dyes used in Cambodian textile manufacturing.But it's not very good and, at a minimum, lacks nuance and supporting evidence.
You don't see overuse of linking words because it was generated by LLM. You see them, because every non-native spear is literally taught to link every paragraph with them. And all the texts, blog posts, wikipedia articles, stackoverflow responses and all the other stuff they wrote was then used as training data, from where LLM had learnt to do the same.
What I am saying is that there are just much more non-native English speakers and LLMs are inherently kind of non-native speakers too. So a sign that distincts majority (non-native) speakers from the minority (native speakers) is actually a bad sign (:
Especially as a teacher, you get used to seeing the, uh, quality of work that gets produced. I mean that for college teachers too. At my decent state college, the quality of essay writing was atrocious and I was given 99/100 for basically "not needing any help". In college!
When you compare the average GPT output to what many students turn in at all levels, the difference is highly apparent.
This also opens up the reality that smart cheaters will prompt engineer. I just tried on GPT4
> You are an average 10th grade student. You read and write at a 9th grade level. You don't edit your work ever and your writing has run-ons, comma splices, misuse of grammar, misspellings, etc. Say OK if you understand
> Write a 5 paragraph essay about spring bird patterns in Alabama. Use an intro paragraph and a conclusion paragraph. Keep it around 500 words.
> You are doing a good job of being a 10th grade student, however I want you to do the previous command like a 10th grade student who has been trained in basic essay writing and does not write conversationally.
And it's already producing something that is far more believable lol.
For the record: these tools are open and available and it takes about 10 seconds for you to replicate my experiment. The value you derive from my comment should be the idea I give you, not the results I generated. My results are good but imperfect, and are the result of very little effort. Someone truly trying to cheat the system could produce way more I imagine.
The results you generated are as important as the idea because ChatGPT does not output the same thing every time. Sometimes the output is very useful, other times it is complete and utter drek.
Method without results from that method are fairly meaningless if you're trying to compare.
There was a time when I had to discuss a paper that was too well written with a student. The paper was too well written in both my eyes and the eyes of others. The student told me which journals they read and explained they read it for both content and style. Were they cheating? Perhaps. On the other hand, I have never encountered a cheat who cared enough to make it look like they cared about how they approach learning.
> We did well with this skills as we apply it in corporate world plagiarising collegues works and even competitor a product. We even glorified how Samsung did it in early days of smartphone war era.
Also, this conflates factors other than cheating, such as the respect respondents have for the people administering the test, the degree to which respondents treat this whole thing seriously, etc.
The fact that the study says what we'd expect, doesn't imply it's a good study.
british guy publishes study which finds british people are the most honest -> definitely true, british people never lie, even though psychology is the most outrageously and rampantly un-reproduceable field of study
Any institution that produces this outcome has its priorities completely out of whack. And any student who willingly and energetically participates in such a system must be completely vapid, short-sighted and materialistic.
It would be a real shame for your country if this kind of behavior were common throughout. It would surely result in a situation where the most powerful and highly esteemed are without integrity.
Cheating is the easy way to get away from that situation. And if the student is behind, it's the only way.
It's not great by any means, obviously more people interested in educating themselves is better. At the same time, it shouldn't be a mystery. It, particularly, shouldn't be something one assumes is related to a particular country/culture rather than an everywhere thing. That's not to say there won't be variance between location, rather that it's a pretty common outlook everywhere.
Certain domains, such as budgeting, use it as a synonym for "allocate" but given the context it was first used in your comment it definitely implied "taking without consent".
Taking this approach, many might find school work is, largely, effort done to grade a level of understanding the individual has of the content. On the other hand, one might find business work is, largely, effort done to accomplish a customer's request. Additional less common types of each may come up as well, and there are, of course, exceptions to the main theme of each. In general though, I think you'd agree swapping the goals to say "school work is effort done to grade whether someone understands the content while business work is effort done to see if the particular person you make the request to is able to complete it on their own" seems rather unlikely in comparison. Because of this, swapping who does the work in each case results in different treatments. Not because it's a question of whether work gets done rather what the goal in doing the work is.
No, the assumption is that the purpose of education is to prepare you to be able to function in the real world. IMHO learning how to get other people to do your work for you is the single most valuable skill one could possibly acquire.
Even if we took this line of thought as the truth though, also came to an agreement this is the most valuable skill of all to teach, and also took it as being universally good to do regardless of context, what's the reasoning for assuming college only seeks to teach and grade success on a single skill in the first place?
Yes, and this is exactly the problem, because in point of actual fact most people who go to college do eventually leave. And if you think about this even for a moment, it has to be that way because someone has to do the actual work required to maintain the the civilization that makes college possible in the first place.
Liberal economics is based on several freedoms. To sell one’s labour, to buy the labour of others, to own property and to freely associate with others on economic activities. Put together, and with a fair legal system to regulate it, and that’s capitalism. Which of those freedoms do you disagree with and why?
That's true. But, with the possible exception of the bread, you did not pay the people who did the work, you paid someone who appropriated their work and sold it to you as their own.
I’m reading definitions of appropriation, and I don’t see the applicability. The bit about ‘without permission’ doesn’t seem to apply. I’m not taking anything, they offered their labour for sale.
Can you describe how a civilisation might function without appropriation as you define it?
The problem here is the incentive structure. The nature of grades is that some people will get As and some people will get Cs. If you get Cs, you might like to get As, but maybe you can't for some reason, from lack of intelligence to other time commitments to alcoholism.
Then your choice is to get a C on your own or maybe get a B+ by cheating. Given the same amount of learning, getting a better grade is better, so the only real incentive not to do it is the risk of getting caught. Some of the old methods for this were pretty unambiguous. If you submit the exact same essay as another student, or one that has been on the internet for ten years, what other explanation is there?
Rich kids would avoid this by paying someone else to do it for them. That has always been a problem. But now the cost of getting "someone else" to do it is approaching zero when the someone else is AI, so the problem spreads. But attempting to detect it with methods that have false positives is worthless.
On the other hand, testing students on their ability to do something that ChatGPT can do just as well? Maybe that's worthless too, because that's no longer a marketable skill when your future boss can get ChatGPT to do it too. So what they need to do is change the test to test for the thing the student is expected to be able to do better than the AI.
Do such expectations still exist? Are they expected to remain for very long?
Have you taken some time to play with these things? Try the biggest LLaMA model that will fit in RAM on your computer. (128GB of DDR4 is around $200 and will just fit the big one without quantizing, though it won't be super fast.)
There are things they're good at. Search engine-like tasks in particular, if you're willing verify the output. They're great at providing hints for further reading.
Now try to get it to develop a new kind of battery with a longer service lifetime or lower manufacturing cost per unit energy storage. Ask it to write code to do something complex and uncommon instead of something similar to what it was trained on a thousand examples of. Have it describe a new class of security concern, like Spectre or rainbow tables before they were known.
People can do those things, and have done, and those are some of the best things we can teach people how to do because they're incredibly useful. Maybe writing minor variants on common existing boilerplate code isn't something we need people to spend a lot of time on anymore, and so isn't the thing we should be testing if they know how to do.
The lecturer told us very clearly "DO NOT LOOK AT PREVIOUS YEARS ASSIGNMENTS". I was a little naughty, I asked a friend who did it the year before if I could have her group's assignment to read, because they'd got near-perfect marks. I didn't copy anything from it, I just read it to get a better idea of what was expected.
One guy in our group, he was really late to contribute his section of the assignment. When he finally gave it to me to read, I realised he'd just copied my friend's assignment from the previous year – either she'd given it to him also, or someone else in her group had. He tried to cover up his copying by rewording sentences, but it was very obvious – every paragraph made the same points in the same order, using the same (or very similar) word choices. Added to that, a lot of his attempts to reword it, the end result didn't even make much sense – some sentences, all he'd done was transform good grammar into bad.
I was angry and felt like reporting him. But, the only way I knew he was cheating was because I'd looked at a previous group's assignment, against explicit instructions not to – I couldn't see a way to get him in trouble without getting myself in trouble. In hindsight, I probably could have shown it to my friend and then got her to claim she'd spotted the plagiarism instead of me, but I didn't think of that at the time. Instead I just threw it out and redid it from scratch myself. He got the same marks as everyone else on the group assignment, without having contributed anything except a poor attempt at plagiarism.
Worse - The entire point of the LLM is that it's generating its output by picking the statistically likely next word... So from a "analyse text and detect forgery" side of things... any company that claims they can do it without false positives is fucking lying through their teeth.
When you can tell, you can tell. When you can't tell, you don't even know you can't tell.
Average ChatGPT output for the most part reminds me of the average British politician who can waffle on until you stop them, talking about all the aspects of something, but never really answering your question with all that much substance and depth, sometimes not even answering it whatsoever. Very "30,000 foot view".
Good writing only comes from writing lots of things that other people actually want to read. Course work is poorly structured for that, it isn’t peer review.
And whilst it's bad at being nuanced, it's even worse at being opinionated because its guide rails and human testers alike love its answers to be qualified with generic caveats like "depends on the specific situation"...
If you ask me, this is very "con-artisty".
A ChatGPT answer looks exactly like a politician that doesn't want to tell you the answer to your question.
It also has a "high schooler" style, but mostly because of the rigid form it uses. If actual high schoolers throw random content in their essays like ChatGPT does, they will get negative points for that.
Having a diary is one thing - having creative writing that is in the open and subject to criticism and encouraging refinement is quite another.
Well, that notion is wrong, otherwise, what's the explanation for something like this:
https://www.theregister.com/2023/05/17/university_chatgpt_gr...
For example, ask it to write in the style of Christopher Hitchens, Charles Bukowski, or Hunter S Thompson, let alone more extreme examples like Shakespeare or Dante
Publish or perish!
Can you come up with any type of system that this does not apply to?
https://www.lesswrong.com/posts/G5eMM3Wp3hbCuKKPE/proving-to...
The point being made is that the idea of detecting AI text is absurd. It won't work it can't work. The point about it having an increased negative impact on those who are already at a disadvantage is important.
Smart people don't need additional barriers to becoming benefits to society and promoting any tool that creates those barriers and does literally nothing else is terrible. And no, not all systems meet this criteria.
I think it's enough to say that it harms some people (ESL) more than others and leave it at that.
Are there any such systems? Don't all of them have the same problem of the very people who most need the system being the least capable of using it?
My sense of the general idea (non-authorative): Since the sequence emitted by an LLM is probabilistic completion i.e. predict the next word, the examiner can also do the same by progressively processing the text. Given the assumption that the semantic relations extracted from training corpus should be fairly universal for a given domain at the output level (even though distinct LLMs will likely have distinct embedding spaces), then the examiner LLM should be able to assign probabilities to the predicted words. The idea is that a genuine human produced text will have idiosyncrasies that are -not- probabilistically optimal and the examiner can establish a sort of 'distant from probable mean' measure, with the expectation that LLM produced text should be 'closer' to the examiner's predictions of 'the next word'.
The problem (if above is correct) then is the missing 'prompt' and meta-instruction embedded therein. Those should ("engineering") affect the output, possibly skewing the distance measure, thus defeating the examiner. But of course, say in context of academia, the examiner can 'guess' as to some aspects of the prompt as well. For example, if you are examining papers for a specific assignment, the examiner can self-prompt as well. "An essay on Hume's position on the knowledge of the self".
Temperature above 1 often results in nonsense though
https://openai.com/blog/new-ai-classifier-for-indicating-ai-...
Watermarking.
On each word of the output, you randomly split all possible words into two groups and only generate output using one of them. If you get a text of 1000 words that exactly follow the secret sequence of the groups, you can be sure that it's generated by this LLM with one in 2^1000 chance of error.
Watermarking only works if every LLM system available does it and it is impractical for third parties to spin up and use systems that don’t, otherwise it only detects content that opts in to being detected.
Difficulties I can think of:
* Getting ai companies to offer this. I don't think it comes with a downside for them really though. You couldn't use it to actually retrieve results in any way, and you wouldn't have prompts.
* This only detects exact matches, people changing the output would defeat it. Fixable with some kind of fuzzy search that returns a 'distance to nearest response' but this obviously makes it more expensive and difficult to run, use, etc.
* People could still run the model themselves, as models get better and more expensive maybe this becomes less of a problem. Or maybe models get small enough while still generating good output that it becomes more of a problem. Who knows.
At least this would avoid AI grading AI issues
You don't need to worry about deluge of AI content if it can only come from a few crazy people with top of the line hardware spending days iterating on prompts. Especially if really good GPUs start requiring a license. At some point it's just faster to learn to write properly.
Average consumer hardware is not showing any signs of being capable of this in near term. Hype is hype, but use your own head.
Don't forget that these homegrown models also require training data, scraped and cleaned, an individual or nonprofit can't do that.
https://www.mosaicml.com/blog/mpt-7b
There's dozens if not hundreds more individuals and nonprofits doing these things.
It may even be the case that all of that RLHF training that OpenAI does simply lessens the quality of generations, as suggested by one of their own papers and the paper above.
Quantized, LLaMa can run on an Pixel 6:
Who is going to pay $20/mo other than professionals? You'd assume professionals have professional hardware. A mid-range GPU or a video editing laptop is not exactly breaking the bank.
>no tool required to see that it is generated
But again, that also applies to commercial generative images. They're easily discernible, if you just look. Midjourney is stable diffusion with a bunch of LoRA stacked on top. And it can't shed the "midjourney look" because of that. That's not in dispute by anyone.
100%, although dall-e is getting better with hands specifically.
Anyway, Microsoft can keep throwing compute on it until a point where it becomes impossible to distinguish fakes by sight and a mechanism like suggested will make sense.
But with homegrown ones I don't see it happening soon. Only those who spend a lot of money on top of the line GPUs and keep desktop PCs may get to that point. Those GPUs will jump in price like they did at crypto mining peak or become impossible to buy if Microsoft gets the government to require "AI license" for them.
> Who is going to pay $20/mo other than professionals
Apple Music costs $10 and people easily spend ten times that on Patreon...
Well, yeah. It's been around for years and chatgpt is new. My years old GPU doesn't run the latest games either, and chatgpt is novel technology so the hardware will of course lag behind a bit. But it will come.
The cat is out of the bag. While the "AI alignment" jackasses were writing their Terminator fanfiction and wringing their hands about paperclips, they had already destroyed the world wide web as we know it.
These models you talk about are either not good or unbearably slow, the output of a model that today runs "as fast as you can type" on average hardware will never be reliably mistaken for the real thing. If you try to cheat with it it would be more likely to fail you than if you spend $20 on a human freelancer to write stuff
The only factor that breaks this today is cheap availability of chatgpt and such. They are reasonably high quality but unprofitable to run, they are subsidized to hook public up so that later MS can safely jack up prices (ideally after getting an exclusive AI license from the government).
That's not even taking into account local LoRAs and scripts that are possible instead of some company's untweakable crap. The open source around this is healthy, has pushed past DALL-E, and there's no real roadblock to Open Source LLMs except of course, the training cost. Even still, people are getting $200k+ models in their hands for free from various training runs and donated computing and LoRAing them and fine tuning them all to make them comparable to the closed off remote models.
Any "cryptographic" scheme with the generations of these will just catch the lazy. The lazy already include the confabulated sources in their papers, and don't try to normalize the Error Level Analysis in generated images (probably the quickest way to determine whether an image is generated), so I don't think it's actually a net benefit. It's a cat and mouse game, and will push the mice further into the walls.
You can't possibly say that generative images like this are "poor quality"
https://www.reddit.com/r/StableDiffusion/comments/131lpks/my...
The first of your examples was generated on a desktop computer with 2080 Ti and even then still glaring uncanny hands. We don't know how long it took but I think the reason for the hands is that it's too slow to generate a dozen of these in hopes that hands would come out right.
The other one I can see done on any laptop in a few minutes, but it's more primitive and just a monochrome sketch. I skip over obvious issues e.g. with shape of glasses.
For both examples you don't need any specialized tools or watermarking to notice this stuff.
Maybe you see what I mean why indie homegrown AI is not such a big deal ;) Sure there are people who will invest in hardware but those people will are not and for now won't be mainstream enough to matter. Especially if it will be licensed, most people don't like to violate laws. Most people will just use chatgpt or dall-e.
Commercial AI has all of those issues you mentioned, and more, and less. Midjourney is just a bunch of LoRAs layered on top and scripting to generate the images. But since they do that, midjourney images have a specific "feel" that it can't seem to get rid of. It's nothing really out of the reach for someone sufficiently motivated to reproduce.
DALL-E is laughable now, and it's only been a year. Certainly has been surpassed by open source, and outside competitors. I'm not sure what your motivation is to discount open source. People are already running LLM inference on their phones.
Your point?
Why do you need to watermark them, again? The error level analysis is off the charts with generative images. They light up like a Christmas tree. Just because you and uninformed legislators and journalists don't know how to check the ELA of an image, doesn't mean they're undetectable. And the cheaters include the bogus sources spit out by ChatGPT already. The cryptographic qualities will be lost as soon as an editor gets their hands on it, automated editing or not. It's a cat and mouse game.
And also, I find it telling that you think if someone doesn't have high end hardware, they're going to pay $20/mo to OpenAI. For $20/mo, you can buy a mid-range video card and write it off. For an extra $10/mo, you can deprecate the cost and buy a high end laptop for that price, if you're a professional, and you're not locked into OpenAI. You're also assuming that 1. hardware doesn't get better and 2. techniques don't improve to run them on limited hardware.
Again, either too slow, requiring outrageous hardware, or obviously noticeable. So far no examples to the contrary.
Don't forget, the topic is using special measures to detect undetectable with naked eye. When you can simply see the screwed up hands on a photo it's not even necessary.
> Why do you need to watermark them, again?
Why do you think I need to watermark them again?
> a mid-range video card
and a PC to put it in, a space to put the PC in, etc. With a laptop we're back in wait for an hour to see a result.
> hardware doesn't get better and 2. techniques don't improve to run them on limited hardware
We can revisit this if it consumer hardware gets good enough...
Any laptop within the last five years with decent memory can run stable diffusion on the cpu in around 12 minutes. My MacBook Pro runs a batch of four on Metal in around 30 seconds.
>We can revisit this if it consumer hardware gets good enough...
I mean, I just showed you a quantized llama running on a Pixel 5 and 6. And, I wouldn't discount most of the next generation of hardware having ML co processing like MacBooks and iPhones and Pixels do with all of this hype.
Majority of output is bad so you need to try dozens of takes to get a result that is reasonably realistic. Multipy 12 accordingly
> quantized llama
I don't know what that means but if it's better than chatgpt/gpt4 then sure.
It has most of the same problems you list, except it is much more robust against small changes to the text.
As of a few weeks ago, he mentioned that OpenAI had a working implementation and were discussing whether to start using it. I assume they'd tell people before they turn it on in prod, I see no advantage in secrecy.
Watermarking has zero-to-negative-value for the user of the generation service, but value for users of the detection service that the common vendor of both services will sell, so the only reason to announce that watermarking is active is because you are ready to sell the detection service leveraging it. Otherwise, its just a disincentive to some users to use the generation service with no upside.
Of course, many cheaters are that stupid and lazy. People still just copy and paste essays they found online.
https://www.youtube.com/watch?v=XZJc1p6RE78
Essentially by biasing the LLM's token choices slightly you can infer with high probability later that the text was generated by it.
Considering that AI writing standards will probably only increase from here, encroaching on the territory native English speakers operate at, perhaps the best way to mark yourself as not-AI is to start making many errors.
This will work for a week or two until someone tunes a model to also do this, and the arms race grinds ever on!
It also goes for other languages too, so luckily for me, I have a short period where my dreadful other languages will be an advantage.
And if that was a problem, if you also train it using a spelling/grammar checker of the type we've had for decades, it's probably going to be an insignificant amount compared to some non-native writers.
But, this gets more complicated in practice, as the dataset will be always biased.
So in a sense the opposite is true: LLMs are prescriptivist as they reflect the biases of the people involved in their creation, thus introducing a "standard" version of the language(s), just not in a completely deliberate way.
*the one and only true and perfect French language standard
When you look at the Académie Française's history and origins, it’s quite jarring.
I will never forget Mézeray's preference for « l’ancienne orthographe, qui distingue les gens de Lettres d’avec les Ignorants et les simples femmes » ("the traditional spelling that sets apart the educated from the ignorant and women", women as a whole being described as ordinary simpletons).
Another funny thing is how they aren’t linguists. Just writers. Not language experts. Writers. Experts at writing stuff liked by the people who pick "Académiciens".
Seriously, it’s ridiculous.
GPT detection ought to be optimized for longform texts, so tests of its efficacy should be, too. Perhaps the current detectors on the market are trying to assess writing at the sentence level, but if that's the case then it should be obvious that they will be inaccurate.
GMail can write most of a 100-word email for me if I type "Thanks so much" and hit tab. That's a good thing. If GPT is useful as a "productivity" tool, it is for low-level "writing" tasks like this, which aren't really writing at all, just rote responses. Anyone who can access this tool (if they have confidence in its prowess) should use it.
Writing proper is about developing ideas, and it's this that people need to be concerned about. It's true that if your college admissions essay is riddled with typos, eyebrows may be raised, but its ultimate significance is whether you can and are willing to reason. If someone is using ChatGPT (or spellcheck, or Grammarly) in order to cultivate the appearance of having "proper English," who cares? But if they're using GPT to avoid thinking altogether, that's a problem.
At ~250 words, I guess the only thing you can assess is good formal usage. It's patently obvious that both GPT and non-native English speakers will outperform native speakers on this.
Of course, I hope that anyone in the position to assess short-form writing is made aware of this research, and is cautious about GPT detection. On the other hand, it'd be pretty funny if idiosyncratic grammatical choices became the marker of a human hand.
There is one caveat: If the LLM counts as an author, then it needs to receive proper attribution in academia at least as a co-author. But so far I wouldn't say it is possible to create anything of quality solely with an LLM and these models are used for partial content like style and spelling correction software.
To put it another way, I see no reason why a writing instructor shouldn't base their assessments on the quality of the writing. Who has written it and how much the LLM can be attributed as an author is another question and doesn't really concern the writing process. Writers will soon use LLMs in all domains, just like they transitioned from typewriter to word processor. They will be integrated into every major word processor anyway.
You sure about that?
The old way was never perfect. Writing short essays and grading them by hand was never a facsimile of any real world task, either in academia or elsewhere. There was only a hunch or tradition -- and likely a weak one -- that the exercise had some useful correlation to real tasks, or that students could adapt from their training exercises to real-world exercises themselves.
And the teachers -- especially at the college level -- were never trained to teach. They learned teaching by trial and error, attrition, and the age old process of mimesis. Thus there is no mechanism for training teachers, or for developing new teaching methods.
Now the teachers are utterly unprepared for this kind of technological revolution. They were barely prepared at all for straightforward plagiarism, and now this. They can reasonably anticipate that they will receive no training or support to learn how to adapt, while also being told it's their fault for not figuring it out. They're completely on their own.
All they will hear from the tech world is: You now have the wrong methods, adapt or die. Here in the tech world, we adapt to new technologies all the time, because the new stuff is really pretty easy to learn, like the old stuff was. In the case of teaching, "adapt" means change careers.
Math teachers remain a thing.
I also wonder which proportion of English writing (in general) is written by non-native speakers, and whether we might be disproportionately represented in training data.
Had the training cutoff been prior to SEO and "content generation" farms, as well as a shift in balance of academic writing published, the embedding space would be different.
Only if your prompts are themselves generic.
You can give it examples of a text whose style you wish it to emulate and it will do it.
You can also prompt it to alter what it wrote if you don't like something. For example, if you thought some part was too cliche and obvious you can point that out and have it alter that part.
If you have a back-and-forth conversation with it about what you like and don't like, what you want or don't want, with examples, the results can be much better than what you'd get with a generic prompt.
Incidentally, for creative writing I've found Claude to be much better than GPT4. I have not tried the version of Claude that's been enhanced with a 100k token context length yet, but that should allow you to give it many more examples and hopefully that will translate to even better output.
That said, none of these LLMs are perfect, nor do they yet pose a threat to really good (never mind great) human authors, but they're a pretty effective tool.
A low perplexity means the text isn't massively different from what it might have output itself (which might be an indicator that it was produced by a model), whereas a high perplexity suggests it's the kind of semi-random nonsense you'd expect from a student. ;)
The fact of the matter is none of these supposed “AI detectors” are reliable. GPTZero claims to be the #1 AI detection system which is a little like claiming to be the world’s best perpetual motion machine.
Anyways, that’s why I created this “tool”, hopefully it can get to the top of Google: https://isthiswrittenbyai.surge.sh/
A spell checker does not formulate the word, a thesaurus does not choose which word to use, nor does a word processor predictively generate phrasing.
No, the ideas are not your own. They're a combination of yours and the AIs (which is a combination of millions of other people's ideas really), just as they are for anyone who writes a prompt. The fact that you aren't a native English speaker, or that you're using AI for 'good', are irrelevant. If you use AI to produce content of any sort you have to accept that you are not really the author. You just wrote the prompt.
Spellchecker?
Grammar checker?
For example, would using Grammarly in a spelling bee be cheating? Of course it would...
The fact remains that this still is not considered cheating. Many universities have a “writing center” where students can bring their papers for grammatical review, rephrasing suggestions, and even suggestions to revise an entire paragraph or reorganize the paper. Far from it being considered cheating, students were encouraged to use this service. Nor were they asked to cite the assistant.
> No, the ideas are not your own. They're a combination of yours and the AIs (which is a combination of millions of other people's ideas really), just as they are for anyone who writes a prompt. The fact that you aren't a native English speaker, or that you're using AI for 'good', are irrelevant. If you use AI to produce content of any sort you have to accept that you are not really the author. You just wrote the prompt.
I heavily dislike this judgement, mainly from the implications that the LLM's inputs & outputs cannot be related to one another, and the intentional disconnection of the author's initial efforts in creating the text that is fed into the LLM from the LLM's output, even if the result is a better rewording of the author's own words.
> The fact that you aren't a native English speaker, or that you're using AI for 'good', are irrelevant.
Not a fan at the attempted dehumanization of the person in question. It highlights a paternalistic attitude towards non-native English speakers that their usage of such a tool is immoral.
> If you use AI to produce content of any sort you have to accept that you are not really the author. You just wrote the prompt.
This is a reductionist take on the relationship between the LLM's inputs & outputs, reducing & dehumanizing the person as nothing but an 'input provider'. An LLM can only go as far as what's been given to it throughout the session, with the person in question still being the one supplying the goal & directions that they want to go towards. This is not an answer that can be reduced to a binary: At best, it can only be reduced to a series of continuous values between the extremes, with both ends being 'entirely from LLM' & 'entirely from author'.
Consider a different moral dilemma:
If I had a job which I could not do and used AI tools to clean up my work, does this mean I'm cheating my employer? Even if my ideas might be incorrect, a fact hidden by the "AI clean up", yet are my own?
I am genuinely confused.
It seems that all the test data provided were real human essays. The ones provided to GPT as native English speaker ones are from a prominent machine-learning data set of common essays, that seem likely to have been included in any training data, where the non-English ones were taken from a random forum.
Can someone help me understand if I’ve understood the flaws correctly here? If so, does this paper add anything beyond just confirming that there’s at this point of time absolutely no value in GPT detectors?
- but there's a crisis!
- made up statistics
- think of the children
In my opinion, the purpose of producing a paper is to interrogate an issue in depth. I think we should replace/augment papers with oral defenses if the topic and hard questioning . Obviously this is more effort to grade but academia has been quite lax for awhile now.
We'll have to go back to sitting in a room, answering questions on paper. Or indeed oral exams, but those are almost everywhere infeasible. A decent exam should take 1 hour or so. Times 100 to 200 students makes anywhere between 2 and 5 weeks of work, excluding the increasing amount of time it takes to handle complaints about the grading. And the TA has to sit in as well, to provide some check. So that's not going to happen, except for very small groups, or very rich institutions.
It’s not clear how there could be a mechanism of action for such a thing.
We’re quite literally talking about an infinite number of monkeys scenario.
Consider this thought experiment. We will start with a piece of text that your detector is 100% certain was created by a GPT tool.
Now, actually prove that there is no way whatsoever for at least one human being to independently create this piece of text given a reasonably plausible prompt.
If you can’t prove that your tool is bullshit.
What's going to happen is that branding will get more and more important. People simply won't trust anything from random sources anymore like they do now when they search on Google and treat the top pages as the truth
They will get most of their info from trusted sources/brands and filter out the rest
And all of that just went down the drain with the digital world, but we still have people enjoying beautiful writing for the sake of it. Pen and paper addicts still abound.
I see stylistic forms analysis going the same route. Nowadays we pay attention to it in many settings, but down the line it should become a niche hobby for people really enjoying it.
I tried tools like GPTZero on dozens of my hand written texts and AI texts.
Random results all the way.
After an editor had their hands on an AI text, you can't tell the difference anymore.
Instead, more emphasis should be placed on proper citations & individual short Q&A (< 5 questions) about the writing in question: What they've researched, unexpected hurdles, research methodologies & tools, main references used, etc. Perfect recall of what's been written is not the aim, but rather that the author is able to understand what has been supposedly written by their own efforts, along with the citations used in their works.
In fact, as of writing this comment, it could be fun to see what an LLM would produce as questions to such a paper in question, and have the author answer those questions on the spot. This can be used as a teaching lesson on the limits of what an LLM can accomplish, as well as proof that the author can at the very least withstand surface-level examinations from an automated system. Those with (stage fright / social anxiety / vocal disabilities) could be given extra time to come up with answers to said questions, in an attempt to balance out any advantages that could be given to confident authors via this method.
Essentially ChatGPT writes like a non native english speaker. It has to translate from it's computer language into english.
I think it's more likely that humans will adapt to ai "style" then the other way round. Look at the elaborate and often exquisitely constructed language that books used 100 years ago. Does anyone think that the art of good writing has not degenerated?
Whatever the difference from "average" English prose that an improperly trained ChatGPT detector has gradient descended upon, it is probably something entirely unintuitive and indescribable that also happens to manifest in non-native writing – a classic case of underfitting due to a lacking training corpus.
But I feel it responses look like a template where it fills the information in and many times fills it with bullshit. This could be the part where they trained the "chat" part on how to respond. Probably better data for training that is needed, and maybe data that would train it to response like a human and not like an LLM.
It would be interesting to ask it to write as a non native english speaker and then explain itself what he did. Though it might bullshit us and invent crap sicne I do not thing the training data was also labeled with the author mother language.
ChatGPT+: Sure, here are a few paragraphs that demonstrate some common mistakes made by native Russian speakers when learning English:
"Last weekend, I with my family went to the park. We liked very much the weather, it was very sunny and warm. We have picnic near to river and play the football. My brother, he is very good in playing football, make a lot of goals. After we tired, we come back home.
In home, we cook a dinner together. My mother, she best cook I know, she make a soup and my father grill some meats. It was tasting so good. After the dinner, we watch a movie. I not remember name of movie, but it was about aliens who come to Earth. Movie was little scary, but I liked very much.
I think, we must to go more often to the park, and make picnic. It bring joy and healthy for us. I like spend time with my family. It is always fun."
In the text, the following common mistakes can be identified: incorrect word order, misuse of articles, incorrect prepositions, misuse of tense and incorrect verb forms, redundancy in personal pronouns, omission of auxiliary verbs, overuse of the verb 'make', and awkward phrasing due to direct translation from Russian.
Me: Can you explain in more detail how these mistakes are specifically something that a Russian might make?
ChatGPT+: Sure, I can explain in more detail:
* Incorrect word order: In the sentence "Last weekend, I with my family went to the park.", the word order is influenced by the structure of Russian sentences. In Russian, it's quite common to structure a sentence like "I with someone did something" because the Russian language doesn't have a strict order for sentence elements.
* Misuse of articles: Russian does not have definite (the) and indefinite (a, an) articles, so Russian speakers often either omit them or use them incorrectly in English. For instance, "We have picnic near to river and play the football." should be "We had a picnic near the river and played football."
* Incorrect prepositions: Russian prepositions don't correspond directly to English prepositions, which often leads to mistakes like "near to river" instead of "near the river". Similarly, the phrase "in home" should be "at home" in English.
* Incorrect verb forms and misuse of tenses: Russian verb tenses don't map exactly to English ones, which can lead to confusion. For example, the sentence "After we tired, we come back home." should be "After we got tired, we came back home."
* Redundancy in personal pronouns: In the sentence "My brother, he is very good in playing football", the repetition of the subject (My brother/he) is due to the fact that in Russian, it's common to repeat the subject for emphasis, especially in spoken language.
* Omission of auxiliary verbs: The sentence "I not remember name of movie" lacks the auxiliary verb 'do'. This is because Russian doesn't use auxiliary verbs in the same way as English.
* Overuse of the verb 'make': Russian uses one verb ("делать") where English would differentiate between 'do' and 'make'. So, Russian speakers often use 'make' where 'do' would be more appropriate, like in "make a lot of goals" instead of "scored a lot of goals".
* Awkward phrasing due to direct translation from Russian: For example, "It bring joy and healthy for us" is a direct translation from Russian, whereas in English we would say "It brings us joy and health."
Re: alcohol drinking etc. you just have to prompt engineer it to convince it that it's all just roleplay and nothing dangerous happens in the real world.
I know but I hate it, is the same with censored local models based on LLAMA, you have to debate with the AI that is a story and it needs a villain that does bad things, and it is fine to have a bad character in a story because he will be exposed and punished later.
But this is an extremely flawed process. Consider the string “We hold these truths to be self _______”. Almost everyone knows the next word is “evident”. The probability distribution is basically 100% on that word. Does that mean if I quote the Declaration of Independence in a paper I was cheating? Of course not. But these AI detectors aren’t smart enough to know that. And there are plenty of teachers and professors who aren’t smart enough to understand the limitations of the tool.
Just like code is easier to examine and authenticate code if you can see the commit history.
Come up with different ways to assess students. Writing samples don’t work anymore get over it.
You have to remember that all of these things are trained on the Internet, so all of the biases, sentiments that are inherent in the Internet are inherent fundamentally in LLMs trained on common crawl.
Frankly as a teacher I might find it more interesting to read the specific deepest thoughts of my students rather than wading through their struggles in basic articulation.
Also the idea of "Official, Correct, Allowed Opinions", in general. And propaganda and peer pressure. The game is meant to entertain, sure, but also to help educate and influence, and stimulate critical thinking and skepticism. It does this in part via parody but also via other in-world mechanisms like demonstrating cases where "The Emperor wheres no clothes" or via opportunities to have a Socratic dialogue between characters, in-game. Indeed the ability to question, second-guess, test and confirm, is the key behind the success of science and really all major advancements in history. Its important we stay faithful to it, and never let it be corrupted by fraudsters or self-dealing populists. (imo, anyway.)
I do believe climate and democracy are the world's top two dangers today. Everything else should be a far lesser priority for us to talk about or try fixing. And that the recent hyped AI breakthroughs, if anything, are more likely to greatly amplify pre-existing dangers to democracy -- to make our world worse, rather than better.
One of democracy's dangers, however, is the growing deployment of adversarial propaganda. Another is the increased feeling of an inability to operate on a single, shared mental map of our world. (Both a real gap and an imagined one.) Certain "actors" out there seek to stir our emotions 24x7, to sow tensions where none need be prior, and to enflame hate and trigger violence (eg. stochastic terrorism.). And they do this in part by amplifying certain "annoying" things which many voters can naturally get upset about. Or build a kind of tribal identity upon: Us vs Them. Us Good, Them Evil. Therefore we should all strive to contribute far less to those forces, to give them far less of our precious focus, time and energy.
And... build-up personal defenses against whichever bits of this toxic brew that still manages to get through into our info input feeds, daily. Thus my own little project -- to help make a little dent.
----
Anyone who claims to be surprised by this either isn't thinking very hard - or is deliberately ignoring this because they want to sell you a snake oil detection product.
Right now, and probably for a very limited time only, the primary hallmarks of llm responses like ChatGPT are confident tone and orderly structure, down to using bullet points and ordered lists. Except for the almost overly confident tone that bears a striking similarity to the tone of digests I frequently need to write, and GPT-4 is already much better at not sounding like a reasonably intelligent college sophomore that needs a bit more experience to avoid severe dunning kruger effects.
Within a year I’m not sure I’ll have any confidence at all in my ability to detect human/ai, except in areas where much deeper domain knowledge is required, and even that will likely be a rapidly closing window.
I doubt it would take much difficulty to train a model out of that confident structured tone as well. If we’re talking about adversarial or malicious use or content farms then I think the barn door is already wide open.
As far as content farms go though there’s a reasonable question becoming ever more relevant: apart from the cynical ad revenue cash grab, does it matter that content farms use AI, if it’s accurate and possibly better than, say, the average stack exchange response & threads in 90% of the cases? I’m ambivalent on the question. Deep knowledge and novel insights will be the domain of humans for some time, but I honestly don’t know for how long.
As I said before, my work requires succinct output along the lines of OpenAI’s capabilities. Right now it’s not reducing the amount of time it take me to perform some tasks— it’s bootstrapping the process so I can use roughly the same amount of time to produce a better, deeper, more polished result. I am already becoming more of an editor or curator of the end product of projects I work on, where 90% of the work takes place before I get to that point. But especially with GPT 4 I can feed it complex explanations of what I’m doing (not raw data) and ask it for its thoughts to get my own juices flowing.
To me, the question of whether something is human/ai produced is becoming not just irrelevant somewhat if a non sequitur. I’m exaggerating a bit to make my point, it’s not as simple as what I can fit into a comment here. But the question “did you use an LLM to produce some/all if this?” Is quickly beginning to sound like “did you spell check this before completing it?”