GPTZero Case Study – Exploring False Positives
gonzoknows.com
gonzoknows.com
I would have thought this, but every attempt I've seen at detecting chatGPT generated text has failed miserably.
Also false positives are typically "this text is likely to contain parts that were AI generated" rather than "This text is higly likely to be AI generated" (which is what GPT-generated content generally produces).
When I've tried to prompt-engineer GPT to produce text that GPTZero will flag as negative it has been pretty tough!
e.g. If you were Sam Altman at OpenAI and your use-case is mostly looking for training data and wanting to tell if it is AI-Generated or not (so you can exclude this from training data), you probably care much more about false negatives than false positives (false positives just reduce your training data set size slightly, while false negatives pollute it).
Of course they matter if you are marking homework (where conversely false negatives aren't actually that important!), but it's pretty trivial to think of use-cases where the opposite is true.
I haven't tried it though.
I feel like for most emails I write, information density is close to a maximum. This means there's no actual gain to be had from a language model. The email I would write myself is going to be about the same length as the prompt I'd have to write anyway.
Even if English is your mother tongue, if your written English is crappy or you need to write in a style you are unfamiliar with (e.g. formal), then ChatGPT can help.
I also think writing yourself might be far better practice. This tool can easily become a crutch. This is unlikely to be free anytime soon. In fact it's likely to be quite expensive.
I can ask ChatGPT to rewrite an email in the style an American news reporter from 1950, and I can judge whether some of the cliches it generates feel correct. I cannot write in that style at all.
Whatever LLM search stuff comes along will only be free as long as it brings in ad revenue. Which involves making the models fundamentally worse most likely. Or they'll use it to collect personal data. Probably both.
Computationally, GPT is wildly expensive. This idea people have that it's gonna be used all over the place for all sorts of tiny tasks, as if it's just another REST API, is nuts. Unless something fundamentally new comes along that makes these models much cheaper, adoption is likely going to end up much more limited than people expect. Or siloed off into expensive business-facing products.
If I'm making (or valuing my free time at) $20 an hour, then ChatGPT only needs to save me one or two seconds and half a minute respectively.
So many experiments to run...
However once you prompt the LLM with a higher temperature, or tell it to roleplay as someone with elaborate personas, or suggest to use certain linguistic styles, or train it on example text... then it becomes much harder.
I imagine pathological cases of formulaic word use, sentence/paragraph structure will only be detectable in longer form text. After all text is already pretty low-resolution, not much for adversarial models to work with.
Well, to be clear, they can put rules based filters and other things on top of the neural net, but the core GPT will never get more accurate since it has no mechanism to understand what words mean.
Our models sizes are a product of our scaling and hardware limitations. There's no reason to believe we are anywhere near optimal.
It also seems reasonable to assume that they will eventually encounter diminishing returns, and that the current issues, such as hallucinations, are inherent to the approach and may never be resolved.
To be clear I don't have a clue which statement is true (though I don't see why scaling would solve the hallucination problem).
If it doesn't produce better results however then they want their competitors to waste lots of money to make the same mistakes, there is really no benefit from publishing that and lots of drawbacks.
Otherwise it seems too much of a coincidence that Google and OpenAI ended up with models of basically the same size. Google could have trained a model 5x-10x larger easily, it isn't that expensive to them, but for some reason we didn't see that, and GPT-4 just never seems to launch.
Now, since the larger model wasn't good enough to replace a human engineer we can rest easy, it wont replace programmers anytime soon. If GPT-4 for example could replace engineers, OpenAI wouldn't need to monetise ChatGPT, they would just rent out artificial engineers to do coding for $10k a year.
Then again, we are a big old bulb of wetware and we can generally learn to apply grammar rules correctly most of the time (when explicitly thinking about them, anyway).
Maybe what we need is some kind of meta cognition: being able to apply and evaluate rules that the current LLMs can already correctly reproduce.
Please say more about what you mean here because I disagree.
It’s certainly more eloquent, but it still can’t multiple 2 4-digit numbers…
The four digit number thing is just the current lower bound of where it gets confused because a lack of training data.
Once you teach a patient 8 year old the rules of multiplication once, they can multiple any two numbers (that they’ve never seen before) with an arbitrary number of digits. An LLM cannot and will not ever be able to do that because it is a specific tool and it is not designed to do that (doing that would be a bad outcome for an LLM since we have different tools that can do multiplication much more efficiently).
So yes, if a LLM learns rules based math (which it is not intended to do) I’ll eat not only my, but every hat in existence.
> "The periodic table is a systematic ordering of elements by certain charcteristics including: the number of protons they contain, the number of electrons they usually have in their outer shells, and the nature of their partially-filled outermost orbitals."
> "Historically, there have been several different organizational approaches to classifying and grouping the elements, but the modern version originates with Dmitri Mendeleeve, a Russian chemist working in the mid-19th century."
> "However, the periodic table is also somewhat incomplete as it does not immediately reveal the distribution of isotopic variants of the individual elements, although that may be more of an issue for physicists than it is for chemists."
GPTZero says: "Your text is likely to be written entirely by AI"
Now I'm feeling existential dread... perhaps I am an AI running in a simulation and I just don't know it?
You just illustrated technical writing. Naturally, your writing style is very similar to that of other technical writing.
Take one guess what kind of writing exists in most of the text GPT is trained on.
The colon in line 1 is clunky, the combination of "but" and "with" in line 2 reads as passive, and line 3 is full person.
A partisan writing cliched slogans and regurgitating tired political statements has a “temperature” closer to 0, likewise an engineer stringing together typical word combinations. A poetical writer with surprising twists and counterintuitive mixtures of words has a temperature closer to 1.0.
I can see a few clichés in your writing, also I did a search on some fragments of sentences which showed a number of results.
If you want to be less “robotic” then you could add whimsical or poetic wording, and less common turns of words (be a phrase rotator).
Being able to speak to machines would likely correlate with “sounding like one.”
There could be a different (mis)categorization here but also por que no los dos.
But the way I see it, if ChatGPT thinks the Python list object should have a .is_sorted() property, that’s a pretty good indication that maybe it should.
I work in PM (giant company, not Python), and one of these days my self-control will fail me and I will open a bug for “product does not support full API as specified by ChatGPT”.
This is a very common feature of delirium in people. Chatting with an LLM seems a lot like what it would be to talk to a clever person with encyclopedic knowledge, who is just waking up from anaesthesia or is sleep talking.
Yes! And when it hallucinates references for articles, often times those articles probably should exist…
When a new person is born their entire life is hallucinated in its entirety by the all great and powerful GPT. Deviation from His plan is met with swift and severe consequences.
Hahaha, Python language fixing itself!!!
Of course, this might mean future Chatbots would successfully emulate that. But it's not impossible an "adversarial style" exists - this wouldn't be impossible to emulate but it might be more likely to cause the emulator to say things the reader can immediately tell are false.
One idea is to "flirt" with all things that people have come up with that AI chokes on. "Back when the golden gate bridge was carried across Egypt..."
Prediction #2: As more people discuss ChatGPT online, by late 2023 discussion of Roko's Basilisk exceeds discussion of ChatGPT. (half /s)
#1 is already happening !
See here (other HN thread) : https://twitter.com/tobyordoxford/status/1627414519784910849
That's already happening I know but it will be amplified to the point that all humanity in writing in lost. All ideas in writing will be a copy of a copy of a copy and merely resemble something once meaningful. Time to go touch grass.
We train AIs but they also train us.
It’s like saying that GPUs can render games, so GPT is a game because it uses GPU.
https://twitter.com/raphaelmilliere/status/16240731504754319...
I don't see how it will be possible to build such a tool either as the combination of words that can come after another is finite.
But it's still fun to deduce that the reason is that the quality of technical writing has sunk so low, that is even below the standards for AI generated text.
Perhaps we can call it the "Synthromorphic principle," the bias of AI agents to project AI traits onto conversants that are not in fact AI.
Which is better than 50% but not nearly good enough to base any kind of decision on.
Unrelated: p-value for getting 12 from 20 correct just by chance is ~0.4 that is there is not enough data for the conclusion "better" in this case.
Null hypothesis: 50%/50%, the result random, normal distribution:
H0: p=1/2
H1: p!=1/2 (two-tail)
import statistics
p0 = 0.5 # proportion of successes according to null hypothesis
n = 20 # sample size
p_sample = 12/n # 12 from 20 are correct
sigma = (p0 * (1 - p0) / n)**.5 # std according to H0
z_score = (p_sample - p0) / sigma # test statistic
p_value = 2*statistics.NormalDist().cdf(-abs(z_score)) # prob. two-tails
# p-value -> 0.4Feels like a good basis tech for something like ChatGPT.
Another challenge is that large language models are highly complex and constantly evolving. GPT-3, for example, was trained on a massive dataset of text and can generate text in over 40 languages. With this level of complexity, it can be challenging to develop detection systems that can keep up with the ever-changing text generated by these models.
To implement a reliable detection system like GPTZero, which is designed to detect text generated by GPT-3, several challenges need to be addressed. First, the system needs to be highly accurate and efficient in identifying text generated by GPT-3. This requires a deep understanding of the underlying language model and the ability to analyze the text at a granular level.
Second, the system needs to be scalable to handle the vast amounts of data generated by GPT-3. The detection system should be able to analyze a large volume of text in real-time to identify any instances of generated text.
Finally, the system needs to be adaptable to the evolving nature of large language models. As these models continue to improve and evolve, the detection system needs to keep up and adapt to the changing landscape.
I am not sure how you would code something to detect an author based on writing style. It feels like something people would have tried to do before. Probably using a similar approach that LLMs use but with a separate predictor for specific authors.
I think this is probably because it doesn't match the conversational style of a forum discussion.
I'm impressed it is ever accurate.
False positives and negatives all over the place.
I really wish it worked, but generally, it doesn't.
[^1]: Which may take a bit to happen, because our current crop of AI researchers have all taken "The bitter lesson"[^2] to heart.
[^2]: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
At least for now,I was thinking it didn’t do that, and maybe the lack of references would be an indicator of unedited gpt output.
...
> However, according to a study published in the Journal of Dairy Science, the diacetyl content of butter can range from approximately 0.5 to 14 parts per million (ppm) (https://doi.org/10.3168/jds.S0022-0302(03)73775-3).
The doi I could not find, so I'm pretty sure is bogus.
I asked it to produce a full reference:
> Sure, the full reference for the study I mentioned is:
> Yvon, M., Chambellon, E., & Bolotin, A. (2003). Effect of pH on diacetyl and acetoin production by Lactococcus lactis subsp. lactis biovar diacetylactis. Journal of dairy science, 86(12), 4068-4076.
I went to the index of the journal 86(12), and that article is not there.
The idea seems quite roboust to me and I can envisage a future where companies that provide access to LLMs would also publish a detection tool for their models.
[0] https://filteroutai.com/validate/a07e081b71b294ba2de236441be... https://filteroutai.com/validate/2c3fa6de32845df02be7a4ff185...
[1] - https://github.com/openai/gpt-2-output-dataset/blob/master/d...
I didn't even have to look at the internals to know GPTZero and OpenAI's solutions would be pointless. I'm still very surprised OpenAI made snake oil.
I could write in the tone of ChatGPT if I tried hard enough. It's an intractable problem and a tool like this probably does more harm than good.
@dang
If fox news deletws the last 10% of reality from it's broadcasting, the people might start catching on. Although these days i am not sure
...nice edit...