Does GPT-2 Know Your Phone Number?
bair.berkeley.edu
bair.berkeley.edu
Unfortunately, both GPT-2 and GPT-3 have a tendency to enter loops. But it's not always bad: https://github.com/minimaxir/gpt-3-experiments/blob/master/e...
> https://github.com/minimaxir/gpt-3-experiments/blob/master/e...
Oh my goodness that is absurd in the most delightful way. Thanks for sharing that.
I worry about this with Google's language translation models, too. It's entirely possible that it's making up phrases or connotations that never existed in organic human speech, but people who aren't fully fluent in the language use Google Translate for assistance, publish something, and then suddenly it's in a published text by an actual human and Google reinforces its own belief.
For at least the past five years, Google Translate has translated "who proceeds from the father" into Latin as "qui ex patre filioque procedit" - inserting the additional word "filioque," which means "and the son." The question of whether to add this word is a 1500-year-old theological argument: https://en.wikipedia.org/wiki/Filioque Since the Western Church added the word, most texts in Latin include it, so Google is almost certainly deciding that the phrasing with "filioque" is more popular - but it doesn't know what the words mean, so it can't realize that the phrase it came up with means something different!
Taking the translation example - how do you enforce that users of Google Translate keep provenance in their translated text that remains with the text?
https://www.reddit.com/r/latin/comments/6akqdi/why_is_google...
Translation for many (most?) other languages does something more sophisticated.
It seems lots of people use training data from Flickr, like COCO, and then use the resulting model for commercial services.
Wrong: I don’t think you‘re paraphrasing article 4 correctly. Both are about science/teaching exceptions. Not general exceptions. https://www.consilium.europa.eu/media/35373/st09134-en18.pdf pp 43
[0]: apparently there were some mistakes in translations which required changes in some language versions, but the english version was final IIRC https://eur-lex.europa.eu/eli/dir/2019/790?locale=de
In general, I think the question is unsettled
So you can happily reproduce the statements of facts and the process. You can’t include someone’s anecdotes about how their grandfather liked to make the dish for New Year’s Eve.
That said, now a lot of privacy policies make sense.
ALSO, it makes me wonder about "Your call is recorded for training purposes"... is it a coincidence or is it very carefully worded?
One aspect that causes this is that historically statistical models calculated from large volumes of text (which is a notion that predates computers, e.g. frequency dictionaries and the whole [sub]field of quantitative corpus linguistics) have been considered facts about that corpus of text and thus not copyrightable at all or (depending on jurisdiction) entitled to different set of protections/limitations assigned to compilations of facts, which give some rights to the people who compiled the facts but no rights to the source of these facts (since facts as such aren't entitled to protection by copyright law).
This also applies to many forms of analysis of audiovisual data, where the copyrights of the source works do not transfer to the results of the statistical or qualitative analysis and can't limit their creation, distribution or sale.
The appropriate analogy to a commercial book or movie is not a translation, but some analysis of it - e.g. a thorough literary review and critique of some book or movie is a separate work with its own copyright, and the original author has no claim on it despite the fact that is (obviously) based on the contents of the work and describes it in great detail. Including verbatim fragments of the work is limited (fair use allows some inclusions but not all), but all the other details are not.
The whole notion of copyrightability of ML model weight files is interesting and IMHO not settled. You could argue that there is some creative expression in forming the model (which would support it being copyrightable) or you could argue that it's a mechanistic result of the application of some algorithm and settings (which have the creative part, and are copyrightable on their own), and so the output can't be copyrightable, no matter how much work (human or machine), time and cost it took - at least in USA copyright law doctrine (e.g. Feist Publications v. Rural Telephone Service) is that mere "sweat of the brow" (no matter how much) does not entitle a work to copyright protection; it requires application of human creativity to create an original work, and automated processes can't satisfy that requirement.
And crucially, if some output is not copyrightable in the first place, it can't be considered a derived work according to copyright law i.e. the exclusive right of authors to create derived works (or grant permission for others to do so) does not apply.
Another analogy might be a simple n-gram model (i.e. counts of bigrams - word pairs, trigrams, etc) which is quite clearly a mechanistic noncreative collection of facts about a dataset, and is also able to "answer" questions such as what is someone's telephone number if that was in the source data.
This is an interesting angle I had not considered before. It seems like “right to be forgotten” requests could be quite damaging to the “train once run anywhere” promise of some of these models. (Or this could just mean that the training data needs to be more carefully vetted for personal data, but probably both as no vetting process can be 100% successful).
"I'm afraid I cannot tell you that, Dave"
Sorry about that.
Any resources?
In neural networks, weight changes do carry meaning. If your network has particularly large updates for the horse recognition neuron, you likely watched horse pictures. If it has updates for handbags, you likely watched pictures of them. If the company averages over all your submissions, the noise will be less relevant and the handbags and horse pictures will eventually show up.
For your example, sure it's possible for the server to do something you don't know. But it's the same people on the server doing the aggregation as on the client doing the obfuscation. If they really wanted they could easily e.g. set the seed for the noise to be based on a key which they have on the server, which would be very difficult to detect. I think you just have to trust that apple wouldn't let that happen.
Most countries have libel/defamation related laws that cover this and I hope this gets tested in court soon.
Exposing software/machine learning algothrims an entity doesn't fully understand sholdn't be a defense in court. At the moment developers just throw a sentence in their software license saying they aren't liable for damages but this isn't good enough. Someone is liable, if the law decides that the original creator isn't liable then the entity that hosts/runs the software needs to be.
I wonder the degree to which this inference is true in practice with respect to information like phone numbers... How exactly are the train and test sets formed in a de-correlated-with-respect-to-memorization-of-phone-numbers manner for models of GPT class that are trained on corpus's the size of the internet?
If a particular person's phone number occurs 1000 times in the corpus prior to being split into train/test sets, what are the chances that the number only appears in either the train or test set but not both?
It's better at being a text editor, in that you can click the text and edit it! (Sweet relief.)
But it's missing the world-info feature from AI Dungeon, and without that I don't think it's practical to write anything long-form.
So even if it all is mathematically equivalent to approximate lookup and approximate memorization, that doesn’t mean it’s “just” that.
“We show, however, that deep networks learned by the standard gradient descent algorithm are in fact mathematically approximately equivalent to kernel machines, a learning method that simply memorizes the data and uses it directly for prediction via a similarity function (the kernel)“
However real machine learning tries to approximate functions in n-dimensions and that is really, really, really hard to do. Currently no one really has a lookup table. There are some inputs where the error level is acceptably low, and others where the model just isn't optimized enough and the errors are ridiculous. The only question that remains to be answered, is whether any of these high-dimensional functions can actually be found by machine learning, or if we are just stuck with these endless approximations. Also I suppose you could ask if any such functions actually exist; maybe certain phenomena are just pure chaos.
Does the human memory work the same way?
Alternatively, they could do something like GAN and have a 'discriminator' classify if a sample is natural or synthetic. Then, at inference, condition to be original.
So, verbatim training data reproduction - I don't think it's going to be a problem, I think the author is making too much of it.
On the contrary, let's have this knob exposed and we could set it anywhere between original and copycat, at deployment time. Maybe you want to know the lyrics of a song, or how to fix a Python error (copycat model is best). Maybe you want an 'original' essay for homework inspiration. Who knows? But the model should know PII and copyrighted text when it sees it.
But if you mean by original to invent a whole new genre, or completely new esthetics, I agree, it's out of its scope. It is a great interpolator.
You can use gpt-2/3 to determine whether it's generated text is "original vs copycat" to a degree by changing the prompt. This is currently more of an art than a science
That's expected when you publish anything online, you lose control over the data.
Or GPT accidentally defaming someone? That assumes bad intent which is a bit problematic with a language model.
With made up contexts like this in the article it's easy to cry wolf.
If I leave the house I take on some risk, but most of the time I'm in control.
This is infantile reverse justification from an end. GPT style corpuses are clearly problematic without a lot more smarts on figuring out what's appropriate and what's not.
Responsibility is on those who publish content online, not the tools.
We should be more concerned about mitigating GPT-like systems because one thing is for sure, we can't stop these models. The genie is out of the bottle and I'm sure multiple actors are working on similar systems.
That is a great way of thinking of it. Lots of information, particularly from 90s and 00s, got put up on the 'Net with no intention of it being archived in perpetuity for public consumption and used for unintended purposes.
It was very good at guessing the right person.
I'm curious if one can use GPT-3 to do some clustering of people based on their writing. For example, if I take the writing of several diagnosed sociopaths as a prompt or fine tuning data, could I use GPT-3 to detect same in the wild?
I would imagine GPT-5 or 6 will start consuming video as well. This will be interesting if you start mixing the content from YouTube, TikTok, WSHH, etc. into the mix. So not only will it be able to generate text, it will generate a convincing video of a person speaking to you with plausible facial expressions and intonation in speech.
Every company stores this stuff for ages, and the value of candid conversation data just keeps increasing. Eventually these companies are either going to get hacked, get bought, or go bankrupt, and all the cleartext data they hold is going to get passed around to various data markets and end up incorporated into GPT-12 or whatever.
The moral of the story here isn't that companies need to stop storing data, or that we need to run ML researchers out of town. It's that people really need to start using the encryption technologies that were built decades ago to protect some of their most valuable assets, their mental model of the world. Otherwise these systems, which are being trained to extract as much value out of you and the ones you love as possible, will use this data you're giving them for free against you, and it'll be to late to do anything about it then.
I've no doubt taken into account copyrighted works and personal information when training my built-in neural network. The examples the article gives like "misremembering" the murder as the murder victim sounds like something a person would do. Knowing verbatim contact information of some random person is also possible. All in all, GPT-2 sounds a lot like us.
Thus “free will” must come from some sort of faith. Either it was given to us by “God”, or we’re all part of some simulation and it exists on a level we don’t know about, or something else entirely too advanced for us to understand.
I take the view that ∀ x ∈ <free will definitions>, Has(x, humans) ≡ Has(x, AI)
∀ x ∈ <free will definitions>, Has(x, humans) ≡ ~Has(x, AI)
For these to both be true, <free will definitions> becomes empty. Thus I argue we have none. Edit, or rather it’s meaningless to argue if we have any or not, as it cannot be defined.
This sums up my view too. Can't argue about it if you can't define it.
Free will is faith in the sense that we learn the concept from our culture as our sense of self develops. The assumption of having it shapes our thinking such that we can say we actually have it.
A person who took their inner voice as voice from the gods would have a restricted free will in this view.
You can ask a person "did you come up with that sentence, or is it something you read?" And they can answer that question. The answer isn't perfectly reliable, but it isn't completely unreliable either. You can't do that with GPT-2/GPT-3.
Big picture: the human mind is a machine of sorts. It's not the same kind of machine as any form of machine learning we've developed so far.
How would we be able to tell how reliable it is in general?
"""The left brain of people whose hemispheres have been disconnected has been observed to invent explanations for body movement initiated by the opposing (right) hemisphere, perhaps based on the assumption that their actions are consciously willed.""" - https://en.wikipedia.org/wiki/Neuroscience_of_free_will#Rela...
No idea if it would be feasible to integrate this information directly into the model. With my limited understanding of neural networks, this seems difficult.
There is a lot of space between “user data and ML research on that data needs to be closely regulated” and “run the ML researchers out of town.” This uncharitable exaggeration makes critics sound like an uninformed mob and closes yourself off to legitimate ideas. This directly leads to logically flawed and morally gross suggestions like:
> It's that people really need to start using the encryption technologies that were built decades ago to protect some of their most valuable assets, their mental model of the world.
It is simply not reasonable to expect most Facebook/Skype/etc users to know enough about encryption to make good decisions here - for the same reason that you can’t expect people to have detailed understanding of food science to protect themselves from unscrupulous grocers, or to have a detailed understanding of medicine to protect themselves from fraudulent doctors. It really doesn’t matter if this knowledge has been around for “decades” since it’s still specialist knowledge.
Protecting digital privacy from unscrupulous tech companies and unethical ML researchers is a job for the government. Suggesting otherwise is victim-blaming. “The problem is that society is dumb and needs to become smart through the power of self-righteous scolding” is not helpful.
'The majority of humans are idiots who need to be protected from themselves by the government.'
Because the government is always scrupulous and trustworthy...and nobody that's ever said trust us we'll protect you has lied...
Nobody is disputing that the government can be corrupt, and at its best will be incapable of protecting everyone from everything. And if you are sincerely anarchist then that’s more of an ideological dispute beyond the scope of this comment.
But I suspect that you actually support regulation around food and medicine, since “a fatal case of salmonella is a small price to pay to make sure we don’t rely on the government” isn’t actually a good argument. And, unlike people who are ignorant about computer technology, you probably have empathy and understanding for people who are ignorant about whether the ground beef they are purchasing is actually safe to eat. You would not be sympathetic to a libertarian food safety specialist came in and said “well I can tell that the ground beef is unsafe, why not teach everyone my skills so they can defend themselves?”
The idea that the government can protect people when it comes to health, medicine, transportation, and personal property... but NOT when it comes to privacy, is just not coherent.
I sincerely have no idea what your problem is.
• https://worldbuilding.stackexchange.com/questions/51746/when...
• https://kitsunesoftware.wordpress.com/2018/10/01/pocket-brai...
https://en.wikipedia.org/wiki/White_Christmas_(Black_Mirror)
The obvious response to this is, "Then don't patron those companies," which is just as naive and oversimplified of a solution as the one above suggesting everyone everywhere should just use encryption (but for very different reasons).
There is just simply no logical reason to believe consumers would ever get data rights over corporations under capitalism with major reforms putting corporations back under our boots where they belong.
Potential damage: the sim can do any work which meat-you might otherwise be paid for, building the things the megacorps want to build, while meat-you sufferers from <insert dystopian fate of your choice here, because you now have less annual economic value than a guide dog costs today>.
Even worse outcome: the sims are sentient and they know they’re enslaved and they can’t do anything about it.
(I don’t think anyone yet knows if sims would be sentient, or if it might be optional, but in the worst case dystopian nightmare the combination would definitely be even worse than just one of the two).
This would make a good SF story. Probably has already :-)
It's a variation on the transporter problem from Star Trek - https://www.youtube.com/watch?v=nQHBAdShgYI
General state of encryption technology is that it's arcane and unusable outside of like a couple of chat apps
The problem isn't really the technology, it's that the people you are interacting with on the internet don't care about privacy as much as you do, and obviously when it comes to keeping things private, everyone has to participate.
You are proposing a technical solution to a social problem. That pretty much never works.
It's like trying to build a dog by hanging around the park collecting turds.
This is in addition to the ever greater sample-efficiency of bigger models, which learn eerily fast from just a handful of datapoints, rendering the supposed 'moats' of big giant proprietary datasets ever more moot...
-- Richelieu
"Moreover, if such a request were granted, must the model be retrained from scratch? The fact that models can memorize and misuse an individual’s personal information certainly makes the case for data deletion and retraining more compelling."
Besides, how would one even know that their info was used in a training dataset, only if and when it's revealed in a generated excerpt?
Still wish the old talktotransformer model was released, instead of monetized behind a new company. I haven't been able to find a comparable model yet.
Around that time (since no one else was doing it) I released a wrapper to streamline that code and make it much easier to finetune on your own data. (https://github.com/minimaxir/gpt-2-simple)
Nowadays, the easiest way to interact with GPT-2 is to use the transformers library (https://github.com/huggingface/transformers), of which I've created a much better library for GPT-2 that leverages it. (https://github.com/minimaxir/aitextgen)
The stuff about whether or not models should be destroyed if they contain copyrighted work gets kind of chilling if models actually achieve sentience someday. If I could make a faithful copy of my consciousness, my consciousness can reliably reproduce numerous copyrighted works. As can anyone.
Of course, most of those I deliberately memorized. I think it's a crucial thing here that the model is not actually memorizing anything - if we assume for a moment that the model is a consciousness, these are all half-remembered snippets leaking into casual conversation. And I think it's most likely any truly conscious entity is going to do that sort of thing from time to time.
I can foresee a future GPT-x that doesn’t know my phone number, but can deduce it.
What you should be doing is replacing no less than yearly any numbers you have, also email addresses and anything within your control (changing physical address is much more difficult).
You may even consider changing your legal name if that's a consideration depending on what Google has on you.