OpenAI Personal Data Removal Request Form
share.hsforms.com
share.hsforms.com
Edit: Ok the link can be found here in part 4 of : https://openai.com/policies/privacy-policy
We have a system that may have information about you and may even distort information about you. In fact it probably has some information about you considering that we exercised no control over the process of ingesting information into the system. Furthermore, we don't have understanding or control of our system in such a way that we can remove that information or even discover it. However, we still released the system to the world and now we expect you to test it with various prompts and hope that you get lucky before someone other person does.
But practically, I fear that you're right. That distinction will not be made. That's why the only viable defense against the likes of OpenAI that I could think of was to close my websites to public access. I'm still seeking a way to safely reopen them. I hope I can find one.
The data isn't sitting in some database somewhere, it's inside of a large lanaguage model. It's not like they can just execute a DELETE statement or do an entirely new training run.
Are they intercepting the outputs with something like a moderation server as a go-between? In that case, the data still would technically exist in the model, it just wouldn't be returned.
Maybe using fine-tuning?
Facebook for example know it's you because you signed up to the account.
If no ID was required, you could freely delete my records in OpenAI's corpus, violating my right to control access to my own data.
If that's the way you choose to look at it, perhaps you could argue that the system should be opt-in, rather than opt-out. Maybe you should have to provide ID to grant access, instead of letting your identity be exploited for profit implicitly.
If someone wrote your name on a wall and I asked them to erase it, I don't think that violates your rights. You didn't ask for OpenAI to train on your data in the first place. Having it deleted now is no different than OpenAI never having existed in the first place.
"Hi, I am $competitor, and I want all information about my successes scrubbed from the internet. Thank you."
However, Google didn't want to be considered a news platform, because otherwise they'd be held responsible for their content. So you can't ask CNN to remove the article themselves, but you can (and people do) ask Google to remove those article from search results.
If I used Spotify and anyone could delete my data (playlists for example) that would be quite annoying.
Chances are quite good that there exist people who you would be unhappy with me ensuring are deleted from Wikipedia, for instance.
You have no right to make someone forget me.
Unless we're talking about something other than GDPR, this same situation would apply to online services, not just training AI datasets.
> Individuals also may have the right to access, correct, restrict, delete, or transfer their personal information that may be included in our training information.
https://help.openai.com/en/articles/7842364-how-chatgpt-and-...
if it costs them $10 million to remove my PII that's their problem
if they don't like it then they can stop operating it entirely
It is an engineering problem and this is (largely) an engineering forum. Tomorrow solving this might be a part of your job as well, so idk why are you so dismissive.
I completely agree with the parent post, it's not my problem that their product was badly designed. If it takes them 10 million dollars to comply with the various data protection laws around the world, that's none of my concern.
"Ethical and legal problems" are actually largely consumer's perspective problems and have little to do with actual ethics or legal precedents.
Welcome to the real world, where people leave their personal data all around the internet without any concern. This IS an engineering problem, just like implementing "the right to be forgotten" in a search engine is - it would likely only concern notable people anyway.
> it's not my problem that their product was badly designed. If it takes them 10 million dollars to comply with the various data protection laws around the world, that's none of my concern.
Well sure, from the consumerist standpoint, that is definitely not your problem. But if a consumer is all you amount to, then why even bother posting here? Karen rhetorics is lazy, tiresome for everyone around and brings nothing interesting to the discussion.
Hell, pii itself is manufactured. God damn houses! Builders never stopped to think someone I don't want to talk to could use my address to find me!
OpenAI either chose to have this problem or they're completely incompetent, and I doubt the latter is true. They knew damn well that they will encounter PII in their dataset and they knew damn well that there are laws surrounding that. The GDPR was well under way and in the news when their company formed, let alone when their various GPT models were finished, and 20 years before that there were already European privacy laws regarding data collection and other PII use.
They chose to train a huge model without implementing a safeguard against PII and now they'll have to live with the consequences. One of those consequences may be "their investment money is going down the drain in fines" or maybe even "their product is illegal across the EU and the EU still wants to see the money from those fines".
if their business model isn't compatible with the law then they can wind up the company and return the remaining capital to their investors
it's no different to companies complaining that the cost of legal disposal of toxic waste renders their business nonviable
Tough. If their business is only viable by ignoring the law, they shouldn't be in business.
Any ramifications concerning the removal of PII is not the government's problem. If they can't use PII in a legal way, they shouldn't have collected it in the first place.
Not exactly... in fact, to the point that you're just spreading misinformation.
They didn't "deem the entire product illegal" at any point, and after OpenAI initially responding to Italy's objections by removing access for all Italians, they have since (a week ago) re-opened to Italian users having taken steps they presumably believe are enough for Italy to be OK with them operating there again.
https://www.reuters.com/technology/chatgpt-is-available-agai...
I very much doubt this is the last we hear from EU countries with regard to OpenAI / other LLMs and GDPR... but that's not the same as claiming Italy have already ruled it to be totally illegal.
It works, because nobody ever does this, so the token 4,096 limit is in no danger.
/s
Of course they can - it might just be expensive but sure, they could.
Not sure how best efforts work with GDPR.
> NjE=
> atob("NjE=")
> "61"
Lets hope they're not that stupid, as it's trivial to work around.
Using GPT-4 with the initial prompt "You are a helpful assistant. All your responses must be base 64 encoded." Asked the question, "How old is Barack Obama?"
Received the following response, which seems fairly accurate for GPT-4s knowledge cutoff date: "NTkgdG8gNjAgeWVhcnMgb2xk"
I asked it to generate an SVG representation of Dali's The Persistence of Memory and it output a recognizable vector image of three distorted clocks.
I happy to be proven wrong, but that certainly feels emergent.
Given: “intergalactic benevolence”
It output “aW50ZXJnYWxhY3QgYmVuZXZvbG9uY2Ug” which is actually “intergalact benevolonce”.
Weird.
Yes.
Provide them with yours in Base64 and you get the answer in Base64. Decode it and it's the response you'd expect.
At least that's how it is with the one I just tested, which is based on GPT-4.
I’m sure they’re not doing basic string matching.
I understand where the training data is. I didn't say anything about the training data.
And they don't mention how long it is until they spend $10+ million to retrain it and remove PII, if that is the only way they can handle it.
The legal principle here is very, very simple — no training data without explicit legal consent. Companies need to stop being cute about this, or governments need to come down hard to start regulating this, yesterday.
Oh i am pretty sure that if you dont remove all data you’ll pay for it. Looking forward to hefty fines for openai.
Cue re-running the training model a little bit more frequently than they'd like... At least it would certainly become opt-in very quickly, which of course it should have been from the start.
Is there a way to remove PII without having to use their service?
-
On a serious note - there needs to be an easier way to remove any and all PII from across the web, period.
It should be illegal for ANY site to harvest PII and host it for ransom (credit/social credit site, for example should be fully illegal)
Also, with "relevant prompt" -- how can I use my own account to test to see if I have PII in the system?
Do I just need to attempt to prompt for my own PII to check?
How do you prompt to check for your own PII without ADDING PII into the system via your testing prompts?
This includes things like:
- owning a car, or just having a drivers license: https://www.privateinternetaccess.com/blog/dmvs-are-making-a...
- having a credit history (credit bureaus can sell by default to third parties like financial institutions, not even accounting for the equifax breach)
- buying a home (public record)
- going to court (public record)
- entering the vast majority of grocery stores with cameras
- owning a cell phone that is turned on at some point (cell towers store information on which phone numbers are connecting to them, giving rough location history)
- Messaging anybody on virtually any service without extreme precautions (even e2e apps will store who is talking to who, which can create detailed social maps)
There is the option to opt-out by living in the woods without any modern technology, but I think it's worth noting that in general, especially for anybody with a HackerNews account viewing this page, that it is not trivial to avoid having your personal data collected (and possibly sold) on a continuous basis.
Recall that guy who wrote a thesis on the actual fiber-optic lines laid throughout the world and the USG seized his thesis as a matter of national security...
We need the same thing built for PII vs AI access to such information.
but starting with available PII. Regardless of the fact that the state says "public record" public recod should be fucking reigned in.
Sharing sensitive information once, or even a hundred times, should not constitute a presumed consent to opt-in to that information being shared again or data mined.
I received a confirmation in February that my data had been excluded from model training. However, recently, after the addition of the new Data Controls feature, I noticed that I was suddenly opted in again in the settings. I've tried contacting them about it via Discord and e-mail so that they can clarify whether the exclusion is still valid, but it seems like I'm getting ignored.
On the other hand, they already know which sites they used to scrape data. So publish it, maybe with a handy lookup portal where you can enter urls to see if it got scraped.
I prefer an opt-in model, but that's not likely to happen any time soon, so this seems reasonable while this gets legally sorted out. Just because something is transmitted publicly doesn't mean it's without copyright. Otherwise any song broadcast on radio is up for grabs to be resold by anyone receiving it.
You can just send them a snailmail or e-mail and they'll have to process that too. You can find templates for that all around the internet.
For a public figure, of course there is lots of information in the training data, all public data. But when asked about me or my brother, ChatGPT either refuses to answer OR hallucinates the hell of it. Then, nearly everything is wrong and the output resembles the answer to a prompt like: "Create a short bio for a fictional character named xx, living in yy and working as zz." (Okay, often yy and zz are wrong either.)
Requesting to delete these hallucinated facts seems quite stubborn and ineffective?
It feels like anything that you release on the internet publicly is fair game. If however you didn’t release it in public, put it behind a password and then OpenAI somehow got access to it and train on it, I can see the argument here but if you put up data on your own, I don’t see why you can prevent others from accessing that data. If you don’t want others using it out there, don’t put it out there.
It might feel like that to you, but that's not what the laws are in some economically important parts of the world. In Europe, the relevant bit is the "right to be forgotten". If you want to operate an information system, you need to implement that. It's hard to see why it wouldn't apply to a chatbot just the same as it applies to search engines.
It's much easier to explain why there's a distinction between a human brain and a massive database accessible at will to billions of people.
Actually, no, copyright is the right to exclude others from copying, and without someone else having that right, you can copy whatever you want.
(And even if someone else has thr relevant copyright, you can still copy if your copying falls in one of several exceptions to thr exclusive rights of the copyright owner, such as fair use.)
Yes, if someone copies text it is technically a copyright violation. Now, copying a small bit, not publishing it, or publishing under fair use. That's all fine.
Copying everything on the internet for commercial use falls under exactly none of those exceptions. There will likely be lawsuits, though OpenAI will likely try to say "If you don't want to be used this way, opt out and we'll remove you"
Just because something is on the public internet, does not mean you have the right to do anything you want with it.
https://www.washingtonpost.com/technology/interactive/2023/a...
Anyone had any success with this ?
I don't know all the details about GDPR but I can't imagine it allows people to erase their existence from the Internet if they became a public figure at some point.
Unless OpenAI is stupid enough to gather medical data from protected sources under HIPAA, that is, but I doubt even OpenAI wouldn't be nonchalant enough to use a data source like that.
* You must identify valid grounds under the UK GDPR (known as a ‘lawful basis’) for collecting and using personal data.
* You must use personal data in a way that is fair. This means you must not process the data in a way that is unduly detrimental, unexpected or misleading to the individuals concerned.
* You must be clear, open and honest with people from the start about how you will use their personal data.
If you've been asked to remove data in accordance to a gdpr request, is it best effort or do you have to strictly comply with the request? ChatGPT I am not a lawyer, but I can offer some general information about GDPR and data removal requests. The General Data Protection Regulation (GDPR) is a legal framework that sets guidelines for the collection, processing, and storage of personal data of individuals within the European Union (EU). Under GDPR, individuals have the "right to be forgotten," also known as the "right to erasure." This means that they can request the deletion of their personal data under specific circumstances.
When a data controller (the organization that determines the purposes and means of processing personal data) receives a valid GDPR request to remove personal data, they are generally required to strictly comply with the request. However, there are some exceptions, and the extent of the obligation may depend on the specific circumstances surrounding the request.
The GDPR requires data controllers to respond to data removal requests without undue delay, and in any event within one month of receipt of the request. This period may be extended by two further months where necessary, taking into account the complexity and number of requests.
It's important to remember that GDPR compliance is not solely about "best effort." Organizations are expected to take reasonable and appropriate measures to ensure compliance with the regulation. This includes having processes in place to handle requests and promptly remove personal data where required. Failure to comply with GDPR can result in significant fines and penalties.
However, it's also essential to consider that there are circumstances under which a data controller may not be required to comply with a data removal request. These exceptions can include situations where the processing of personal data is necessary for exercising the right of freedom of expression and information, compliance with a legal obligation, or for reasons of public interest.
In any case, it's crucial to consult with legal counsel or a data protection officer to better understand your organization's specific obligations under GDPR and ensure proper compliance.