Employees are feeding sensitive data to ChatGPT, raising security fears
darkreading.com
darkreading.com
- there’s no way they’re manually scrubbing out sensitive data so its bound to spill out from the training data when prompting the model
- OpenAI is openly storing all this data they’re collecting to the extent that they’ve had several leaks now where people can see others’ conversations and data. We are one step away if it hasn’t already happened from an exploit of their systems (that likely weren’t built with security as the top priority as opposed to scale and performance) that could leak a monumental amount of data from users.
In the most innocent case they could leak the personal info of naive users. But largely if Linkedin is any indication, the business world is filled with dopes who genuinely believe the AI is free thinking and better than their employees. For every org that restricts ChatGPT use, there are fifty others that don’t, most of which have at least one of said dopes who are ready to upload confidential data at a moments notice.
Wouldn’t even put it past military personnel putting S/TS information into it at this point. OpenAI should include more brazen warnings against providing this type of data if they want to keep up this facade of “we can’t release it because ethics” because cybersecurity is a much more real liability than a supervised LM turning into terminator.
Source: worked there after the meltdown.
Where can I, a random employee, get that? I know how to get ChatGPT.
In that case, you wouldn't know to block them until it was too late.
Ultimately either you must watch/block all outgoing traffic, or you must train your people so thoroughly that they become suspicious of everything. Sadly, being paranoid is probably the most economical attitude these days if IP and company secrets have any value.
Well, GP was referring to blocking ChatGPT as a federal contractor. I suspect that as a federal contractor, they are also vetting other people that they share data with, not just blocking ChatGPT as a one-off thing. I mean, generic federal data isn’t as tightly regulated as, say, HIPAA PHI (having spent quite a lot of time working for a place that handles both), but there are externally-imposed rules and consequences, unlike simple internal-proprietary data.
It's a stupid and patronizing position, but corporate IT are sadly incentivised to be stupid and patronizing.
By the way, OpenAI says they wont use data submitted through its API for model training - https://techcrunch.com/2023/03/01/addressing-criticism-opena...
Companies can use Azure OpenAI Services to get around this -- there's data privacy, encryption, SLAs even. The problem is it's very hard to get access to (right now).
Point is, for a meaningful subset of high-value use-cases you don't need to move your important private stuff across any trust boundaries, and it still can be pretty helpful...so just calling that out in case that's useful to anyone...
The fact that the can't do this is the whole reason they have to use ChatGPT.
Besides, it's not like it puts out great code (or even always working code), so I still have to read everything and debug it. And sometimes it writes code that is just fine and fit for purpose and horrendously ugly, so I still have to scrap everything and do it myself.
(And then sometimes I spend 10x as long doing that, because it turns out it's also just plain good fun to grow an aesthetic corner of the code just for the hell of it, too — as long as I don't have to.)
And even after all that extra time is factored back in: it's still way faster and more fun than the before-times. I'm actually enjoying building things again.
And I agree it’s fun. Maybe it’s the simulated social interaction without consequences. I can be completely honest with my robot friend about the shitty or awesome code and no one’s feelings are going to get hurt. ChatGPT will just keep trying to be helpful.
If you know exactly what you want to ask of it, and have the ability to evaluate and verify what it produces, it's incredible what you can get out of it. Sure it's nothing I couldn't have done otherwise... eventually. The productivity it enables is worth every cent.
Easily the best $20 I've spent in ages, they should have run with the initial idea of charging $42.
But holy moly anyone putting confidential information into it needs to stop
If people are using the free version of chatGPT then it’s unlikely there is a contract between the companies and more likely just a terms of use applied by chatGPT and ignored by the users.
to take a stab at your question, though : my cell phone doesn't learn to get better by absorbing my telecommunications; it's just used as a means to spy on my personal life by The Powers That Be. The primary purpose of my cell phone is for the conveyance of telecommunications.
chatGPT hordes data for training and self-improvement in its' current state. It's whole modus operandi involves the capture of data, rather than it being used for that tangentially. It could not meaningfully exist without training on something, and at this stage of the game it's the trend to self-train with user data.
Until that trend changes people should probably be a bit more suspect about what kind of stuff gets thrown into the training bin.
What is different about ChatGPT (if anything)?
ChatGPT seems to be used more like a fast stackoverflow, except people aren't thinking of it like a forum where others will see their question so they aren't as cautious. We're just waiting for some company's data to show up remixed into an answer for someone else and then plastered all over the internet for the infosec lulz of the week.
For every company like yours there are hundreds that don't. People use free gmail address for sensitive company stuff, paste random things in random pastebins, put their private keys in public repos, etc.
Yes, data leaks from OpenAI are bound to happen (again), and they should beef up their security practices.
But thinking people are using only ChatGPT in an insecure way vastly overestimates their security practices elsewhere.
The solution is education, not avoiding new tools.
And, IIRC, pay-as-you-go API requests are explicitly not used for training data. I'm sad GPT-4 isn't there yet - except for those who won the waitlist lottery.
If people were smart and performed according to best practices, articles like this one would not be necessary.
I am honestly confused how people can use this thing with the interface OpenAI runs. The app has been near-unusable for me, for months, on every device I tried it on.
And what sort of understanding do you have with the alternative frontends/integrations about how they handle your API keys and data? This might be a better solution for a variety of reasons but it doesn't automatically mean your data is being handled any better or worse than by openai.com
Some of those are document tools working on language / knowledge. Others are infrastructure, working on ... whatever your infra does, and your infra manages your data (knowledge).
If you read their data policies, you'll find they are not the same.
I am unsure if the so called AI can think in models but so far, not but still an impressive assisting tool if you take care of its limitations.
Another point where it lacks is in logic, my daughter has a lot of fun with the book "what is the name of this book?" but she was struggling with the "map of baal" explanation, her the answer was a certain map, yet the book had another answer, I had a third one as I interpreted a proposition. I never got an answer without a contradiction in chatgpt reasoning, and the book had been mistranslated to French so one of its propositions was changed (C, both A and B were knaves) but not the answer.
> I am unsure if the so called AI can think in models but so far, not but still an impressive assisting tool if you take care of its limitations.
I don't know. I'm using it for exactly that ("here's a problem, come up with a data model") and it gives a great starting point.[0]
Not perfect, but after that it's easy to tweak it the old-fashioned way.
I find its data modelling capabilities (in the domain I'm using it for - API services) to be rougly on par with a mid-level developer (for a handwavy definition of "midlevel").
X and Y are not alike, and should not be compared. X is a benefit to you(r employer), whereas Y is a risk to the customer who has entrusted you with their data.
I do and if it could be leaked through ChatGPT I would have it blocked.
Risk isn't a single dimension, it's a combination of exposure (chance of happening) and impact (how much will you lose)
Nothing you’ve said negates anything OP said. It’s simply an elaboration wrapped in elitism.
And yes, if the bypassing the block is combined with disciplinary action, it does work. It’s not worth getting fired over. This is likely what heavily regulated industries like financial services and defense are doing.
It's a fairly standard practice. I wouldn't associate it with overreaching surveillance.
Hey, they need someone to proofread their War Thunder forum posts to make sure they're using correct spelling and grammar when leaking classified info. ;-)
(Ref if you don't get the joke: https://taskandpurpose.com/news/war-thunder-forum-military-t...)
As the leaks ware more inadvertent.
I was under the impression OpenAI weren't using questions as training data for future models. I recall Sam Altman saying they delete questions after 1 month, but I can't locate the source for that.
This is still a serious data loss risk.
MS also "helpfully" asks you if you want to use that enhanced grammar check in MS Word(as far as I have seen, might be there in other office products too). I cannot imagine sending all my documents to MS. But I am not sure most users will realize what is happening.
All these companies offer helpful services but are hovering up data and no one knows the consequences yet. It feels like ChatGPT is just one symptom of a bigger problem.
I think MS Edge is getting even worse about this, with the big fucking Bing icon in the corner and making it impossibly hard to get rid of it.
They will go to OpenAI directly and do stuff because "they want it now" and they don't understand why not.
Microsoft is already building it into Office 365 with the same enterprise-grade agreements so it might get easier that way.
Do you have a source for this? I know some people have claimed to see others' data, but I haven't seen any evidence that that's what's actually being seen, vs LLM hallucinations. OpenAI claims, and I can't imagine they're lying, that the training data is fixed and ends in 2021, so I don't see how it would be possible for user prompts to be leaking into output, absent a massive and very unlikely bug (compared to the much more likely AI hallucination explanation).
Some kind of concurrency bug in a library they were using to retrieve cached data from Redis led to this leak.
> We took ChatGPT offline earlier this week due to a bug in an open-source library which allowed some users to see titles from another active user’s chat history. It’s also possible that the first message of a newly-created conversation was visible in someone else’s chat history if both users were active around the same time.
Do you also block pastebin? Anything else that has a web form? How is ChatGPT special compared to any other service on the Internet where people can paste data in a form?
I mean... I see the problem, but I think one needs to realize that it's a far more generic problem that has basically nothing to do with ChatGPT and AI. If people paste confidential data into random webpages that's of course bad. But if you block ChatGPT because you fear that, it means you expect that people might do that. And then your problem is not ChatGPT, but lack of awareness what is confidential data and what to do with it.
I don't think the idea that you can give people access to the www and at the same time preventing them from putting things in forms can be done. That's simply not how it works. And if you're blocking access to a few services where they might do that, well, they have a million others, and you're deceiving yourself that you've done something.
pastebin and indeed most things that has some sort of public webform is blocked in all the companies I have worked with.
It is probably a losing battle though, as it is very hard to block everything without default deny.
Paradoxically, maybe GPT could be used to veto websites on first access :)
Search engines too? And these days, that means web browsers, because the (IMHO stupid) idea of combining address and search bars into one means everything you type while trying to open a website gets leaked to some party (most likely Google).
Corporations constantly put their most sensitive data in 3rd party tools. The executive in the article was probably copying his company strategy from Google docs.
Yes, there are good reasons for concern, but the power of the tool is simply too great to ignore.
Banning these tools will go the same way as prohibition did in the US, people will simply ignore it until it becomes too absurd to maintain and too profitable to not participate in.
Companies which are able to operate without these fears will move faster, grow more quickly, and ultimately challenge companies restricted to operate without.
Now I think the article should be a wake-up call for OpenAI. Messaging around what is and what is not used for training could be improved. Corporate accounts for Chat with clearer privacy policies would be great and warnings that, yes, LLMs do memorize data and you should treat anything you put into a free product on the web as fair game for someone's training algorithm.
* Their contractors can (and do!) see your chat data to tune the model
* If the model is trained on your confidential data, it may start returning this data to other users (as we've seen with Github Copilot regurgitating licensed software)
* The site even _tells you_ not to put confidential data in for these reasons.
Until OpenAI makes a version that you can stick on a server in your own datacenter, I wouldn't trust it with anything confidential.
OpenAI just hasn't started to try adding privacy and security yet.
https://help.openai.com/en/articles/5722486-how-your-data-is...
> OpenAI does not use data submitted by customers via our API to train OpenAI models or improve OpenAI’s service offering. In order to support the continuous improvement of our models, you can fill out this form to opt-in to share your data with us. Sharing your data with us not only helps our models become more accurate and better at solving your specific problem, it also helps improve their general capabilities and safety.
> When you use our non-API consumer services ChatGPT or DALL-E, we may use the data you provide us to improve our models.
It's also in the FAQ: https://help.openai.com/en/articles/6783457-chatgpt-general-...
> Will you use my conversations for training?
> Yes. Your conversations may be reviewed by our AI trainers to improve our systems.
Google tries hard to sell you on their auto-answers for emails ('smart reply'), wonder how those got trained...
There's a huge difference between trusting a third party service with strict security and data privacy agreements in place vs one that can (legally) do whatever they want with your corporate data.
Google Workspace solves the issues of data privacy both by having extreme user data & datacenter access controls[0,1], a robust terms document that details how data is collected and used[2], and enterprise customers can access an audit report that details what and when things are accessed by Google employees[3].
0: https://storage.googleapis.com/gfw-touched-accounts-pdfs/goo...
1: https://workspace.google.com/security/
Or the fears are real and companies that operate without them will be exploited, or extinguished for annoying their customers.
And on the other hand, if your company doens't use Github etc due to security concern, it's a very good sign telling you need to ban ChatGPT too.
Sorry, but I really struggle to see how a non AI company will actually become more profitable simply by getting their employees to use ChatGPT. In fact, the more companies that use it, the more demand there will be for "human only" services.
If you’re an artist, just send all your work to DALL-E. Why have money or fame?
The enterprise version of gmail was just an additional step to instill trust. In practice it is still a decision based on trust rather than physics. An "enterprise version of privacy guaranties" for Huawei 5G modems or tiktok apps would not make governments suddenly happy with the risk model where sensitive data would have a minor risk of ending up in China.
And corporations have strict agreements with their providers. They are even required to in many cases due to GDPR and the likes. Users connecting to ChatGPT on their own accounts bypass this.
The original Gmail TOS explicitly stated that they scanned the content. They only stopped for the rollout of Gsuite.
If employee sets random GMail account that is not covered by agreement that is personal account. Sending company data to personal email account might be grounds for firing person.
Setting up some account at random with OpenAI and putting company details like customer names or else there is data breach.
Companies will let people use the tools - but it is not like one can start setting up random accounts without approval from management. Of course there are different types of companies with less or more red-tape.
Which is why a private option is so critical. To not fight against human nature, means providing an ability to use the tool in a safe way.
Unless your company really has nothing to hide, it's easy to accidentally dump a company secret or an API key in a chat session. Of course if everyone is aware of this and constantly careful then you may be OK.
If you're copy pasting API keys or such into ANYTHING, you probably shouldn't be a programmer to begin with.
It's like people who use root account key/secret credentials in their codebase. It's not AWSs fault you got a large bill or got hacked, its because you're dumb.
And if you think that there is no special code then you're wrong.
Lots of code expresses buisness strategy that is a competitive advantage/ sensitive.
My org has done a risk assessment and accepted the risk of using such tools; arguing there is no risk is short sighted.
It literally is exactly that. You don't think the code a business creates is "business data"?
Not my place legally or ethically to share code with 3rd parties that I've been paid to read and write.
Your code is not special, but customers data may be. Also, some companies needs to comply to various certifications, and proven leak of source code that was put into some third party tool may be a reason to revoke such certification. Which can cause a serious financial harm to a given company, as it can lead to ex. losing government clients.
This is just the tip of the iceberg.
Uploading code to ChatGPT can be done by trainees.
Productivity isn't everything.
FWIW I don't think the employee should be fired for this or anything, if anything a company could embrace these new technological advances and provide training on using ChatGPT in a more secure manner(ie don't paste your customer's PII into a prompt, etc...).
With this particular employee, using chatGPT has not increased his productivity or the quality of his work by any noticeable degree.
> I don't think the employee should be fired for this or anything
The problem isn't using the technology. The problem is sharing confidential information with an unapproved entity. That is specifically and clearly spelled out as a firing offense, for pretty obvious reasons.
Even if some people feel that it's an overly tight policy, it's a the stated policy and the company has every right to put and enforce whatever rules it wishes about the use of its own data.
2.) Of course, you can tell people not to use it. Unlike people at SV companies apparently, people in government and government contractors accept restrictions like not having phones in secure labs all the time. Start firing or even prosecuting people and people will discover very quickly they don't really need some tool.
And, yes, private versions of this sort of thing helps a lot.
Uh, why can't you tell people not to use it...? If security is that important for your company, of course you can tell your employees which tools to use.
A fun fact: in many areas of TSMC, smart phones are banned. No one says "you can't just tell people not to use smart phones."
This does not surprise me at all. What I want to know is how they enforce it.
Unless they have something better than "fear of somebody seeing you using the smartphone", it isn't getting enforced. If they do have something better I want to know what.
No, I'm not joking. One of my high-school classmates works as R&D there. They ask you to pass through a metal detector gate, take away your phone if found, then give your a company phone for emergency call only. It's that strict (at least for R&D. Probably not for management and others).
This is what I'm interested in.
I get it, the guys in bunny suits will probably tolerate being groped and wanded every workday for the rest of their career. I have a hard time beliving the scientific staff and executives tolerate that.
They put a special sticker on all of your cameras and inspect if it is still there on the exit.
We've granted some exceptions to the team responsible for determining how to stand up something internal. Lots of shooting in the dark going on here, so I figured we would need some divulgence of our IP against public tools to gain traction.
There is way more context to our business than can fit in 4/8/32k tokens. Even if we could fit the 32k token budget, it would be very expensive to run like this 24/7. Fine-tuning a base model is the only practical/affordable path for us.
> ChatGPT is now available in Azure OpenAI Service
https://azure.microsoft.com/en-us/blog/chatgpt-is-now-availa...
That's all I use ChatGPT for. I don't need it to be able to write poetry.
./main -m ./models/7B/ggml-model-q4_0.bin -t 4 --temp 0.7 --top_k 40 --top_p 0.5 --repeat_last_n 256 --repeat_penalty 1.17647 -n 1024 -p $'Here is a handy short form regex to validate an email address: '
Here is a handy short form regex to validate an email address: ^([a-zA-Z0-9_.+-]+@[a-zA-Z0-9.-]+\.[a-z\.]{2,6})+$ The first character can be either uppercase or lower case. The second group must contain at least one letter and the third group may contain any number of characters (up to 5). The last part ensures that it ends with @ followed by two more letters separated by dots. If you want to make sure that your input string contains only valid characters for emails then use this regex instead: \A[\w.]@[\w.]\.\w{1,4}\z
And the second one is confusing me. It seems to be matching a single character only for the initial portion?
Neither of them seem good, and especially the last one.
And the way it describes both seems off as well. I would have to say it brings more harm than good based on that.
What it emitted accepts a large number of invalid addresses (due to things like not checking dot placement, and the inexplicable (…)+ wrapping around the entire thing), and doesn’t accept a large number of valid addresses (some comparatively esoteric, like local parts containing any of !#$%&'*/=?^`{|}~ or IP addresses for the domain name, and some very reasonable, like TLDs of more than six characters, or internationalised TLDs even in Punycode form).
The description it emits does not match the regular expression at all well, either.
The second regex it emits is even worse than the first, unnecessarily uses PCRE-specific syntax, and is given with a nonsensical description. (Note: the asterisks got turned into italics, backslash-escape them here on HN. With this fixed, the regex was \A[\w.]*@[\w.]*\.\w{1,4}\z.)
> on the very surface scan seems OK
And there’s the danger of this stuff. As a subject-matter expert on regex and email, I glanced at the regular expression and was immediately appalled (… quite apart from the whole “here we go again, this is certain to be terrible” cringe on the prompt). But it looks plausible enough if you aren’t.
email_pattern = r"^(?=.{1,256})(?=.{1,64}@.{1,255}$)(?=\S)(?:(?!@)[\w&'+._%-]+(?:(?<!\\)[,;])?)(?<=\S)@((?=\S)(?!-)[A-Za-z0-9-]{1,63}(?<!-)\.?)+[A-Za-z]{2,19}(?<=\S)$"
Also completely spitballing, I expect that a big chunk of OpenAI's 'secret sauce' is simple processing layers above and beyond the model. If you input gibberish to llama does it give you an output? If OpenAI is artificially tokenizing inputs (as opposed to just sending inputs straight to the software), it would both dramatically limit the input domain, thus improving output tuning, as well as give "it" the ability to say when it doesn't know something. I put "it" in quotes since that response would not becoming from the LLM, but from the preprocess tokenization system returning an error code in natural language.
I think there's some weak indirect evidence for this in the service itself, since incoherent inputs are instantly rejected, whereas even simple queries take dramatically longer to output even the first word. It's like the input is not even being sent to the LLM software for processing.
I've been debating the idea of building tiers or layers of models to accomplish the same.
It very well could be that this go/no-go pre-processor is simply another ML model trained on a binary classification task. Stack a few of these and you can wind up with some interesting programming models.
One developer opening a folder in VSCode with Copilot enabled aaaaand it’s gone. You never know what part of the folder left your building.
You use Windows, VSCode, etc, all of this has access to your code.
And Windows and VS Code don't upload your data to Microsoft unless you choose to do so.
The problem is the verification of this of course.
If we want to be intellectually honest with ourselves, we can either be fearful and have a plan to contain data from ALL of these companies, OR, we address the risk of data leaks through bugs as an equal threat. OpenAI uses Azure behind the scenes, so it'll be as solid (or not solid) as most other cloud-based tools IMO.
As for your data training their data: OpenAI is mostly a Microsoft company now. Most companies use Microsoft for documentation, code, communications, etc. If Microsoft wanted to train on your data, they have all the corporate data in the world. They would (or already could!) train on it.
If there's a fear that OpenAI will train their model on your data submitted through their silly textbox toy, but NOT through training on the troves of private corporate data, then that fear is unwarranted too.
This is where OpenAI should just get a "corporate" tier, charge more for it, and is basically make it HIPAA/SOC2/whatever compliant, and basically do that to assuage the fears of corporate customers.
2. Azure infrastructure still likely has far better security/privacy by virtue of all their compliance, (HIPAA, FedRAMP, ISO certifications etc.) than whatever startup-move-fast-ignore-compliance crap OpenAI layers on top of it in their application layer.
There is zero fear. OpenAI openly writes that they are going to use ChatGPT chats for training. On the popup modal they show you when you load the page. That is not a fear, that is a promise that they will do leak it whatever you tell them.
If i tell you “please give me a dollar, but i warn you you will never see it again” would you describe your feeling over the transaction as “fearful that the loan won’t be repaid”?
While this is not a new problem -- employees share sensitive data with Google all the time -- the data leakage will be more clear than ever. With ads-based tracking and Google search, the leakage was very indirect. With generative AI, it can literally regurgitate memorized documents.
The security risk goes beyond data exfiltration. Folks are already trying to teach the AI incorrect information by spamming it with something like 2+2 = 5.
Data exfiltration + incorrect data injection are super underrated risks to mass adoption of generative AI tech in the B2B world...
Worst case the CISO gets fired and then they all play musical chairs and end up in new roles.
Heck, even Lastpass, ostensibly a security company, doesn't seem particularly affected by their breach.
My point is, especially with ChatGPT, where it can reasonably 10x your productivity, most people will be willing to take the risk.
Endpoint security software is a security risk too.
Well I have so speak for yourself.
> once you have a lot of your sensitive data outside in a third party system you have lost control.
Every time you search “how do I do this with this software stack” in Google you are leaking data which is nominally sensitive to a third party system. Every time a technical staff member goes to stackoverflow without obscufating their IP address they are leaking sensitive data about what software stacks a company uses. Let’s not even get into people posting their resumes on LinkedIn or cloud services in general.
The goal of security is not to stop all data leakage, it’s to stop the leakage of certain high value data, and LLMs can aid in this end if you use them intelligently and avoid leaking any high value data to them and only feed them with low value data. Attackers are not going to have any qualms about using LLMs both to come up with attacks and as part of attacks. Many many people are in situations where using LLMs to advance in security maturity as quickly as possible is more than worth the risk incurred. Don’t win the battle, win the war.
There is obviously a very big gap between searching for information and providing your internal code or documents to a third party. One reveals only your search terms, the other gives an attacker your actual proprietary information.
LLMs are ground breaking and have enormous potential. I am not saying that they should not be used. Only that there are huge security issues when employees of most companies post confidential information to third parties.
No way I’m giving Google any of my data! I will use 5 different browsers in incognito mode and never log in.
To ->
Sure I will login with my name and email and feed you as much of my most personal thoughts and data as I can dear ChatGPT!
This situation is much dumber than that. ChatGPT is very clear that you shouldn't give it private data and that anything you type into it can/will be used for training.
Google is nowhere near that level of transparency.
Because both types of people have always existed. Heck, lack of vigilance among the ancient Greeks is what put the Trojan in Trojan horse.
Sometimes couriosity beats caution.
Even if they are not doing it now(?), what makes you think that they will not do so in the future? It's not like your data has an expiration date.
I see no reason to think OpenAI would leave that money on the table.
If it uses user data to train their models other users could ask "Show me the code for gmail spam filters", and if it was trained on engineers refactoring that spam filter in ChatGPT chances are it would give you the code. If that doesn't count as "selling user data" I don't know what is. They not only sell it, they nicely package and rewrite it to make it easy to avoid copyright claims!
In other words, having been burnt once by touching a flame, the conclusion these people draw is that the problem was with that particular flame and they're fine with reaching for a different one?
But it doesn't, does it? It sells the fact that it knows everything about everyone and can get any ad to the perfect people for it. It's not going on the open market and telling people I regularly buy 12 lbs of marshmallow fluff and then use it in videos I keep on my google drive.
No, Google takes money to present ads to people of different demographics, and uses your data to do that. It doesn’t sell your data, which is, in fact their competitive edge in ads – selling your data would be selling the cow when they’d prefer to sell the milk.
Last time I looked (maybe things have improved…?) Grammarly would automatically attach itself to any text box you interacted with, and immediately send all of your content to their servers for processing. How this software gets past IT departments is a mystery to me.
I am working with customers that are looking to train a homegrown LLM that they host and have blocked access to ChatGPT.
https://www.datanami.com/2023/03/24/databricks-bucks-the-her...
Any company that doesn't want to feed data into ChatGPT should need to proactively block both ChatGPT and any website serving as a wrapper over it.
I agree this would be a good move, but it's going to be harder and harder to do that definitively.
I shudder to think what they are doing now.
I had an interview awhile ago at a place where during the phone screen "they can't talk about their tech stack in detail" so I looked on linkedin and figured out their entire tech stack before on the onsite interview. Come on guys, according to linkedin, you have an entire department of people doing AWS with Terraform and Ansible, you don't have to pretend you can't say it in public.
https://learn.microsoft.com/en-us/azure/cognitive-services/o...
If it works well it will be a big deal.
It was quite amazing what people would submit unprompted, so I'm not at all surprised that people would feed sensitive data into ChatGPT. The next cycle will be that ChatGPT gets - surprise - trained on that data, and may start using fragments of it - which may well still be sensitive enough to cause trouble - as its output.
Don't paste confidential information into a textbox, in fact don't trust anybody or any company with your confidential information unless there is a strong contractual relationship backed up by penalties if it gets broken. And even then: the ultimate responsibility is yours, you may be able to recover some $ for damages but your reputation may well be toast.
Gmail has a vested interested in keeping any knowledge it gains about you secret - it's competitive advantage is knowing more about you than anyone else does.
ChatGPT's strength is its ability to clearly communicate the knowledge it has (including training data it gains from people it interacts with) to give you good responses.
Most corps, companies store a lot of internal data with corps like Google, Microsoft, Amazon and others.
Hahahahahahaha
There's really not much you can do here. This is complete lack of very basic common sense. Having someone like this in your business, particularly at the executive level, is a liability regardless of ChatGPT.
For example, I was having some issues with my LTO-6 drive recently, and I had to finagle through a bunch of arcane server logs to diagnose it. I had the idea of simply copypasting the logs into ChatGPT and having it look at them, and it quickly summarized the logs and told me what things to look for. It didn't directly solve the problem, but it made the logs 100x more digestible and I was able to figure out my problem. It made a problem that probably would have taken 2-3 hours of Googling take about 20 minutes of finagling.
I'm not doing anything terribly interesting or proprietary on my home server, so I didn't really have any reservations sharing dmesg logs with it, but obviously that might not be the case in a company. Server logs can often have a ton of data that could be useful for a competitor (whether it should be there or not), and someone not paying attention to what they're pasting into ChatGPT could easily expose that data.
It really does feel like YC was a plot to fund the harvesting of all with an OpenAI climax. Not a serious conclusion I have, its just funny to watch it unfold, as if nobody even cares about the optics.
Yammer pre-Microsoft and nowadays Blind — lots of “insider” information seemingly posted.
As usage goes up the target size, and opportunity cost, both go up.
Though RDIMMs on eBay are even cheaper than UDIMMs (just over $1/GB) and Broadwell-era Xeon workstations aren't that expensive if you want to run the unquantized version.
And Github Copilot will gather that tiny amount of data to who knows where.
I guess another way to say that is, OpenAI (or another service provider with a better security track record) could broker this service in the cloud, with guarantees around not using the session data for RLHF, not storing session data, stronger auth (OpenAI has had a couple of incidents that show that they have pretty lax security in their backend), etc. and could make a killing selling or re-selling ChatGPT to businesses.
I'm hoping OpenAI will implement something like this on their end soon like data monitoring apps do (sentry etc), but if they don't client-side is an option.
- Can you summarize that product roadmap?
- Can you add comment to this piece of (proprietary) code?
- Please extract the birth dates from this employee list
I could go on and on and on
> You can request to opt out of having your content used to improve our services at any time by filling out this form (https://docs.google.com/forms/d/1t2y-arKhcjlKc1I5ohl9Gb16t6S...). This opt out will apply on a going-forward basis only.
It goes to a google form, which is I guess better then them building their own survey platform from scratch that may have more vulnerabilities.
And IP pirates banned from using the internet. Actually, that one I remember I voted against my local MP after they passed a law to make that the norm.
We don't yet have a social, let alone legal, norm for antisocial use of LLMs; even with social media, government rules are playing catch-up with terms of service, and Facebook is old enough that if it was human it could now vote.
So, yes, likewise computers/internet/social media, being banned from an LLM if it's a monopoly is going to seriously harm people.
But that is likely to be a big "if". The architecture and the the core training data for GPT-3 isn't a secret, and that's already pretty impressive even if lesser than 3.5 and 4.
Funny. This nobel prize winner raises an interesting question:
If your AI is so great at coding, why is your software so buggy?: https://paulromer.net/openai-bug/
And the thumbs up/down are there on the chat interface because it's partly trained by reinforcement from human feedback.
GPT uses the internet to connect to users, but rather more importantly chatGPT in particular has a layer on top of GPT which is trained from human feedback.
Keywords search "RLHF".
That feedback mechanism is, if anything, becoming more detailed as time passes, so I must infer that it's still considered highly important, probably even for the 3.5 model.
That it does any of this is the specific reason for the story you're commenting on, and why putting data into it isn't like putting the same data into e.g. a Google Docs spreadsheet.
Clearly they store information from ChatGPT, but if I hit their chat/completions API with the same exact question do they store that as well?
Not to say it isn't a problem, but it's not an AI/chatGPT specific one.
Also I guess chat with Microsoft teams or slack is ok.
Imagine if everyone knew your inner secrets just by looking at you at the bar..
Why is it different online? i have no idea.. well i kinda know but.. oh well.. we deserve it i guess
i) In this world, there are very few people whose private conversation is worth anything to anybody (celebrities, journalists -- so around 10,000 people)
ii) A tiny tiny %age of information is truly secret (mostly private keys).
iii) Business strategies are mostly a result of execution, not any 'trade secrets'. Meta will succeed because it has executed it's metaverse strategy, not because they kept the metaverse strategy secret.
People who take risks and not care about irrelevant details (just like how they took risk with internet shopping, cloud, SaaS) will win. Losers like the ones who thought AWS will steal their data will be left behind