Google warns staff about chatbots
reuters.com
reuters.com
But rather it seems like they're warning against information leaking out?
> The Google parent has advised employees not to enter its confidential materials into AI chatbots, the people said and the company confirmed, citing long-standing policy on safeguarding information.
Is there any example of this happening in the real world? I've never heard of one.
Even if you were to give me access to some google source code free and clear I'm not sure what I could do with it. It's often only useful with the build systems, other services, external documentation, and tribal knowledge contained within the company.
Sure if I had unfettered access to all of the Google source for a product for long enough I might be able to find a security vulnerability or something but random chunks of code pieced together from the inputs to some chat bot?
Just seems like an impractical attack vector.
Keep in mind, this isn't one employee asking one question to a chatbot, this will be tens of thousands of employees typing in multiple prompts each day.
All of those should be included when training users not to divulge secrets, with chatbots just being one more added to the list.
So I think this recent buzz must be more about the model getting trained, than about discrete leaks.
Second, I think the potential exposure via chatbots is a lot higher. Part of the advantage of chatbots over search engines is that you can easily provide more information to refine the output... but that just means that you're encouraged to provide more information.
Certainly training is a factor in this concern though. There are situations where LLMs seem to output training data almost verbatim, which produces a very real possibility that any secrets in the training data could be retrieved by someone who spends time working out a prompt or just stumbles onto it.
If anything, the underappreciated vector of accidental data leaks is the (annoyingly idiotic) combination of two distinct "features" - merging separate address box and search box into an omnibox, and "instant search". This means that whenever you're using your browser bar to e.g. force the text in your clipboard into plaintext, you're also sending all of it to Google (or whatever your browser is). So is the case if you're trying to type in an URL part - it may end up being treated as search before it starts matching something in your history.
As for the rest of things - en/decoders, base64 en/decoders, prettifiers, format converters - I'd say they're in the same category as ChatGPT: you shouldn't be using it with any sensitive data, because there is a very good chance they're collecting it. ChatGPT is at least run by a high-profile company (read: easy target for scrutiny and lawsuits). Random helpful utility for data conversion? There's a good chance it's been made specifically to collect data pasted by careless people.
They're to some extent battle tested where Openai et al are unknown quantities.
Except for this: https://www.refererheadersettlement.com/
Chatbots are explicit in that they use the data submitted to train their models. They pose a much bigger risk.
Your position is that room 641A operated according to contractual agreements? lmao
Corporations do, however, care about their data being handled in accordance with the law, and safe from competitors and malicious third parties (be they cyber-criminals, kids, foreign intelligence agencies, or journalists). Microsoft, as one of many providers of services to companies of all sizes, offers exactly such safe environment, which is also auditable, and comes with an actual SLA that makes them liable for fuck-ups. This is what serious companies want, and it's why they pay Microsoft and use their services - whether it's Office or SharePoint or Azure AI platform. OpenAI itself doesn't offer any of it, which is why any serious company has put it on a top spot of its shit-list.
1) How would you know if your idea or code were ripped off, unless it was an exact bug-for-bug copy?
2) These things have only existed for 5 minutes.
3) Google knows exactly how these models and their interfaces work, and has probably speculatively designed many ways to mine inputs for interesting items.
> Even if you were to give me access to some google source code free and clear I'm not sure what I could do with it.
They're not afraid of you. I wouldn't know what to do with five pounds of gold or a bag full of stolen jewels. Doesn't mean they're not valuable to someone.
https://www.techradar.com/news/samsung-workers-leaked-compan...
It's not really about security in terms of attacking the site.
It's purely about IP protection and potentially leaking user/employee data. Any data you give one of these chat bots can potentially be regurgitated by the bot later.
We have the same kind of policies at my megacorp company. We can't even use things like chrome extensions that will let you paste in data to parse json or things like that.
I'm rather surprised those are considered an attack vector, but I guess the concern is the extension might be running remotely?
Welp, fortunately, I've never had the slightest desire to install a json parser, the Firefox built-in formatting/parsing is quite adequate for my needs. Anything really complex filter-wise I just turn to jq.
If your confidential and clever trade secret technique leaks into somebody else's code (or could have leaked) you may lose trade secret status. If you have a corporate policy to point to, you can claim that any leaks are due to unauthorized activity and thus the trade secret still stands.
On the other hand, if you import somebody else's proprietary code, or GPL'ed code, or buggy code without realizing it, you face different risks. Being able to point to an established policy against using code from these tools helps with the defense against trolls who assert some vague kind of theft.
There is a reason private repos are private
Bonus: the code was pretty tough to understand alone, but with public api documentation it gave a lot of insight to how google works. Sure I wasn’t going to find a buffer overflow or a way into their network at quick glance, but enough time and I can siphon ideas, data, or full chunks of code to do what I want with.
I reported the vulnerability through a colleague and they have since patched it.
Get your popcorn ready, all the silly vulnerabilities of the 90s are coming back again
Most of the concern is going to be on customer data, business relationship data, and confidential internal data. Companies don't want you running your quarterly business review through an unapproved tool. Even though, it's unlikely to happen. It does happen.
Pasting your code into OpenAI is definitely giving people outside of the company access who have not entered into agreements to leak it (intentionally or not).
This has happened to Samsung. Some of their codebase and meeting notes were leaked by CGPT.
The company I work for has the exact same policy for the exact same reasons.
It can be difficult to tell who has been effected by this because you might need to know what to put into the prompt in order to have the leaked data leveraged. So this can go undetected by the victim.
The point is likely more... biz people, don't put in confidential financial & usage numbers when using the writing assistant, don't use code produced by them because it is not trustworthy, and might contain GPL stuff which would infect our code.
https://www.jacksonville.com/story/news/nation-world/2015/02...
Google also buys credit card transaction histories...
I don't think this is true. Quite the opposite, they can do what they want with your data including training models with data you've given them in the past.
Generally companies don’t let other companies play with their data. This may be happening on personal gmail accounts but it’s not happening on corporate accounts.
Source:
https://workspace.google.com/learn-more/security/security-wh...
Considering their chat product is less than a year old, and has had multiple bugs that expose chats to other customers, and has been shown to not be honoring “delete” requests… I think it’s safe to say that there’s evidence that OpenAI is sufficiently negligent even if not nefarious.
Oh that’s to say nothing of the actual data being used in training. Imagine if a bing engineer can just ask a GPT for info on what google is planning and it spits out data it was trained on? Big risk.
OpenAI has clearly stated that they will use user input for training so it's now Google's responsibility to keep their confidential information from ChatGPT inputs, assuming that OpenAI can make unintentional mistakes. What's wrong in here? And your claim is a completely different story, Google and other cloud services clearly state that they will handle user data separately so it's their liability if something goes wrong. And you're trying to insist that they will secretly use user data for whatever they want, breaking their public promises and putting their own business at severe legal risks. So tell me, why would they do that?
> We already know that state level actors are hoovering up civilian data en masse, so this is probably already happening.
Please don't put your own precious conspiracy unless you have plausible evidence. That only harms the credibility of your claim and deteriorate signal to noise ratio in this forum.
This isn't speculative. We know that the government uses Google (and others) to access pretty much all data on the internet:
https://en.wikipedia.org/wiki/PRISM
From the article:
> Internal NSA presentation slides included in the various media disclosures show that the NSA could unilaterally access data and perform "extensive, in-depth surveillance on live communications and stored information" with examples including email, video and voice chat, videos, photos, voice-over-IP chats (such as Skype), file transfers, and social networking details.
^ all of that is going to be going into models.
Yes I think this ongoing thread about AI in corporate environments logically casts a light on re-evaluating how much we trust cloud services (Gmail, Google Sheets, etc).
Perhaps the "meta" is the pendulum is swinging back to self hosted/managed email? sysadmins rejoice!
An arbitrary person can't ask a public gmail endpoint, "what websites is u/bluefishinit subscribed to and what are their other usernames?" and expect to get a reply.
The concern is that the information added to CGPT is getting incorporated into the public models for everyone in the world to access. That's the difference between Docs or whatever and CGPT.
But there is exactly zero control Google has over the confidential data entered in tools and forms that don't belong to them.
Your use of ChatGPT is not protected by any such provisions.
Google is claiming that sending confidential data to third parties is not allowed and will leak that information.
A more apt quote would be "Tobacco companies caution employees against giving cigarette recipe to competitors"
The infinitely bigger point is that this company is gladly pushing out a suite of products that it knows perfectly well to be carry a significant risk of harm. Of course the harm in this case is done to the -companies- where it is used, not the employees. But that's just a minor, and entirely tangential detail.
You should read the article.
I'm sorry that the pervasive toxicity of this industry has rubbed off on you to such an extent that feel you need lash out at others in this way.
I believe it was clear enough from the original snippet that the "smoking" is being done by the companies / organizations that choose to use the products -- not the employees.
This is not news.
Naturally, you could also interpret it as an indictment of its own tool ( Bard ).
Only takes a couple to be a big risk. Being smart and having common sense aren't the same thing; remember, Google's where an engineer fell in love with a chatbot and decided it was sentient. (https://www.washingtonpost.com/technology/2022/06/11/google-...)
That's way higher than I would have thought.
Isn't Grammarly[1] an "AI tool", for instance?
> Alphabet also alerted its engineers to avoid direct use of computer code that chatbots can generate, some of the people said.
It's telling, and says imho a lot about the "quality" of AI generated code.
- they're trying to prevent proprietary knowledge leaking;
- they're trying to prevent dodgy shit they're doing leaking.
My guess is the second as it's highly unlikely that they have any proprietary knowledge that would give their competitors a significant advantage to know. Most large companies' strategies are fairly obvious from the outside and proprietary knowledge leaks through employees moving. It's the dodgy shit they're doing that can damage them.
- The risks (technical, legal) of using AI generated code are not worth assuming.
- They pay a lot of talent a lot of money. Efficiency isn't that important to them. Heck, over-employing talent to prevent them from working for start ups that might compete may be worth it to them.
microsoft is collect chatgpt info like google collects info off searches
They probably use Google at Google anyway.
In a similar fashion if you accidentally typed your password into a non-Google service, it would notice and expire that password. (I can't recall if this was only at the browser level, or if it happened in terminal sessions, too.)
Later they loosened up some of the latter stuff because the importance and weight of passwords themselves was reduced.
(I also started just a few months before they implemented the e-mail "expiry" policy which prevented people from keeping old emails, cuz, y'know, they kept leaving stupid paper trails there that got them in legal trouble.)
you can infer what disease someone might have, what they're doing based on their searches/queries, what they want to buy, etc