Why won’t Google give an answer on whether Bard was trained on Gmail data?
skiff.com
skiff.com
Product manager and CEO _should_ know better. It's very understandable that they don't - and I have empathy for them, but unfortunately they're wrong.
In exactly the same way, there is also a difference between 'is' and 'was', and I think it's totally plausible that a Google lawyer might use that to hide the fact they used GMail data to train AI in the past.
That doesn't mean they did. It only means I wouldn't be surprised if someone proves they did, and that their lawyer used the tense of a response to try to hide it.
4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model.
What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, and I have used this dataset to train my language model. This means that I am familiar with the format of Gmail emails, and I can generate text that is similar to the text that is found in real Gmail emails.”
Now do I think they have done the nasty? I don't know.
Should it be reviewed by an outside team? I think yes.
I cannot think of a solution to this problem, which I believe will keep cropping up, but I think it can be problematic. I think it needs to be prooven true to be safe.
How exactly are you interpreting that statement?
In this case I really think it prudent to assume the worst from Google as they don't really have a positive history for walling off users data, be it personal email or phone meta information.
> "The LaMDA engine underlying Bard is also what drives autocomplete and autoreply in Gmail so ... yeah Bard's training data includes Gmail. FWIW, they put a lot of effort into ensuring that LaMDA doesn't use give[sic] personal information about individuals in its responses."
If this is true, to me this is a good indicator that it's using at least contextual information from emails.
This reminds me of not communicating with Gmail users. Gmail has been evil forever since "personalized" ads.
https://twitter.com/cajundiscordian/status/16382433030356705...
I mean, single handedly destroying the browser market by deceit and abuse of market position in broad daylight, you'd think sooner or later EU or someone would force them to pay and put up a browser ballot on Google.com, but so far, no.
Luckily the French consumer protection agency has at least forced them to implement the cookie question thing almost so we can now reject all right away.
It's not single handedly, it's every developer handedly.
First thing any dev does is install Chrome, and make their normal friends install Chrome, even when -- for example -- it destroys the battery life on their Macbook while breaking a variety of usability conveniences. When pressed, it turns out most of this is cargo culting, they actually haven't turned on developer tools or compared performance or memory handling, much less experienced the remarkable security and privacy integrations* available in co-shipped Safari browser.
Having observed the browser wars since before NCSA Mosaic, I'd argue not Google but a set of "tech influencers" well represented on HN made this happen, not Google.
In the longer context, HN's clamor to force iOS off Safari, the only significant bastion against a monoculture and one ad company's complete grip on visibility of all web use, is shocking.
---
* For instance, various anti-tracker and IP anonymizing capabilities, plus a built-in and cross device syncing password manager that even supports TOTP codes.
Maybe frontend devs. For me FF is invaluable, especially the containers.
No. They are not available in the EU.
At least, I assume this is how Google thinks about it.
If anything, bard not being available in the EU would prove that using the data wasn't necessary for EU's gmail service.
This is not evidence-based thinking and leads to heavy biases.
"Did you send an email to Fred?" -- "No, I'm not sending an email to Fred" doesn't answer the question.
It's not a "conspiracy theory" to have realized that big corps have teams of people to frame their public statements with carefully chosen words to present issues in the best light for them, even if it's deeply misleading.
To me this unambiguously says both that I did not send an email to him and I don't intend to. Is it really ambiguous to you?
I think most native English speakers - if they were reading/listening carefully - would interpret the mismatch of verb tense as an intentional attempt at not answering the question while simultaneously sounding as if it does.
If someone gave me that answer to that question I would repeat the question to them with an emphasis on “did.”
If a corporate PR expert or a politician said that, the conclusion would be very different.
They can say it was not trained on Gmail data because it was trained with Google’s Smart Compose, which it was trained on Gmail data.
Gmail Data -> Google’s Smart Compose -> Bard
It all depends on where you draw the line to stop reporting. Language is a powerful tool of deception.
I can easily imagine people in charge with the mentality of "there's no way that anyone can prove we did it."
It's very improbable, but looking at the "AI integration / product" race it is still a non-zero chance it could have happened.
Everyone else: "But what does that mean? Was it trained on Gmail data or not?"
Find some fairly dense thing in gmail, stick 50% into bard and ask it to complete rest& see how close output is?
There aren't any techniques I know of to prevent it either; when training an image model the recommendation is to dedupe the input so nothing is weighted over anything else, but that's not an absolute defense.
People would be wise to de-Googlify their digital presence as best they can.
The new Gmail ads are atrocious too, IMHO. Peppering ads that "appear" as emails into my Gmail feed seems over the top. I'd rather pay $6/mo, and do, for a single-user Microsoft 365 tenant, which nets me 1TB of cloud storage, in addition to 50 GB of email. I trust Microsoft more than I do Google at this stage.
And they already have access to the whole internet, Gmail conversations would be one tiny part of it.
Also wonder if they got to actually train on github data (considering the Microsoft angle)
Considering all the Github data is on Google BigQuery, probably: https://cloud.google.com/blog/topics/public-datasets/github-...
See also: .zip domain
If something personal - like Gmail/ drive goes out in any of bard responses, it will be the end of bard.
And even if bard contains training data recent enough to include discussions on bard, how would it be able to tell speculation from facts?
The only way I can think of is through deliberate alignment. After training on source data, the model is fine tuned by human curated chat dialogue.
It's safer for companies to not comment at all, unless forced to.
“Google replied to the tweet directly, saying, “Bard is an early experiment based on Large Language Models and will make mistakes. It is not trained on Gmail data. -JQ”.
Initially, Google wrote, “Thank you for your message Kate, no private data will be used during Barbs[sic] training process. We always take good care of our users’ privacy and security.”
"That seems like a clear and heartening assurance. It’s notable, then, that Google quickly deleted that tweet and didn’t amend it with any additional clarification"
That makes me wonder if another model was trained and then pulled in. Example being the Gmail one.
1. "is" versus "was" 2. This tweet is from Google Workspace. It might pertain to Google Workspace gmail only and say nothing about public gmail.
Google has been very explicit that they will not analyze emails of Google Workspace subscribes for ad targeting, but says little about how in-depth they analyze for Gmail users.