Following pushback, Zoom says it won't use customer data to train AI models
darkreading.com
darkreading.com
Paying customers absolutely HATE the idea of their data being used to train AI models without their permission - this Zoom story is just the latest example of that.
Companies that try to do this will get burned. Zoom just got burned really badly, and I personally don't think they actually intended to even do this - they just didn't make it clear enough that they were NOT going to do it, which sparked a PR nightmare firestorm for them.
I think the incentives for companies are very much the other way round: if paying customers hate this, then the incentives are NOT to do it.
Except to improve services (ML training!), Advertising (We'll sell your data to advertisers!), and by government order (pick your favorite three letter agency.)
Alternatively: They'll just use your data anyway and not tell you about it. How is anyone going to prove their data was used as the source? It's like all the NDAs people sign when they join a company and they pinky promise not to use it at the next job where they land a big fat raise and promotion... suuure they aren't going to take what they've learned and improve upon it to try and get more promotions and raises in the future.
It can't be stopped.
Maybe. Maybe not. Hard to tell without trying.
I applaud the EU and California for giving it a go with their data protection laws. I really hope their crackdown on this stuff is effective.
Uh no. An NDA would cover proprietary intellectual property, not tools everyone else also uses. Unless you're now working for the previous employer's competitor, it's unlikely that proprietary tech would ever be used. Working for competitors and partners is usually also forbidden for some period of time after leaving.
1) whistleblowing 2) compliance audits (think soc2)
They're asking for evidence that your admin accounts are reviewed for permissions needed each quarter, that you are doing your DR tests, that you are following documented change management processes, etc, not asking to look at what your running containers are doing.
There is absolutely no way that any compliance audit I can think of would or would even attempt to uncover that kind of info.
We can agree that Zoom did a terrible job of rolling out their new terms, regardless of what their intention was. What other companies will learn from this is to improve the roll out.
Once local/private inference becomes more viable, there will be even more of an incentive for the companies who store unencrypted data to use it as a competitive advantage.
Paying customers will be the ones to resist for the longest maybe but that frog too will be boiled eventually.
I'm assuming it's safe because they make a big deal about compliance. But on the other hand MS have a huge incentive to obtain data for AI since they are going all in on it.
Also, modern Microsoft’s _whole thing_, more or less, is “we are your trusted enterprise partner who definitely won’t do bad stuff with the data you put in our cloud services”. They are unlikely to throw that away for a bit of flavour-of-the-month AI boosterism. Note that they’ve recently released a private ChatGPT thing; they can’t credibly acknowledge the problem with one hand and exploit it with the other.
This makes it difficult for B2B companies like Zoom to use customer data for AI training.
Now we'll be forced to use Teams for online trainings after pretty much universally using Zoom since the pandemic. Our customers are gonna love that.
It's above my pay grade but I wonder if we've already signed our data over to MS for a certain price, with that stipulation.
just call it "fair use", like OpenAI and GitHub
but then they would not be able to do this https://news.ycombinator.com/item?id=37100140
But I can see them being okay with bringing copyright back to 10 years or something, they make most of the money from publishing a new television series early on. If they made it ten years, paying for streaming old television series doesn't require paying the copyright owners anything so it's pure profit for the streaming service.
Of course people can go elsewhere, but if they can only get things made in the last 10 years on some particular service, they'd just use the same to view old things... and pay the website without the creators getting anything.
So every distributor wins because they can all offer the older material as a pure profit for themselves and an enticement to join the service (on top of the actual unique material from the last 10 years).
I absolutely believe that companies should have the freedom to change their ToS moving forwards, and that "promises" to "never" change a ToS are worthless (your example is a perfect reason why).
BUT, I simply don't see how it could ever be fair to use data retroactively. If you change your ToS, you should only get to monetize new user data going forwards. It seems like a basic legal principle.
Is there any existing law/precedent that suggests this is already the case, i.e. that such a company can be successfully sued but that people don't usually try? Or do we need new laws around this, and are there any government reps pushing for this?
I just don't love the idea that my dinky little app has to ask every customer every time I add a new feature significant enough (debatable) or different enough (debatable) that uses their data in a way either I or they didn't anticipate (debatable). God forbid I try to monetize it (debatable). 'Control over your data' is meaningless in our current paradigm and I'll rue the day something like GDPR comes to the US in a meaningful way. No wonder the EU can't build.
As for this specific article, Zoom's (rightfully) getting heat for this but I don't blame them or any company for exploring how they can monetize every last morsel of data. In zoom's case (and many enterprise software companies), customers are paying a shit load of money and they didn't sign a contract and consent to give data for training an LLM.
I absolutely do, if it's customer data that the company previously promised not to monetize. It's not their data to do with as they please, after all.
But the tech sector has fallen very far in terms of ethics so no company can be trusted. It's just a shame. The public views our industry in a very, very poor light and that view is 100% earned.
A TOS is a contract. It literally stands for "Terms of Service." Meaning, you give me money and here are the terms under which I will offer you the service you are paying for. How enforceable that "contract" is depends on a ton of things, differing in various jurisdictions (law is complicated), but it is - at the end of the day - a contract.
So I don't know how actionable it is, but the OP said that the company considered changing their TOS for currently active users. That could, in theory, be breach of contract and the customers might have a claim (again IANAL).
[There could have also been a clause in the TOS saying that they could change the terms at any time for any reason - though I suspect in many if not most jurisdictions, that would make the entire contract unenforceable].
In your case, don't make [potentially] contractually binding promises that you can't or don't want to keep.
The TOS are generally broad enough from the start that you can do anything you want with user data as is necessary to provide product features. Nobody's updating TOS every time they add a new feature.
Realistically, this is specifically about situations around selling data to third parties, and/or training for AI that is not related to product features. (There's a big difference between Zoom using chats for building LLM's, versus Google training on Gmail messages to build Gmail autocomplete.)
That was unnecessary.
> I don't blame them or any company for exploring how they can monetize every last morsel of data.
That's how a company works: try to do everything they legally can to make as much money as they can. Society has to decide of the framework into which companies optimize, and that is materialized with laws that the companies must follow. In the EU, there is a tendency to believe that users have a right to some kind of privacy.
Of course, this constrains what companies can do, and you could say "no wonder the EU can't build". I just call that cultural differences. In most countries in the EU, people don't have to start a crowdfunding campaign when they go to the hospital, because they actually have some kind of social security. I am all for GDPR.
No, companies don’t need to be like that. This is a meme that needs to die. Companies can have a set of values (principles) and act according to those principles. Any investors can be told ahead of time the principles by which the company operates, and if they don’t want to buy stock on that basis, they’re welcome to stay out.
Bryan Cantrill has had some excellent rants about this over the years. Eg: https://youtu.be/bNfAAQUQ_54 . His take is that money for a company is like fuel in a car. You don’t go for a road trip (start a company) because you want to get more fuel. You go because there’s some place you want to get to. And fuel (money) is something you need along the way to make your journey possible.
Don’t let sociopathic assholes off the hook. They aren’t forced to be like that. They’re choosing to abandon their ethics and common decency. Everyone would be better off if this sort of behaviour wasn’t tolerated.
Well, they don't need to. But the people at the top make more money if they are. And they are not at the top because they have principles: they are at the top because they want power or money.
> Companies can have a set of values (principles) and act according to those principles.
I would love it, but I just can't buy it. Like at all. How many big companies do you know where the executives don't get a much higher salary than the employees? Humans can't help it: if they are in a position of power, they will think they are worth more.
> Any investors can be told ahead of time the principles
IMO, if you have principles, you are not an investor. And investors want to get ROI, which is more likely from companies that don't have principles.
> His take is that money for a company is like fuel in a car.
Sounds exceedingly naive to me :-). The driver does not get fuel at the end of every month.
> Everyone would be better off if this sort of behaviour wasn’t tolerated.
Yes. We need laws, set by the society. We need the people to understand that they will never be one of those rich executives, and to vote for laws that prevent them to become indecently rich.
Monetization by adding paid features falls well within those boundaries. Monetization by selling user data to whomever will buy it does not.
I'd really love to have a GDPR specifically for people like you who feel entitled to do whatever they want with collected data. I'd love to have had it when reddit decided to charge outrageous prices for the API.
Adults realize other adults do what benefits them.
But most software does not honor those licenses, and nobody cares. Enforcing such a law takes money.
I guess at least the GDPR can be enforced, to some extent, with Big Data. It seems like the fines are usually ridiculously low (they don't seem like an intensive for the company to change anything), but that's better than nothing.
That's a lot of compute power to waste on it, but I would guess that that's what bot networks are going to be used for in the future (or already are, right now, if they're done mining bitcoin).
I wonder if Teams would face similar uproar, assuming that bit isn't already in the ToS.
Maybe, but Teams is very good at ignoring uproars. While I assume that there are people who feel differently, everybody I know already loathes Teams and only uses it when their employer forces them to anyway.
People use it only because someone up the chain sees it's included in Office and they're like "we're not paying for something else if we get this for free". I hope Slack and others bring MS to court to stop this, this is exactly what happened during the browser wars.
Aside from the usual (and understandable) MS hate, I don't see the problem. Features and performance? Nothing else is better, most are worse. Cisco is a mess, Skype is horrible, Zoom lies about their security, etc, etc
Well, we have very different experiences with Teams. It's perhaps not the worst, but I think it's pretty bad in the sense that it's painful to use and gets in my way.
Zoom's TOS Permit Training AI on User Content Without Opt-Out - https://news.ycombinator.com/item?id=37038494 - Aug 2023 (35 comments)
How Zoom’s terms of service and practices apply to AI features - https://news.ycombinator.com/item?id=37037196 - Aug 2023 (177 comments)
Ask HN: Zoom alternatives which preserve privacy and are easy to use? - https://news.ycombinator.com/item?id=37035248 - Aug 2023 (16 comments)
Not Using Zoom - https://news.ycombinator.com/item?id=37034145 - Aug 2023 (194 comments)
Zoom terms now allow training AI on user content with no opt out - https://news.ycombinator.com/item?id=37021160 - Aug 2023 (510 comments)
They do automatic captioning/transcription of meetings, so there is a model for that; they do automatic background blur/cutout, so there is a model for that; they are probably working on a "meeting summarization" product for that.
Those are features that people love and use all the time. I would be curious to know how anyone expects Zoom to improve on these features without collecting data from real users on the platform.
I think it's fine using "usage data", but the contents of a private conversation should be considered to be... private.
LLMs and generative image models have shown the ability to leak/reproduce training data. That's a big deal.
Yet...
One of those explanations seems much more likely than the other to everyone, but curiously I think some people will disagree about which side is implausible.
That's a new fun and exciting definition of E2E a lot of people are pushing.
Zoom, like Meet, Teams, WebEx, and many others to my knowledge is "encrypted" but not by default "end-to-end encrypted" in the normal meaning of the term. (Some of these have options for E2EE but it's buried in the service configs and not easy to enable.) So they can and do see audio and video on their servers (as can anyone who breaches their infrastructure) by design. The encryption in this default mode only prevents your ISP from seeing the content of the call.
As a distinction, Signal calls are E2EE -- Signal doesn't see unencrypted video/audio for calls, even ones that are relayed through Signal servers. And even in that case, Signal still knows the participants of the call, just not what is being said.
(As a side note, this is why we built Booth.video -- to demo that this isn't a fundamental tradeoff and it's possible to have E2EE, metadata-secure video conferencing in the browser.)
According to their own claims, right? There's no way for anyone to verify that they're actually E2EE, just Meta's word that it is so.
now i wonder how you did that. Is the key exchange of participants happening out of band?
It’s kind if what it means? OP’s question is w.r.t to the receiving party’s ability to consume the data. The point that’s being made is that E2E doesn’t mean encrypted at rest and receiver can’t consume the data.
I see a lot of comments nitting on the wording for a lack of specificity but, IMO, OP’s question was more about understanding what goes on at the two ends of the pipe. The point being made is that the recipient can still chose to do whatever it is they want with the content.
In context of video conferencing software (WebRTC specifically) this is actually somewhat interesting, because typically the signaling server is the one who hands out the public key of the other peer and needs to be trusted, so they could by all means deliver public keys to which they posses the keys for decryption and it therefore would allow them to play man in the middle in a typically relayed call. So even if E2EE is implemented, it might be done poorly without figuring out how to establish trust independently.
Separately, most Zoom meetings are not E2EE. That’s why features like live transcription work.
I think zoom probably have a defence against the fraud accusation that no reasonable person would believe end to end encrypted meant zoom doesn't have the data as that's the whole point of the service existing.
https://support.zoom.us/hc/en-us/articles/360048660871-End-t...
It seems pretty weird that if your office used Zoom, that you would need to agree to all these terms that aren't part of your employment contract to actually be employed.
i see nothing indicating data wont be provided to third parties.
i see nothing indicating third parties will be prevented from using aquired data to train AI
i see nothing indicating zoom will not aquire trained models from third parties that use Zoom harvested data in training.
There is no such thing as a training model auditor.
> Finally, the company must obtain biennial assessments of its security program by an independent third party, which the FTC has authority to approve, and notify the Commission if it experiences a data breach.
https://www.ftc.gov/news-events/news/press-releases/2020/11/...
https://fortune.com/2023/07/26/pcaob-audit-completely-unnacc...
a lot of things can happen in 2 years though
It's better than nothing (assuming you're still using Zoom).
Our whole system is based on assuming a degree of trust, based on both social norms and reputation of prior interaction, with a confidence of remedy in the case of a failure. If we really had to have much stronger confidence up-front in commercial interactions there would be a lot more friction and overhead in every transaction. Dealing with the occasional fraud seems like a better tradeoff.
New Zoom Subscription Tier:
Virtual agent with perfectly fine tuned domain-specific knowledge performs 99% as well as your sales/support person.
24/7/365
$400 per month
Editor’s note: This blog post was edited on August 11, 2023, to include the most up-to-date information on our terms of service. Following feedback received regarding Zoom’s recently updated terms of service Zoom has updated our terms of service and the below blog post to make it clear that Zoom does not use any of your audio, video, chat, screen sharing, attachments, or other communications like customer content (such as poll results, whiteboard, and reactions) to train Zoom’s or third-party artificial intelligence models.
It’s important to us at Zoom to empower our customers with innovative and secure communication solutions. We’ve updated our terms of service (in section 10) to further confirm that Zoom does not use any of your audio, video, chat, screen-sharing, attachments, or other communications like customer content (such as poll results, whiteboard, and reactions) to train Zoom’s or third-party artificial intelligence models. In addition, we have updated our in-product notices to reflect this.
You can't expect to train AI models without some sort of storage mechanism to train on. If they made a 'ninja edit' to their TOS, does this mean they've also backtracked on their data collection?
Google posted about Federated Learning years ago: https://ai.googleblog.com/2017/04/federated-learning-collabo..., not sure how widely it has caught on though.
Also what if they break the law? Who is monitoring that? If detected, who is enforcing it?
I was more concerned about the wording, which implied they would give themselves the right to use the data to train "AI models" more generally.
I have few problems with them building a better noise cancelling solution for their platform, but lots of problems with them selling it for improving third party surveillance and fingerprinting.
LLMs can regurgitate training data unpredictably so you really can’t have any enterprise data flowing through such a system.
I guess my point is that “their AI models” will very likely include more than noise cancelling before too long. It’s too juicy a dataset to ignore.
Provide me with clear uses, and the ability to withdraw or restrict my data contribution in the event of the company deciding to "expand" to other AI "solutions", and I'll feel respected as a user and allow that specific use of my data for training.
But reply with vagueness giving them a carte-blanche to use my data on anything under the sun, and I'm just gonna look somewhere else and encourage others to do the same.
Now are you willing to abandon the rest of the other companies using your information to train their AI models? (Looking at Google, Microsoft (GitHub), Meta, Instagram, etc)
Now should be the time to self-host then. Whether if it is a GitLab, or Gitea instance for Git, or a typical Mastodon server with a single user that controls the instance for full ownership.
How are you going to serve your content? What is the best tool out there — OwnCloud?
But still it didn't matter, nobody remembers or cares (except me, I refuse to install their native app. The day they fully block browser access will be a bad day).
https://infosecwriteups.com/zoom-zero-day-4-million-webcams-...
Companies are happily exposing all their data to those services, I don't understand why anybody would pretend to be surprised of the results.
Honestly, I don't understand why you wouldn't want the most accurate AI models available. The LLM is only as good as the data set it's trained on, and the more I read about LLM's and the advent of AI evolving from them, the more I'm starting to think if we don't jump both feet into the pool, then we'll never get to the promised land of:
"AI model, spin me up a T-shirt company that's scaled to 10mm users a month, and spin it down after 6mo. if sales don't increase by n% month-over-month"
or
"AI model, get me [A,B,C, ...n] groceries so I can throw a housewarming party on Friday. I can only accept the delivery Tuesday or Thursday. I don't care which store(s) those ingredients come from or how the internals are orchestrated."
What's the threat model here, specifically? What nefarious things would happen by using customer data? Most companies exist to make money, which honestly, is a pretty benign objective, all things considered.
and similar (presumably more sophisticated) exfiltration of commercially valuable information obfuscated away within the language model
An employee can blackmail another person, but the model simply has no reason to, or am I misinterpreting the "whys" of needing privacy here?
The concern isn't judgement from the AI, but that products from the model trained on your data could expose sensitive information.
Since it's never quite clear exactly how the data could be used in situations like this, there's a chance that very sensitive data could be parroted back to people who were not the intended audience.